stackchain-dashboard/docs/ops-fundamentals.md
timmy 1b448fef23
All checks were successful
CI / build-release (pull_request) Successful in 10s
CI / lint (pull_request) Successful in 3m53s
CI / browser-journey (pull_request) Successful in 6m55s
CI / release-candidate (pull_request) Has been skipped
docs: add disk capacity incident runbook (Closes #1287)
2026-08-23 01:10:16 +00:00

1.8 KiB

Release Engineering & Continuous Delivery

  • Treat every commit as a deployable artifact.
  • Use CI to run tests and build artifacts; use CD to push to staging/production.
  • Pin versions, sign artifacts, and audit changes.
  • Release trains > hero deploys: small batches, fast feedback.

11 Factor Apps (12 Factor methodology adapted)

  1. Codebase: one repo per service, tracked in version control.
  2. Dependencies: explicit manifests, lockfiles, isolated environments.
  3. Config: env vars/secrets outside code.
  4. Backing services: attach databases/queues as resources.
  5. Build/run separation: immutable artifacts, separate runtime.
  6. Processes: stateless, share nothing, push state out.
  7. Port binding: self-contained HTTP service.
  8. Concurrency: scale by process model.
  9. Disposability: fast startup, graceful shutdown.
  10. Dev/prod parity: avoid one-off configs.
  11. Logs/logging: treat logs as event streams.

Microservices Architecture

  • bounded context per service
  • async communication where possible
  • API contracts + versioning
  • independent deployability
  • observability: metrics, traces, structured logs

Disk capacity incident

Root filesystem usage at or above 85% is an operations incident.

  1. Capture the initial state with df -h / /var/lib/gitea.
  2. Inspect directory sizes and active processes before cleanup. Reclaim only disposable caches and abandoned temporary environments. Responders must not delete Gitea data, repositories, databases, secrets, or active release artifacts.
  3. Re-run df -h / /var/lib/gitea; usage must be below 85%.
  4. Confirm the release path with systemctl is-active gitea act_runner, then run the repository test suite.

Record the before and after usage in the incident ticket. If safe cleanup cannot restore capacity below 85%, keep the incident open and escalate storage expansion rather than removing durable data.