Ops-dashboard. deployen van repos naar docker
  • TypeScript 57.1%
  • Go 34.6%
  • Shell 5.4%
  • JavaScript 1.7%
  • Python 0.8%
  • Other 0.4%
Find a file
Janpeter Visser e43c8a3744
Some checks failed
CI / Select checks (push) Successful in 19s
CI / Root app checks (push) Failing after 2m1s
CI / Ops-agent checks (push) Successful in 30s
CI / DB access operator (push) Successful in 1m17s
CI / Deploy artifact checks (push) Successful in 20s
CI / Docker image build (push) Successful in 3m15s
CI / Mac foundation hermetic checks (push) Successful in 4m55s
CI / Required checks (push) Failing after 18s
Merge pull request 'test(docker): guard the ← Containers back-link (SP-4)' (#282) from test/sp4-docker-back-link into main
Reviewed-on: #282
2026-10-05 10:46:58 +02:00
.forgejo/workflows ci: reproduceerbaarheidseis Mac-foundation vervalt (besluit JP 2026-09-28, optie C) 2026-09-28 18:41:49 +02:00
.githooks chore(ci): pre-push-hook controleert het CI-testregister 2026-09-29 16:06:39 +02:00
.superpowers fix(IDEA-187): stabilize atomic secret switch proof 2026-08-23 14:52:44 +02:00
app feat(T-201): tekstknoppen in pilvorm, ember op inloggen en koppelen 2026-10-04 15:43:43 +02:00
components feat(T-201): tekstknoppen in pilvorm, ember op inloggen en koppelen 2026-10-04 15:43:43 +02:00
deploy fix(mac): logmapfout laat write niet falen 2026-10-04 20:44:16 +02:00
docs docs(T-202, T-203): routecontrole increment 2 en stylingdocument 2026-10-04 15:49:01 +02:00
hooks fix(control-room): preserve live and dialog state 2026-08-01 16:18:29 +02:00
lib fix(worker-logs): retain detail links after gzip rotation 2026-09-30 14:32:58 +02:00
ops-agent fix(db-access): keep database URLs out of argv (ISS-41) 2026-09-28 14:28:58 +02:00
prisma feat(IDEA-187): T-144 contract 4 — central Mac release projection writer with sequence/head CAS 2026-09-14 21:33:15 +02:00
public style: use ops dashboard icon set 2026-06-10 16:54:25 +02:00
reviews feat: Add DR-drill plan reviews and updates 2026-05-21 23:35:08 +02:00
scripts test(ci): register the docker back-link tests in the base group 2026-10-05 10:20:11 +02:00
test test(docker): render the detail page per container state 2026-10-05 10:18:34 +02:00
test-fixtures fix(T-1775): gate schema writers with pinned policy and shared host lock 2026-09-11 11:59:43 +02:00
vendor chore(vendor): bump scrum4me-copilot 4bef512 -> ca913d2 (M23 kit) 2026-07-09 18:15:23 +02:00
.dockerignore feat: Dockerfile, deploy configs en Caddy-block voor ops.jp-visser.nl 2026-05-13 17:12:37 +02:00
.env.example feat(copilot): catch-all SSE route + 2 user-scoped FlowRun app-tools (Fase 5 B3) 2026-06-12 10:02:25 +02:00
.gitignore feat(IDEA-187): T-145 D10a contract 3b-ii — Go preflight evaluation and jp-cutover-preflight over the shared golden vectors 2026-09-17 00:06:22 +02:00
.gitmodules feat(copilot): vendor @s4m-kit submodule + tsconfig/next.config/zod (Fase 5 B1+B2) 2026-06-12 09:58:56 +02:00
AGENTS.md docs(T-203): stylingregel in AGENTS.md 2026-10-04 15:49:12 +02:00
CLAUDE.md docs: apply agent guide startup contract (T-156) 2026-09-30 19:13:24 +02:00
components.json feat: Next.js + Tailwind + shadcn/ui project skeleton 2026-05-13 16:59:21 +02:00
Dockerfile fix(idea-187): repair PR CI contracts 2026-08-24 08:12:04 +02:00
next.config.ts chore(IDEA-187): recover reviewed implementation content 2026-08-22 19:00:26 +02:00
package-lock.json fix(deploy): enforce immutable MCP release tags 2026-08-24 03:10:52 +02:00
package.json package.json bijwerken 2026-09-21 12:58:50 +02:00
postcss.config.mjs feat: Next.js + Tailwind + shadcn/ui project skeleton 2026-05-13 16:59:21 +02:00
prisma.config.ts feat(worker-logs): persist non-idle runs to Postgres via Prisma 2026-05-17 20:51:52 +02:00
proxy.ts feat(qr-login): serve /m/pair as public shell 2026-06-15 18:14:59 +02:00
README.md docs(audit): README telt zes CI-domeinjobs (PR #269) 2026-09-29 07:43:25 +02:00
recommendation.md docs(IDEA-187): T-145 part A — scope the tailnet comments to their own grant 2026-09-15 14:35:51 +02:00
task-2-fix-rereview.md feat: Add DR-drill plan reviews and updates 2026-05-21 23:35:08 +02:00
tsconfig.json chore(qr-login): revert onnodige tsconfig ES2018-bump (regex /s was redundant) 2026-06-15 17:59:50 +02:00
vitest.ci.base.config.ts ci: give portable, server and Mac tests one owner 2026-09-12 12:42:21 +02:00
vitest.ci.mac.config.ts ci: select measured worker profile and record sprint A evidence 2026-09-12 16:01:45 +02:00
vitest.ci.server.config.ts ci: select measured worker profile and record sprint A evidence 2026-09-12 16:01:45 +02:00
vitest.config.ts fix(IDEA-187): close package and stable MCP integrity 2026-08-23 22:47:09 +02:00
vitest.db-access.config.ts fix(db-access): validate adoption proof target before partial resume 2026-09-22 01:22:14 +02:00
vitest.mac-package.config.ts fix(IDEA-187): close package and stable MCP integrity 2026-08-23 22:47:09 +02:00

Ops Dashboard

Single-user ops dashboard voor jp-visser.nl.

See docs/runbooks/ for setup, deployment, and operational procedures.

CI

Forgejo is the leading forge for this repo. CI lives in .forgejo/workflows/ci.yml and runs on pull requests and pushes to main.

The required merge gate is the CI workflow:

  • root app: npm ci, npx prisma generate with a build-time placeholder DATABASE_URL, three disjoint test groups (npm run test:ci:base, npm run test:ci:server, npm run test:ci:mac, each with its own vitest.ci.*.config.ts; group ownership lives in scripts/ci/test-groups.json), npm run typecheck, npm run build. Locally npm test still runs the full suite in one go.
  • ops-agent: npm --prefix ops-agent ci, npm --prefix ops-agent run check
  • deployment artifacts: parse deploy/workflow YAML
  • production image: docker build -t ops-dashboard:ci .

The workflow also runs a DB access operator job (npx vitest run --config vitest.db-access.config.ts) against a throwaway postgres:17 service, covering the Scrum4Me db-access operator wrapper. It is one of the required status contexts below.

The workflow additionally runs Select checks and Required checks around the six domain jobs. Both are in shadow mode (CI_MODE=shadow): every domain job still runs on every event and the required contexts below are unchanged. See docs/runbooks/ci-selection.md.

The root app and Docker image jobs fetch vendor/scrum4me-copilot over HTTPS with the S4M_COPILOT_READ_TOKEN Actions secret, because @s4m-kit/* resolves from that private submodule. The workflow pins the submodule URL to git.jp-visser.nl before credentials are offered.

Do not add a duplicate .github/workflows/ workflow for the GitHub mirror.

Branch protection

main is protected and requires these exact status contexts before a pull request can be merged:

  • CI / Root app checks (pull_request)
  • CI / Ops-agent checks (pull_request)
  • CI / Deploy artifact checks (pull_request)
  • CI / Docker image build (pull_request)
  • CI / DB access operator (pull_request)

Direct pushes to main are allowed (enable_push = true, no push whitelist; enabled 2026-07-09). Pull requests remain the default route, but git push origin main is no longer rejected by the pre-receive hook.

Be aware of what that means: the five required contexts above only gate the merge of a pull request. A direct push bypasses them entirely — CI still runs on main afterwards, but nothing blocks the push if it fails. Use a pull request whenever you want the gate to actually hold.

To reinstate the hard gate, set enable_push back to false on the main branch protection (Forgejo: repo → Settings → Branches, or PATCH /api/v1/repos/janpeter/Ops-dashboard/branch_protections/main). Note that Forgejo does not exempt site admins from a protected branch.

Installation

Prerequisites

  • Docker + Docker Compose (plugin) installed on the host
  • A PostgreSQL service named postgres already running in the same Compose stack
  • The repository cloned to /srv/scrum4me/ops-dashboard
  • /srv/scrum4me/compose/docker-compose.yml as the shared Compose file

1. Configure environment

cp deploy/ops-dashboard.env.example /srv/scrum4me/ops-dashboard/.env
# Edit /srv/scrum4me/ops-dashboard/.env — set DATABASE_URL, AUTH_SECRET, etc.

2. Install ops-agent

sudo deploy/ops-agent/setup.sh

This creates the ops-agent system user, installs /opt/ops-agent, generates /etc/ops-agent/secret, and enables the systemd unit.

For an existing install, add or update the Docker inspection module:

sudo deploy/ops-agent/install-docker-inspection-module.sh

The installer restarts ops-agent by default. When batching this with a larger setup/update, run it with OPS_AGENT_SKIP_RESTART=1 and restart the service once after all changes are in place.

setup.sh also installs the Caddy write module (ISS-23). For an existing install, add it on its own:

sudo deploy/ops-agent/install-caddy-module.sh

It installs ops-agent/wrappers/caddy/write-caddyfile.sh root-owned under /usr/local/lib/ops-agent/wrappers/caddy/ and adds one exact sudoers rule to /etc/sudoers.d/ops-agent that lets ops-agent run it as root without arguments (the wrapper only reads stdin); the rule is validated with visudo on a candidate file first, and the step is idempotent. The wrapper backs the command_key caddy_write_config used by the Caddy editor and the update_caddy_config flow: it copies the candidate Caddyfile into the running scrum4me-caddy container (docker cp to /tmp) and validates it there — a throwaway container would reject valid configs that reference files only present in the real container (e.g. a tls_trust_pool certificate under /data) — keeps the last 20 versions in /var/backups/caddy (root-only), writes /srv/scrum4me/caddy/Caddyfile in place and asserts that host and container see the same inode and SHA-256 — an atomic rename would leave the single-file bind mount of scrum4me-caddy on the old content. It does not touch commands.yml; the matching keys come from deploy/ops-agent/baseline/commands.yml, so restart ops-agent after updating the live commands.yml.

setup.sh also installs the Scrum4Me DB-access operator module. For an existing install, update it on its own:

sudo deploy/ops-agent/install-db-access-module.sh

It installs /usr/local/lib/ops-agent/wrappers/db-access/, the config dir /etc/ops-agent/db-access/ and the log root /var/log/ops-agent/db-access/. Outside test mode it also installs /etc/tmpfiles.d/scrum4me-schema-lock.conf and runs systemd-tmpfiles --create for it, which provisions the shared schema lock used by both the Prisma and Watch routes. It also writes six exact sudoers rules to /etc/sudoers.d/ops-agent that let ops-agent run policy-bundle-flow.sh prepare, … approve, prisma-operator.sh adoption-precheck, … adopt-idea-213, … adoption-precheck-staged and … adopt-idea-213-staged as root; the rules are validated with visudo on a candidate file before the real file changes, and the step is idempotent. Since ISS-25 the installer also creates the evidence directories /etc/ops-agent/db-access/adoptions/idea-213/, installs /etc/ops-agent/flows/adopt_idea_213.yml and /etc/ops-agent/flows/adopt_idea_213_staged.yml, and merges the adoption_precheck, adopt_idea_213, adoption_precheck_staged and adopt_idea_213_staged blocks into /etc/ops-agent/commands.yml (the previous file is kept as commands.yml.prev); it does not restart the service, so restart ops-agent afterwards to load the new command keys. The credentials file /etc/ops-agent/db-access/scrum4me-prisma.env is not created by the installer — provision it separately (root:ops-agent 0640) from deploy/ops-agent/prisma-migrator.env.example.

The command keys prisma_migrate_deploy, prisma_migrate_status and prisma_migrate_precheck route through that wrapper, so Prisma migrations for Scrum4Me web run on isolated migrator credentials instead of the web runtime's. Four further keys run the same wrapper as root for the fixed IDEA-213 adoption: adoption_precheck (read-only) and adopt_idea_213 against the active bundle, and adoption_precheck_staged (read-only) and adopt_idea_213_staged against the staged bundle from pending-policy.env — see docs/runbooks/db-access-policy-bundle.md. The wrapper checks the policy hash before the first database call and verifies the migrator identity next; precheck runs before and apply after migrate deploy. Since Release II, POLICY_HASH in the env file is mandatory rather than an opt-in gate: both prisma_migrate_precheck and — outside test mode — prisma_migrate_deploy refuse to run without it, and a production POLICY_HASH additionally requires an explicit immutable POLICY_ROOT. Since PBI-164 the deploy flows prepare that bundle themselves: redeploy_all and update_scrum4me_web run prepare_policy_bundle after the pull, which switches POLICY_ROOT over automatically when the policy hash is unchanged and otherwise stops with exit 3 before any database action, pending the approve_policy_hash flow. See docs/runbooks/db-access-policy-bundle.md for bundle provisioning and the host-activation window.

Docker inspection uses /etc/ops-agent/container-inspection.yml as the allowlist for containers, expected env keys, and config file paths. Standard views show missing expected keys, redacted logs, and redacted config content; raw env values and config files are only returned by explicit reveal actions.

Copy the generated secret into the web-app env file:

sudo cat /etc/ops-agent/secret
# Paste the value as OPS_AGENT_SECRET= in /srv/scrum4me/ops-dashboard/.env

3. Build and start the dashboard

sudo docker compose -f /srv/scrum4me/compose/docker-compose.yml build ops-dashboard
sudo docker compose -f /srv/scrum4me/compose/docker-compose.yml up -d ops-dashboard

The dashboard is now reachable on 127.0.0.1:3001 (proxied by Caddy).

4. Install the self-update script

sudo deploy/ops-dashboard-updater/install.sh

To enable scheduled updates (daily at 03:00):

sudo systemctl enable --now ops-dashboard-updater.timer

To trigger a manual update via SSH:

sudo systemctl start ops-dashboard-updater.service
# or:
sudo /opt/ops-dashboard-updater/update.sh

Never trigger updates through the dashboard UI — the script restarts the container that serves the UI.

5. Register in the fleet (heartbeat)

Keep this host visible/"online" in the shared fleet registry via a systemd timer that POSTs to the in-container heartbeat endpoint (the containerized app can't run the ts-node CLI heartbeat used on native hosts like the Mac):

sudo cp deploy/ops-dashboard-heartbeat/ops-dashboard-heartbeat.service /etc/systemd/system/
sudo cp deploy/ops-dashboard-heartbeat/ops-dashboard-heartbeat.timer   /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now ops-dashboard-heartbeat.timer
sudo systemctl start ops-dashboard-heartbeat.service   # register once now

The endpoint reads INSTANCE_SLUG and INSTANCE_BASE_URL (and optional INSTANCE_CAPABILITIES, INSTANCE_HOSTNAME, INSTANCE_TAILSCALE_IP, APP_VERSION) from the container env and authenticates with OPS_INGEST_SECRET (the same secret the worker-logs ingest uses). In a container os.hostname() is the container id — set INSTANCE_HOSTNAME (e.g. scrum4me-server) in the env for a readable name.

Configuration

File Purpose
/srv/scrum4me/ops-dashboard/.env Web-app environment (DATABASE_URL, AUTH_SECRET, OPS_AGENT_SECRET, …)
/etc/ops-agent/secret Shared HMAC secret between web-app and ops-agent
/etc/ops-agent/commands.yml Whitelist of commands the ops-agent may run
/etc/ops-agent/container-inspection.yml Allowed containers, env keys and configfile paths for Docker inspection/reveal; review and customize before use
/etc/ops-agent/db-access/scrum4me-prisma.env Migrator credentials + policy config for the Scrum4Me DB-access operator (root:ops-agent 0640). Credentials are provisioned by hand; POLICY_ROOT/POLICY_HASH are rewritten by policy-bundle-flow.sh during the deploy flow
/etc/ops-agent/flows/ Flow YAML files (backup, caddy reload, etc.)
/srv/scrum4me/compose/docker-compose.yml Main Compose file (ops-dashboard service is part of this stack)
deploy/ops-agent/commands.darwin.yml macOS command whitelist (git, processes, MCP, S4M Hub-hooks) for the Mac instance
deploy/ops-dashboard-heartbeat/ systemd timer + service that POST the fleet heartbeat (keeps this host "online")

Forgejo release-candidate polling

Release-candidate verification uses three distinct server-only credentials:

  • FORGEJO_RELEASE_WEBHOOK_SECRET is the webhook HMAC verification secret in /srv/scrum4me/ops-dashboard/.env.
  • FORGEJO_RELEASE_READ_TOKEN is a least-privilege Forgejo API read token in the same environment file. It must not be the webhook secret.
  • OPS_RELEASE_CANDIDATE_POLLER_SECRET is the dashboard-side credential for the local systemd poller endpoint. Its matching protected host header file is /etc/ops-dashboard/release-candidate-poller-authorization; it contains the corresponding Authorization header and is readable only by the account running the timer.

The dashboard never renders, persists or logs any of these values. Install the periodic poller after configuring the application environment:

sudo cp deploy/release-candidate-poller/release-candidate-poller.{service,timer} /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now release-candidate-poller.timer

Ops-agent auth

The web-app communicates with the ops-agent via a shared secret stored in /etc/ops-agent/secret (mode 0640, owner root:ops-agent).

  • The ops-agent reads the secret at startup via OPS_AGENT_SECRET_PATH.
  • Every request from the web-app carries Authorization: Bearer <secret>.
  • The agent validates using a constant-time comparison to prevent timing attacks.
  • The web-app reads the secret value from the OPS_AGENT_SECRET environment variable.

Fail-closed behavior

The agent refuses to start when no secret file is present at OPS_AGENT_SECRET_PATH (it logs the reason and exits non-zero). For local development without a secret, set OPS_AGENT_ALLOW_INSECURE=true to explicitly run in an unauthenticated mode — never use this in a deployed environment. Even at runtime, a missing secret makes every request return 503 (not a silent bypass).

Secret rotation procedure

  1. Generate a new secret on the server:
    openssl rand -hex 32 | sudo tee /etc/ops-agent/secret
    sudo chown root:ops-agent /etc/ops-agent/secret
    sudo chmod 0640 /etc/ops-agent/secret
    
  2. Update OPS_AGENT_SECRET in the web-app's environment file (/srv/scrum4me/ops-dashboard/.env) with the new value.
  3. Restart both services:
    sudo systemctl restart ops-agent
    sudo docker compose -f /srv/scrum4me/compose/docker-compose.yml restart ops-dashboard
    
  4. Verify the dashboard is operational and that systemctl status ops-agent shows the service running without errors.

Mac production host

An independent, hardened Mac production host is prepared under deploy/mac-production/ (source only — nothing there mutates a host when checked out). The reviewed operator cutover is driven by the fixed-verb deploy/mac-production/cutover.sh state machine (25 states, 21 verbs, one generated transition table), shipped to the Mac as a digest-verified cutover kit built by deploy/mac-production/build-cutover-kit.sh, and documented step by step in docs/runbooks/ops-voor-mac-cutover.md. D10b landed the root side — the root observer jp-cutover-observe, the root-mutating jp-cutover with admission, preflight, activator sync and compensation, and the verb recipes — so the blanket CUTOVER_OBSERVER_REQUIRES_D10B refusal is gone: cutover.sh now runs every verb through the closed verb→command table. That makes the verbs runnable, not run — until the operator gate opens, cutover.sh kit-verify is still the only command an operator runs. The live cutover (G1/G2) runs only under its own explicit operator gates in Tasks 14–15; this repository ships the closed contracts, the expected tailnet grants and the read-only database proofs, all unit-tested and mutation-free.