No description
  • TypeScript 74.2%
  • Python 10.3%
  • Shell 9.8%
  • JavaScript 4.4%
  • Dockerfile 1.3%
Find a file
Janpeter Visser 7e4eafe2b1
All checks were successful
CI / Compose config (push) Successful in 5s
CI / Build-arg coverage (push) Successful in 4s
CI / Docker build (push) Successful in 7s
Merge pull request 'fix(dispatch): verhoog supervisor mem_limit naar 512m (OOM-wees-job)' (#109) from fix/dispatch-supervisor-mem-512m into master
2026-10-05 03:38:52 +02:00
.forgejo/workflows fix(ops-agent): ops-dashboard bouwen met de vereiste build-args, plus een gate 2026-09-10 02:09:32 +02:00
__tests__ test(dispatch): reset-test van de inspect-grens deterministisch (review #105) 2026-10-03 18:14:16 +02:00
bin feat(dispatch): usage van een poging meesturen met het resultaat (T-1972) 2026-10-03 14:46:48 +02:00
codex chore(codex-config): spawn de MCP zonder npx-wrapperlaag 2026-07-16 11:33:26 +02:00
deploy fix(dispatch): verhoog supervisor mem_limit naar 512m 2026-10-05 03:12:01 +02:00
docs test(dispatch): inspect-tolerantie volledig gedekt + gedocumenteerd (review #104) 2026-10-03 18:04:54 +02:00
etc fix: lokale Docker build werkend krijgen 2026-05-02 19:18:35 +02:00
lib fix(dispatch): tijdelijke unknown-inspect maakt een poging niet direct uncertain (T-1973) 2026-10-03 17:53:46 +02:00
scripts feat(max2): worker-redeploy-flows wachten op healthy containers 2026-09-26 11:12:27 +02:00
skills docs: remove trailing blank line from worker review 2026-09-30 13:27:51 +02:00
tests feat(dispatch): usage van een poging meesturen met het resultaat (T-1972) 2026-10-03 14:46:48 +02:00
vendor feat(dispatch): __dispatch_input als bron voor het kind (M41 T7) 2026-10-02 14:52:28 +02:00
.env.deploy.example chore(cli): pin claude-code 2.1.287 en codex 0.160.0 2026-10-02 22:25:19 +02:00
.env.example chore(cli): pin claude-code 2.1.287 en codex 0.160.0 2026-10-02 22:25:19 +02:00
.gitattributes fix: lokale Docker build werkend krijgen 2026-05-02 19:18:35 +02:00
.gitignore feat(IDEA-213): supervise isolated job and host attempts 2026-09-15 14:31:28 +02:00
.gitmodules feat(IDEA-213): supervise isolated job and host attempts 2026-09-15 14:31:28 +02:00
AGENTS.md docs: clarify worker guide handoff (ST-051) 2026-09-29 21:14:44 +02:00
CLAUDE.md docs: clarify worker guide handoff (ST-051) 2026-09-29 21:14:44 +02:00
docker-compose.yml chore(cli): pin claude-code 2.1.287 en codex 0.160.0 2026-10-02 22:25:19 +02:00
Dockerfile chore(cli): pin claude-code 2.1.287 en codex 0.160.0 2026-10-02 22:25:19 +02:00
Dockerfile.dispatch feat(dispatch): model-claude image stages, probe stub and operator docs 2026-10-03 00:23:26 +02:00
mcp-config.json chore(mcp-config): spawn de MCP zonder npx-wrapperlaag 2026-07-14 17:16:01 +02:00
package-lock.json feat(IDEA-213): supervise isolated job and host attempts 2026-09-15 14:31:28 +02:00
package.json test(dispatch): live review smoke voor M41 U7 2026-10-02 16:27:08 +02:00
README.md docs(audit): model-claude proxy proof in README (PR #102) 2026-10-04 07:32:48 +02:00
tsconfig.dispatch.json test(dispatch): live review smoke voor M41 U7 2026-10-02 16:27:08 +02:00

scrum4me-agent-runner

Headless Claude Code worker die de Scrum4Me job-queue (M13) leegtrekt vanaf een QNAP NAS via Container Station. Geen Vercel, geen browser, geen toetsenbord — Claude Code draait als daemon, claimt jobs uit mcp__scrum4me__wait_for_job, voert ze uit in een per-job clone, en pusht nooit zelf.

Architectuur in één plaatje

┌─ QNAP TS-664 (Container Station) ─────────────────────────────┐
│                                                                │
│  ┌─ container: agent-runner ────────────────────────────────┐  │
│  │  PID 1: tini → run-agent.sh (daemon-loop)                │  │
│  │            ├─ health-server.js  (8080 → host 18080)      │  │
│  │            └─ claude -p (per-batch, met MCP via stdio)   │  │
│  │                  └─ scrum4me-mcp → Neon Postgres         │  │
│  │                                                          │  │
│  │  /tmp/job-<id>      ephemeral working trees              │  │
│  │  /var/cache/repos   bare git mirrors  (volume)           │  │
│  │  /var/cache/npm     npm cache         (volume)           │  │
│  │  /var/log/agent     run + job logs    (volume)           │  │
│  └──────────────────────────────────────────────────────────┘  │
│                                                                │
│  /share/Agent/cache  /share/Agent/logs  /share/Agent/state     │
└────────────────────────────────────────────────────────────────┘
                              │
                              ▼  HTTPS
                   Neon Postgres (Scrum4Me DB)
                              ▲
                              │
                   Vercel ─── Scrum4Me UI (gebruikers enqueueen jobs)

Eén claude -p-invocation roept intern wait_for_job aan totdat de queue leeg is (≈600 s lege block-time → afsluiten). De wrapper start claude -p opnieuw zodra hij eindigt, met exponentiële backoff bij fouten.

Wat zit waar

Bestand Doel
Dockerfile Ubuntu 22.04 + Node 22 + Claude Code + scrum4me-mcp + scripts
docker-compose.yml Service-definitie, volumes, env-file, restart-policy, limits
package.json Npm-dependencies van de runner zelf (alleen scrum4me-mcp pin)
mcp-config.json Claude Code MCP-config (verwijst stdio naar scrum4me-mcp)
CLAUDE.md Agent-rol-instructies, auto-geladen door claude -p
bin/entrypoint.sh Container-startup: dirs, health-server, daemon-loop
bin/install-skills.sh Container-startup: publiceert de APPROVED skills-inventory naar de discovery-mappen van Claude/Codex/Agents
bin/run-agent.sh Daemon-loop met backoff, exit-code-routing en state-writes
bin/check-tokens.sh Pre-flight: Scrum4Me-token, Claude OAuth-token, DB-bereikbaarheid, DB-rolcapaciteiten
bin/check-database-role.cjs Pre-flight-helper: weigert verhoogde PostgreSQL-rollen op de aanwezige worker-DB-URL's
bin/job-prepare.sh Per-job: bare-fetch + clone-via-reference naar /tmp/job-<id>
bin/job-cleanup.sh Per-job: logs naar /var/log, working tree weg
bin/health-server.js HTTP-endpoint op 8080 (intern) dat state.json en marker-files leest
bin/rotate-logs.sh Compress/cleanup van oude .log-bestanden
.env.example Alle env-vars met uitleg

Vereisten op de NAS

  • Container Station 2+ (Docker compose v2)
  • Agent als QTS Shared Folder op een echte volume (bv. CACHEDEV1_DATA). Niet een mkdir /share/Agent — /share zelf is een 16 MB tmpfs en handmatige directories overleven geen reboot. Aanmaken via Control Panel → Privilege → Shared Folders → Create. QTS legt dan automatisch de symlink /share/Agent → /share/CACHEDEV1_DATA/Agent.
  • Drie subdirs onder die share: /share/Agent/cache, /share/Agent/logs, /share/Agent/state. Aanmaken via File Station of via SSH na share-creatie.
  • Internet-uitgang naar api.anthropic.com, git.jp-visser.nl (Forgejo HTTPS-clone/push), cli.github.com (build-time voor de gh CLI), je Neon-host, registry.npmjs.org.

Verifieer vóór je deployt dat /share/Agent echt op disk staat:

ssh admin@<nas> 'ls -la /share/ | grep Agent; df -h /share/Agent'

Verwacht een symlink (l...Agent -> /share/CACHEDEV1_DATA/Agent) en een df-uitvoer met TB-grootte op cachedev1/cachedev2. Als je hier tmpfs 16M ziet, is de share geen geregistreerde QTS Shared Folder en zal elke transfer

16 MB falen met scp: write remote ... Failure.

Deploy

# 1. Op je werkstation: token's regelen
#    a. CLAUDE_CODE_OAUTH_TOKEN  →  draai `claude setup-token` (browser-flow)
#    b. SCRUM4ME_TOKEN           →  log in als de dedicated agent-user in
#                                   Scrum4Me, /settings/tokens, label "NAS-runner"
#    c. DATABASE_URL/DIRECT_URL  →  Neon dashboard
#    d. GH_TOKEN                 →  Forgejo → avatar → Settings →
#                                   Applications → Generate New Token; scope
#                                   minimaal `write:repository` op de twee
#                                   repos (janpeter/Scrum4Me + janpeter/
#                                   scrum4me-mcp). Wordt gebruikt voor clone
#                                   en push naar Forgejo. PBI-86 (hybride
#                                   model): `gh pr create` is uit de
#                                   worker-flow verwijderd — de GitHub-PR
#                                   komt via de handmatige promote-Action
#                                   in Forgejo.

# 2. Repo op de NAS plaatsen
ssh admin@nas
cd /share/Agent
git clone https://git.jp-visser.nl/<jij>/scrum4me-agent-runner.git
cd scrum4me-agent-runner

# 3. Env aanmaken
cp .env.example .env
chmod 600 .env
vi .env   # vul alle waarden in

# 4. Build + start
docker compose build
docker compose up -d

# 5. Verifiëren
curl http://nas.local:18080/health
docker compose logs -f

QNAP-port: host-poort 8080 is bezet door de QTS-webinterface; daarom mapt deze stack standaard 18080:8080. Override via AGENT_HEALTH_PORT_HOST in .env als je een andere host-poort wilt.

Snelle redeploy — bin/deploy-to-nas.sh

Voor een bestaande deploy die je opnieuw wil bouwen + deployen (bijvoorbeeld na een merge in scrum4me-mcp of een aanpassing aan CLAUDE.md):

# Eenmalig: NAS-target instellen
cp .env.deploy.example .env.deploy
vi .env.deploy   # zet NAS_HOST=admin@<nas>

# Pin in .env exact één bestaande immutable MCP-release
vi .env          # zet MCP_GIT_REF=deploy/<immutable-release-tag>

# Daarna: één commando voor de hele cyclus
bin/deploy-to-nas.sh

Het script verifieert eerst de persistente MCP_GIT_REF uit .env zonder het bestand te sourcen. main, kale SHA's, dubbele assignments en onbekende tags stoppen vóór de build. Daarna doet het:

  1. docker buildx build --platform linux/amd64 --load
  2. docker save | gzip → scrum4me-agent-runner-amd64.tar.gz
  3. scp van tarball + docker-compose.yml + (eerste keer) .env naar NAS
  4. ssh op NAS: docker load + sanity-check op .env + docker compose up -d --force-recreate
  5. docker compose logs -f — lokaal-volgbaar terwijl pre-flight + eerste batch starten

.env op de NAS wordt niet overschreven als 'ie er al staat. Bij een verse NAS-installatie wordt 'ie wél geüpload + ge-sed't (NAS_BASE, AGENT_UID, etc.). Voor de eerste deploy of een schoon volume zie de volledige procedure hieronder.

Deploy — cross-build vanaf Mac (Apple Silicon → amd64-NAS)

Alternatief voor de in-place build hierboven. Bouw de image op je Mac voor linux/amd64, schrijf 'm naar een tarball, transfer naar de NAS en laad daar. Handig als de NAS te langzaam is om te builden (npm install op een N5095 met NAS-storage is traag) of als je geen git push wilt voor elke iteratie.

Vóór je begint: controleer dat /share/Agent een echte QTS Shared Folder is (zie Vereisten op de NAS). Dat is de meest voorkomende val.

1. Tokens en .env op je Mac

Zelfde tokens als in Deploy. Hou er rekening mee dat je twee .env-bestanden kunt willen: één voor lokaal Mac-testen (AGENT_PLATFORM=linux/arm64, AGENT_UID=501, paths onder /Users/...) en één voor NAS-runtime. De NAS-versie wordt bij stap 4 ge-scp'd en met sed geschikt gemaakt.

2. Image bouwen voor amd64

cd /Users/<jij>/Development/scrum4me-docker

docker buildx build \
  --platform linux/amd64 \
  --build-arg MCP_GIT_REF=deploy/<immutable-release-tag> \
  --build-arg CLAUDE_CODE_VERSION=2.1.287 \
  --build-arg AGENT_UID=1000 \
  --build-arg AGENT_GID=1000 \
  -t scrum4me-agent-runner:local \
  --load \
  .

AGENT_UID/GID=1000 zijn de NAS-UIDs, niet je Mac-UIDs (vaak 501/20). Bij verkeerde UIDs kan de container niet schrijven naar de bind-mounts.

Verifieer architectuur:

docker image inspect scrum4me-agent-runner:local --format '{{.Architecture}}'
# verwacht: amd64

3. Image naar tarball

docker save scrum4me-agent-runner:local | gzip > scrum4me-agent-runner-amd64.tar.gz
shasum -a 256 scrum4me-agent-runner-amd64.tar.gz

De tarball is ~580 MB voor de huidige image-size. Hij staat in .gitignore, dus geen risico op accidenteel committen.

4. NAS voorbereiden + bestanden transferren

# .env naar de NAS (de Mac-versie; we patchen path-velden zo direct)
scp .env admin@<nas>:/share/Agent/scrum4me-agent-runner/.env

# image-tarball + bijgewerkte compose
scp scrum4me-agent-runner-amd64.tar.gz docker-compose.yml package.json README.md \
    admin@<nas>:/share/Agent/scrum4me-agent-runner/

Zet de NAS-specifieke waarden in .env:

ssh admin@<nas> "
  source /etc/profile
  cd /share/Agent/scrum4me-agent-runner
  sed -i \
    -e 's|^NAS_BASE=.*|NAS_BASE=/share/Agent|' \
    -e 's|^AGENT_BASE=.*|AGENT_BASE=/share/Agent|' \
    -e 's|^AGENT_PLATFORM=.*|AGENT_PLATFORM=linux/amd64|' \
    -e 's|^AGENT_UID=.*|AGENT_UID=1000|' \
    -e 's|^AGENT_GID=.*|AGENT_GID=1000|' \
    -e 's|^AGENT_HEALTH_PORT_HOST=.*|AGENT_HEALTH_PORT_HOST=18080|' \
    .env
  chmod 600 .env
"

source /etc/profile is verplicht in non-interactieve ssh op QNAP. Zonder die regel staat docker niet op $PATH (Container Station's binary zit onder /share/CACHEDEV*_DATA/.qpkg/container-station/...) en faalt elk docker-commando met command not found.

5. Image laden + container recreaten

ssh admin@<nas> "
  source /etc/profile
  set -e
  cd /share/Agent/scrum4me-agent-runner

  # checksum-verificatie (optioneel maar aan te raden)
  sha256sum scrum4me-agent-runner-amd64.tar.gz

  # image laden — overschrijft bestaande :local tag
  gunzip -c scrum4me-agent-runner-amd64.tar.gz | docker load

  # container recreate; --no-build voorkomt onbedoelde NAS-side build
  docker compose up -d --no-build --force-recreate agent

  docker compose ps
  curl -fsS http://localhost:18080/health | head -c 800
"

Verwacht in de health-output cache_free_bytes met een groot getal (TB-orde) — dat is je signaal dat /var/cache op echte disk zit. Een waarde rond 16 MB (16777216) betekent dat je per ongeluk nog steeds op tmpfs draait.

Updaten (handmatig, bewust)

SCRUM4ME_TOKEN of CLAUDE_CODE_OAUTH_TOKEN rouleer je via een rebuild:

cd /share/Agent/scrum4me-agent-runner
git pull
vi .env                   # nieuwe waarden
docker compose build      # nieuwe scrum4me-mcp-versie als dat veranderd is
docker compose up -d

Dezelfde flow voor schema-drift in scrum4me-mcp: maak eerst een immutable deploy/*-tag, pin die als MCP_GIT_REF in .env, verifieer hem met scripts/verify-mcp-release-ref.sh .env, en rebuild. Een kale commit-SHA werkt niet: de Dockerfile gebruikt git clone --branch.

CLAUDE_CODE_VERSION — pin van de Claude Code CLI

De claude-stage van de Dockerfile installeert @anthropic-ai/claude-code via npm (sinds 2026-06-08, eerder via claude.ai/install.sh). De ARG is gepinned op 2.1.287 (npm latest per 2026-10-02; claude-opus-5-5 eist

= 2.1.280, dus stable mag pas als die daar voorbij is) — bewuste bumps gaan via een PR die de Dockerfile-ARG, .env.example én deze paragraaf simultaan updatet zodat verse rebuilds reproducerbaar blijven.

# huidige dist-tags bekijken
npm view @anthropic-ai/claude-code dist-tags

# bump naar een specifieke versie via .env
sed -i 's|^CLAUDE_CODE_VERSION=.*|CLAUDE_CODE_VERSION=2.1.287|' .env
docker compose build agent
docker compose up -d --force-recreate agent

De claude-build verifieert de geïnstalleerde versie: bij een exacte numerieke pin controleert grep -F dat het binary die versie rapporteert; bij CLAUDE_CODE_VERSION=latest (of een andere non-numerieke waarde) volstaat een niet-lege versistring. Een typo of niet-bestaande npm-versie faalt zo alsnog hard.

PRISMA_CLI_VERSION — pin van de Prisma-CLI in de mcp-clone-laag

De mcp-clone-laag draait na npm ci --omit=dev een npx prisma generate. --omit=dev verwijdert de prisma-CLI (die is een devDependency van scrum4me-mcp), waarna npx de CLI ophaalt van npm. Ongepind pakt npx de dist-tag latest, en die staat sinds 2026-08 op een release candidate (8.0.0-rc.12) die geen generate-commando meer kent — waardoor élke verse worker-build hard afbrak met CLI.UNKNOWN_COMMAND. De Dockerfile pint daarom via ARG PRISMA_CLI_VERSION=7.8.0 en roept npx prisma@${PRISMA_CLI_VERSION} generate aan. 7.8.0 is dezelfde versie als @prisma/client in scrum4me-mcp's package-lock.json; CLI en client horen dezelfde major/minor te hebben. Bump deze pin samen met scrum4me-mcp wanneer daar de prisma-versie wijzigt.

Wijzigingen in docker-compose.yml (volumes, tmpfs, env_file, ports)

Let op: docker compose restart herstart alleen het proces in de bestaande container met de oude config. Wijzigingen in volumes, tmpfs-mounts, env_file of ports worden daarmee niet doorgevoerd.

Gebruik altijd --force-recreate als je docker-compose.yml is veranderd:

docker compose up -d --force-recreate agent

Verifieer daarna dat /var/cache op de NAS-overlay staat en niet op tmpfs:

docker exec scrum4me-agent df -h /var/cache
# Verwacht: Filesystem op /dev/mapper/cachedev* of een NAS-share
# Fout:     tmpfs  16M  ... (dan is force-recreate niet uitgevoerd)

Health-endpoint

GET http://<nas>:18080/health retourneert:

{
  "status": "running",            // running | idle | unhealthy | token-expired
  "lastBatchAt": "2026-05-01T12:34:56Z",
  "lastBatchExit": 0,
  "consecutiveFailures": 0,
  "tokenStatus": { "anthropic": "ok", "scrum4me": "ok", "db": "ok" }
}

HTTP-status: 200 als running/idle, 503 bij token-expired of als de laatste heartbeat ouder is dan 5 minuten.

Ook 503 met {"status":"unhealthy","reason":"tmp-low","tmpUsedPct":<n>} zolang de marker TMP_PRESSURE in de state-dir staat. bin/tmp-sweep.sh draait vanuit run-agent.sh tussen iteraties (bij de start, na een overload en na elke iteratie). Het wist alles van de agent-gebruiker direct onder /tmp, behalve tsx-*, node-compile-cache en verborgen entries. Zit /tmp daarna nog op of boven AGENT_TMP_WARN_PCT (default 80), dan schrijft het de marker.

Filesystem-grenzen

De agent-user heeft geen SSH-keys en geen toegang tot andere shares dan /share/Agent/*. Wel een ~/.git-credentials met de GH_TOKEN voor HTTPS-clone/push (zie volgende sectie) — die token is scoped tot de twee configured repos en mag worden gerouleerd door rebuild + redeploy.

Repo bootstrap (clone-on-start)

Bij elke container-start runt bin/repo-bootstrap.sh (als de agent-user, ná drop-privileges) en zet zo'n setup neer:

  1. Configureert git's credential-helper met GH_TOKEN zodat git clone/push naar https://git.jp-visser.nl/... (Forgejo) zonder prompt werkt.
  2. Voor elke repo in GH_PRECLONE_REPOS (komma-gescheiden owner/name):
    • Bestaat ~/Projects/<name>/.git al? → git fetch origin --prune
    • Anders → fresh git clone

Daarna vindt scrum4me-mcp's resolveRepoRoot (in wait_for_job) de clone via z'n convention-fallback ~/Projects/<name>/.git. Worktrees voor jobs landen vervolgens onder ~/.scrum4me-agent-worktrees/<jobId>/ zodat de hoofd-clone niet wordt aangeraakt.

Push gaat over dezelfde token: git push -u origin feat/story-<id> slaagt zonder prompt. gh pr create is in PBI-86 (T-1005) verwijderd uit de worker-flow — de GitHub-PR ontstaat via een handmatig getriggerde promote-Action in Forgejo (zie de Scrum4Me-repo docs/runbooks/forgejo-hybrid-flow.md).

Veelvoorkomende issues

Verzameld uit echte deploy-sessies op een TS-664 met QTS 5.x. Eerst checken voor je gaat tweaken.

/share/Agent is tmpfs 16M — scp faalt op grote files

Symptoom:

scp: write remote ".../scrum4me-agent-runner-amd64.tar.gz": Failure

Eerst 100 % bij kleine files, dan blijvend "Failure" zodra een file de tmpfs vol maakt.

Oorzaak: /share/Agent is geen geregistreerde QTS Shared Folder maar een gewone directory in de root-tmpfs. Na elke reboot is alle inhoud weg en is /share/Agent slechts een entry in de 16 MB /share-tmpfs.

Fix: maak Agent aan via Control Panel → Privilege → Shared Folders → Create. Daarna verschijnt automatisch lrwxrwxrwx /share/Agent → /share/CACHEDEV1_DATA/Agent. Als die symlink ontbreekt omdat er al een tmpfs-directory /share/Agent staat:

ssh admin@<nas> "
  source /etc/profile
  docker stop scrum4me-agent 2>/dev/null
  docker rm scrum4me-agent 2>/dev/null
  rm -rf /share/Agent
  ln -s /share/CACHEDEV1_DATA/Agent /share/Agent
  mkdir -p /share/Agent/{cache,logs,state,scrum4me-agent-runner}
  ls -la /share/ | grep Agent  # moet nu een symlink tonen
"

rm -rf /share/Agent is veilig zolang /share/Agent nog tmpfs is — alle content stond in RAM en zou de volgende reboot toch verdwenen zijn. Maar verifieer eerst met df -h /share/Agent dat 't echt tmpfs is.

docker: command not found in non-interactieve ssh

Symptoom:

ssh admin@<nas> 'docker ps'
# sh: line 1: docker: command not found

Oorzaak: QNAP's admin-user heeft docker op $PATH via /etc/profile.d/*.sh van Container Station. Login-shells laden die scripts; non-interactieve ssh user@host 'cmd' doet dat niet.

Fix: source /etc/profile aan het begin van je remote command:

ssh admin@<nas> "source /etc/profile && docker ps"

yaml: control characters are not allowed

Symptoom:

docker compose down
# yaml: control characters are not allowed

Oorzaak: een eerdere scp van docker-compose.yml faalde halverwege en liet een file met NUL-bytes / partial writes achter. De file-grootte klopt maar de inhoud is corrupt.

Fix: scp opnieuw vanaf je werkstation. Voor het stoppen van een al- draaiende container heb je de yml niet nodig:

docker stop scrum4me-agent && docker rm scrum4me-agent

Daarna scp docker-compose.yml admin@<nas>:/share/Agent/scrum4me-agent-runner/ en pas dan docker compose up -d --no-build --force-recreate.

QTS-console-menu Python-traceback bij ssh-commando

Symptoom: lange Python-traceback over consolemenu_q/prompt_utils.py:31 en Inappropriate ioctl for device, gevolgd door je daadwerkelijke output.

Oorzaak: QTS' admin-shell start een Python-TUI (qts-console-menu) die crasht zonder echte TTY. Niet-fataal — je commando wordt alsnog gerund.

Fix: negeer de traceback. Als je 'm écht weg wilt, gebruik ssh -tt admin@<nas> '...' voor een geforceerde pseudo-TTY, maar pas op: sommige scripts hangen daarop omdat ze interactie verwachten.

.env is weg na rm -rf /share/Agent

Symptoom: na het opruimen van een tmpfs-/share/Agent weigert docker compose up:

env file /share/Agent/scrum4me-agent-runner/.env not found

Oorzaak: .env zit in .gitignore en wordt nooit door git pull of scp . automatisch teruggezet. Bij het wegnuken van een corrupte share verdwijnt-ie mee.

Fix: scp 'm vanaf je werkstation, en patch path-velden op de NAS-zijde (zie Deploy — cross-build stap 4). Zorg dat je .env op je werkstation alle secrets bevat — vooral SCRUM4ME_TOKEN, dat in een Mac-only .env makkelijk ontbreekt omdat lokale tests soms zonder Scrum4Me-API draaien.

Environment variables

WORKER_CAPABILITY (optional during rollout, required after Phase C)

Hardware tier of this host. The MCP claim-filter biases jobs to higher-tier idle workers (pure preference race, no waiting). Set per host in .env:

  • HIGH_P — top performance (current: max2)
  • MEDIUM_P — mid performance (current: mac)
  • LOW_P — entry performance (current: scrum4me-server)

See design.

INTERNAL_PUSH_URL + INTERNAL_PUSH_SECRET (optioneel — web-push, PBI-55)

push-trigger.ts in de MCP-subprocess POST't fire-and-forget naar <app>/api/internal/push/send bij ask_user_question (o.a. grill-vragen) en update_job_status naar done/failed. Beide vars leeg of afwezig = feature stil uit. De URL moet wijzen naar de app-instantie die dezelfde database gebruikt als deze worker (anders bestaan userId/push-subscriptions daar niet); het secret moet byte-gelijk zijn aan INTERNAL_PUSH_SECRET in de app-env. Doorgifte aan het MCP-proces is per runner expliciet geregeld: mcp-config.json (Claude, ${VAR:-}-expansie) en codex/config.toml env_vars (Codex) — alleen .env vullen is dus niet genoeg ná wijziging van die bestanden; rebuild vereist omdat beide in het image gebakken worden.

AGENT_TRANSCRIPT_RETENTION_DAYS (optioneel, default 7)

entrypoint.sh installeert de user-settings van Claude Code en roept daarna bin/set-transcript-retention.sh aan. Dat script zet cleanupPeriodDays in /home/agent/.claude/settings.json, zodat Claude Code transcripts onder ~/.claude/projects die ouder zijn dan dit aantal dagen zelf opruimt. Alleen positieve gehele getallen zijn geldig. Bij ongeldige invoer blijft het bestand ongewijzigd en logt de entrypoint een WARN; de worker start gewoon.

Queue dispatch rollout (IDEA-213)

This repository supplies the managed runtime side of the central dispatch service: the operator-owned broker (bin/run-dispatch-runtime-broker.ts), the egress proxy and the supervisor image built from Dockerfile.dispatch. deploy/queue-dispatch.compose.yml is an opt-in template behind the profile dispatch-not-activated; read docs/dispatch-runtime-operator.md first, it holds the trust model, the egress layout and the lifecycle this section only points at.

Nothing here is activated. The supervisor entrypoint no longer refuses outright: it reads its configuration, registers its own incarnation and polls for managed work over the real dispatch routes. The service-side surface it needs now exists, and the opt-in cross-repo gate (docs/dispatch-runtime-operator.md › Cross-repo contract gate) drives one whole read-only attempt against it. What is still missing is named below under Unmet gates; the broker, the proxy and their confinement are real and testable today, and no attempt has yet run end to end against a live service.

A refused attempt: precise failure or uncertain

A refusal that the broker proves never started a container — its stop observation for the journalled scope reads status:'created' — is closed centrally as a failed result carrying only the bounded DISPATCH_* code, so the reservation is released instead of held. Everything else stays uncertain, including a scope that actually ran and any refusal without a bounded code; raw error text, paths and credentials never become a central failure reason. A refusal that happens before any scope exists cannot be closed by the supervisor at all: the service accepts no result without stop evidence, and stop evidence binds to a scope id. That request keeps its slot until an operator recovers it. docs/dispatch-runtime-operator.md › Unstarted attempts has the phase-by-phase rule, the measurements behind it and the open gap.

Credentials stay outside the image

The supervisor holds exactly one credential — its own bound API token — and gets it as a file the operator owns (DISPATCH_SUPERVISOR_TOKEN_FILE → /run/secrets/dispatch_supervisor_token). No database URL, no Forgejo token and no dispatch key is an environment value, a build argument or a layer in the image: Dockerfile.dispatch copies only named source paths, installs no Docker CLI and no model CLI, and both targets drop to uid 10001 before their fixed entrypoint. The broker owns Docker access; the supervisor sees only the operator's Unix socket, mounted read-only. The child sees none of it.

Every other value in the compose template comes from the operator's own env file and is required (${VAR:?}), so a missing one fails the command instead of starting a half-configured slot: DISPATCH_SUPERVISOR_IMAGE, S4M_DISPATCH_URL, DISPATCH_MODE, DISPATCH_SLOT_ID, DISPATCH_BOOT_ID, DISPATCH_BROKER_SOCKET_DIR, DISPATCH_JOURNAL_DIR, DISPATCH_PROFILE_DIR, DISPATCH_ATTEMPT_OUTPUT_DIR, DISPATCH_HOST_SOCKET_DIR and DISPATCH_SUPERVISOR_TOKEN_FILE. DISPATCH_HOST_SOCKET is the host route's own listener path and stays empty on a job slot. DISPATCH_WORKER_LOG_DIR is optional and stays unset in the env file: the overlay deploy/queue-dispatch.worker-logs.compose.yml sets it and mounts DISPATCH_WORKER_LOG_POOL_DIR (the host's /srv/scrum4me/worker-logs/dispatch, created once with sudo install -d -o 10001 -g 10001 -m 0755 /srv/scrum4me/worker-logs/dispatch) as the one writable log directory, so every attempt with a job id gets a run-log in Ops Worker Logs (pool dispatch, one instance per slot). See docs/dispatch-runtime-operator.md › Run-log for Worker Logs.

SCRUM4ME_WORKER_INSTANCE_ID was removed from this service: the managed-only supervisor never runs the ordinary worker loop and never sends an instance id, and the slot's pairing with a managed worker lives server-side in queue_dispatch_slots.config. Keeping it here suggested a binding this process does not make. The ordinary worker service still uses it, unchanged.

lib/dispatch-config.ts reads exactly these keys and refuses before anything registers, with one code per cause (DISPATCH_CONFIG_MODE_INVALID, DISPATCH_CONFIG_URL_INVALID, DISPATCH_CONFIG_SLOT_INVALID, DISPATCH_CONFIG_BOOT_ID_INVALID, DISPATCH_CONFIG_BROKER_SOCKET_INVALID, DISPATCH_CONFIG_JOURNAL_INVALID, DISPATCH_CONFIG_OUTPUT_ROOT_INVALID, DISPATCH_CONFIG_HOST_SOCKET_INVALID, DISPATCH_CONFIG_TOKEN_FILE_INVALID, DISPATCH_CONFIG_TOKEN_INVALID, DISPATCH_CONFIG_TOKEN_IN_ENV, DISPATCH_CONFIG_PROFILE_FILE_INVALID, DISPATCH_CONFIG_PROFILE_INVALID, DISPATCH_CONFIG_DURATION_INVALID, DISPATCH_CONFIG_WORKER_LOG_DIR_INVALID). A refusal is always cheaper than a claim this slot would have to abandon after the model already ran.

The pinned profile is the digest

DISPATCH_PROFILE_DIR holds dispatch-profile.json: the exact profile revision this slot serves. The supervisor derives profile_sha256 from those bytes with the same canonical form the service hashes into queue_dispatch_profiles.sha256, so there is no second configured digest to drift. A profile the service would accept but this assembly cannot honour is still refused outright. Every profile, of any action or access, needs the prepared-sources producer — every request carries at least the service's signed __dispatch_input source — and a slot that has not armed it gives DISPATCH_PREPARED_SOURCES_PRODUCER_REQUIRED before it registers.

Arming the prepared-sources producer

Every slot arms it. Two operator keys do so, and they are all-or-nothing (DISPATCH_CONFIG_PREPARED_SOURCES_INCOMPLETE):

  • DISPATCH_PREPARED_SOURCES_ROOT — the staging area, mounted writable at /run/sources and mounted into the broker as its preparedSourcesRoot. The supervisor writes one immutable, hash-verified directory per attempt here; the broker copies its mounts out of it.
  • DISPATCH_PERMIT_PUBLIC_KEYS — the service's trusted Ed25519 public keyset as comma-separated kid:base64url pairs (DER SPKI), the same keys the broker pins. It verifies the signed source manifest by kid and grants nothing, so it rides in the env value like DISPATCH_CREDENTIAL_KEYS, not a file. A private key is refused (DISPATCH_CONFIG_PERMIT_PUBLIC_KEY_INVALID): this slot verifies, it never signs.

Because Compose profiles gate whole services rather than a single volume, the writable /run/sources mount lives in the producer overlay deploy/queue-dispatch.producer.compose.yml, not in the base template. The overlay is mandatory for every slot: it adds the mount and turns both variables above into required (:?) render inputs. The base file alone still renders, but a slot started from it refuses to register (DISPATCH_PREPARED_SOURCES_PRODUCER_REQUIRED). See Validating the template without deploying for the exact command.

With both armed the supervisor fetches POST /dispatch/v1/attempts/sources/manifest under its own attempt proof, downloads exactly the artifact ids that signed manifest names through GET /dispatch/v1/artifacts/:id, and re-hashes every byte against its pin. A changed, missing or substituted source is DISPATCH_PREPARED_SOURCES_REFUSED — there is no Git fetch, no URL from the model, and no forge or database credential anywhere on that path. A repo_write attempt additionally requires a signed repository source, which the broker materialises into /work on the pinned base and the request's own codex/queue-<request-id> branch, from local bytes alone. A read attempt (a review) whose manifest carries a signed repository gets the same checkout in /work, but broker-owned with no group/other write bit, so the child can read it and never change it.

Pinnable image digests

DISPATCH_SUPERVISOR_IMAGE is required and belongs in the release manifest as a digest:

docker build -f Dockerfile.dispatch --target supervisor -t <registry>/<repo>:<release> .
docker build -f Dockerfile.dispatch --target egress     -t <registry>/<repo>-egress:<release> .
# The manifest pins what the registry returns, never a local build id:
docker image inspect --format '{{index .RepoDigests 0}}' <registry>/<repo>:<release>

Build it locally to verify it, but pin only what the registry returns: on a containerd-backed daemon a local build already shows a manifest digest, and that digest is neither a registry reference nor activation approval. Measured on macOS Docker Desktop, both targets build, the image runs as uid 10001 with the fixed entrypoint, and it contains no Docker CLI and no env or key file. The egress proxy image digest is registered with the broker config in the same way (see docs/dispatch-runtime-operator.md › Egress layout).

Broker configuration

BrokerConfig in lib/dispatch-runtime-broker.ts is JSON owned by root and read by the operator's own broker process. maxDurationSeconds is required (integer, 1 to 86400) and is new: the supervisor enforces the profile's max_duration_seconds only while it is alive, so the broker keeps its own deadline and stops a scope even when the supervisor is gone. The deadline is journaled before docker start, armed when the container runs, and re-armed from the journal after a broker restart. Set it no lower than the longest max_duration_seconds of the profiles this slot serves, or the broker stops work the profile still allows. A config without it is refused outright.

socketGroup is required too (numeric gid, 1 or higher): the broker socket is mode 0660, so its group is the whole authorization boundary between the supervisor and the broker. The broker chowns the socket to that group after it listens, instead of inheriting whatever the parent directory grants — which would make access depend on a setgid bit nobody configured. A missing, zero or unusable group refuses with DISPATCH_BROKER_SOCKET_GROUP_REFUSED and the broker does not serve.

The same group reads attempt output (M41 T-1948). The supervisor collects <attempt>/output from the broker root (DISPATCH_ATTEMPT_OUTPUT_DIR, read-only). With the root and every attempt directory at root 0700, uid 10001 got EACCES on Linux (measured on scrum4me-server), so no attempt could ever be collected. The broker therefore makes the root and each attempt directory root:<socketGroup> 0710: the supervisor reaches an attempt by its exact id and never lists the root. At startup the root must be exactly 0700, or 0710 with exactly that group; any other mode (owner, group, other or special bits) is still DISPATCH_BROKER_JOURNAL_ACCESS_REFUSED.

Release. Besides create/start/stop/inspect the broker serves one cleanup operation, release({attemptId}). After the supervisor has a terminal answer for the result it asks the broker to remove exactly that attempt's stopped container (docker rm, never -f, after checking the exact id and both s4m.dispatch.* labels) and its attempt directory — including the operator sources copied into it, such as the child's own provider key. It is refused before the broker journalled its own stop evidence, it keeps the journal and that evidence, and a second call is a no-op. See docs/dispatch-runtime-operator.md › Release after the terminal answer.

Per-host confinement gate

npm run test:dispatch-runtime builds the real proxy library into a deterministic fixture and runs the actual broker against local Docker: confinement, egress binding, deadline and stop evidence. Run it on every host before enabling that host's profile. It proves the broker and confinement on that host, not model or provider compatibility.

  • Pass → record the runtime, image and boot evidence with the release manifest, then enable the profile for that host only.
  • Fail, or Docker unavailable → the host's profile stays disabled and the host is ineligible. Report it as an unmet acceptance gate; never select a broader profile to work around it.

Validating the template without deploying

# Every slot: layer the mandatory producer overlay, which adds the writable /run/sources mount and
# requires DISPATCH_PREPARED_SOURCES_ROOT + DISPATCH_PERMIT_PUBLIC_KEYS.
docker compose -f deploy/queue-dispatch.compose.yml -f deploy/queue-dispatch.producer.compose.yml \
  --env-file <operator env file> --profile dispatch-not-activated config

config renders the template and fails on any missing required variable. Pass the profile: without it the rendered service list is empty and the check proves nothing. Render with the overlay: it is what makes both producer variables required and mounts /run/sources, and lib/dispatch-config.ts refuses every profile without them. It starts nothing, and on a host without the operator's broker socket and journal directory this is the only check that should be run.

The child's own output

The child writes result.json (a DispatchResult) in its /output bind mount, and may write transcript.jsonl next to it; the transcript only feeds the optional run-log (see the DISPATCH_WORKER_LOG_DIR paragraph above) and is never part of the result. The supervisor reads result.json from the broker's attempt root (DISPATCH_ATTEMPT_OUTPUT_DIR → DISPATCH_OUTPUT_ROOT=/run/attempts, read-only) after the stop it observed, validates it with the shared result validator and never executes or interpolates any of it. A cancelled run without a report still terminalizes, because cancellation is the supervisor's own observation rather than a model claim; a normal run without one fails with DISPATCH_RESULT_OUTPUT_MISSING.

Those exact bytes are then staged as an immutable report artifact through PUT /dispatch/v1/attempts/collected/:key before the result is submitted. That is the only artifact path open at this point: the stop this supervisor already submitted revoked the attempt that PUT /attempts/artifacts/:key requires, so collection is bound to the historical start binding instead. It is idempotent on identical bytes, and a failure there stays a failure — it never becomes a local completion. A code result goes through the same route. On a repo_write attempt the supervisor collects the stopped child's committed change from the prepared work area into the exact JSON the central publisher accepts — a Git bundle of base..HEAD with the diff and file list — and stages it as code. Base, branch and origin come from the pinned request and its signed manifest, never from the child: a result naming a different base, branch or head is refused (DISPATCH_CODE_ARTIFACT_REFUSED), and the artifact_id in the submitted result is the service's own receipt rather than anything the child wrote. A code result on a read request, or on a slot with no producer armed, is still refused (DISPATCH_CODE_ARTIFACT_REFUSED / DISPATCH_CODE_ARTIFACT_UNSUPPORTED). Publication itself stays central: the supervisor pushes nothing, holds no forge credential, and never merges or deploys.

Unmet gates

  • No child ever receives its capability. The service mints an attempt-scoped agent_token beside the start permit and serves GET /dispatch/v1/agent/sources/:key and PUT /dispatch/v1/agent/outputs/:key, but nothing here places that token in the container environment, and the broker has no mechanism to. Neither served profile needs it: the child writes its one result.json into its own /output mount, and code is collected post-stop by the supervisor under its own historical binding. The child's only credential is its own provider key (M41 decision B1), delivered as an operator source; see "Model image" in the operator doc. Do not set DISPATCH_AGENT_OUTPUT_KEY: the service then adds agent_token to the start permit and this broker refuses any permit with fields other than expiresAt,permitId.
  • checks is still staged by nobody. Collection knows report and code; the child's checks travel inside result.json and no separate checks artifact is produced.
  • Non-launch recovery is not driven end to end from here. POST /attempts/recovery/{lookup,stop,result} exists, the transport and assembleManagedPorts are wired to it, and the service's own tests cover it; recoverNonLaunchAttempt against the real routes is not in the cross-repo gate.
  • No attempt has run against a live service, and no attempt has ever run through this compose template. The cross-repo gate boots the service in-process against a disposable test cluster with a fake runtime: no container, no image, no model, no egress. Its repo_write case publishes into a local bare repository over file://, so Forgejo remote-head conflicts and draft-PR creation stay unmeasured, and the root broker's own copy/verify/mount step is simulated rather than executed.
  • Linux confinement is proved on scrum4me-server only (2026-10-02, M41 increment 3: npm run test:dispatch-runtime and npm run test:dispatch-model-image green); max2 has not been measured.
  • A meaningful Unix-socket liveness healthcheck for the egress proxy is still operator work; file existence is not readiness. Provider proxy support in the release model image is proved for the model-codex adapter (T-1869) and, since M41 5b, the model-claude adapter (probe child in the same test: npm run test:dispatch-model-image; live review on Docker Desktop: npm run smoke:dispatch-live-review, 2026-10-02, verdict NO-GO in 34 s, key absent from every output, container and attempt directory released).

Bekende grenzen

  • Eén actieve job tegelijk. De wrapper-loop is sequentieel. Voor parallellisme zou je meerdere containers met dezelfde SCRUM4ME_TOKEN kunnen draaien — wait_for_job gebruikt FOR UPDATE SKIP LOCKED dus dat is veilig op DB-niveau, maar dan moet je je node_modules-cache per container scheiden.
  • OAuth-token: 1 jaar geldig. Bij verloop schrijft de wrapper een TOKEN_EXPIRED-marker en wordt de container unhealthy. Geen auto-rotatie.
  • npm install per job kost op een N5095 ~30–60 s per Next.js-clone, óók met de pnpm-store. Voor zeer kleine fixes is dat de dominante factor. Kan later vervangen worden door een persistente warm-node_modules per repo als dat een knelpunt wordt.