- TypeScript 74.2%
- Python 10.3%
- Shell 9.8%
- JavaScript 4.4%
- Dockerfile 1.3%
| .forgejo/workflows | ||
| __tests__ | ||
| bin | ||
| codex | ||
| deploy | ||
| docs | ||
| etc | ||
| lib | ||
| scripts | ||
| skills | ||
| tests | ||
| vendor | ||
| .env.deploy.example | ||
| .env.example | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| docker-compose.yml | ||
| Dockerfile | ||
| Dockerfile.dispatch | ||
| mcp-config.json | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| tsconfig.dispatch.json | ||
scrum4me-agent-runner
Headless Claude Code worker die de Scrum4Me job-queue (M13) leegtrekt vanaf
een QNAP NAS via Container Station. Geen Vercel, geen browser, geen
toetsenbord — Claude Code draait als daemon, claimt jobs uit
mcp__scrum4me__wait_for_job, voert ze uit in een per-job clone, en pusht
nooit zelf.
Architectuur in één plaatje
┌─ QNAP TS-664 (Container Station) ─────────────────────────────┐
│ │
│ ┌─ container: agent-runner ────────────────────────────────┐ │
│ │ PID 1: tini → run-agent.sh (daemon-loop) │ │
│ │ ├─ health-server.js (8080 → host 18080) │ │
│ │ └─ claude -p (per-batch, met MCP via stdio) │ │
│ │ └─ scrum4me-mcp → Neon Postgres │ │
│ │ │ │
│ │ /tmp/job-<id> ephemeral working trees │ │
│ │ /var/cache/repos bare git mirrors (volume) │ │
│ │ /var/cache/npm npm cache (volume) │ │
│ │ /var/log/agent run + job logs (volume) │ │
│ └──────────────────────────────────────────────────────────┘ │
│ │
│ /share/Agent/cache /share/Agent/logs /share/Agent/state │
└────────────────────────────────────────────────────────────────┘
│
▼ HTTPS
Neon Postgres (Scrum4Me DB)
▲
│
Vercel ─── Scrum4Me UI (gebruikers enqueueen jobs)
Eén claude -p-invocation roept intern wait_for_job aan totdat de
queue leeg is (≈600 s lege block-time → afsluiten). De wrapper start
claude -p opnieuw zodra hij eindigt, met exponentiële backoff bij
fouten.
Wat zit waar
| Bestand | Doel |
|---|---|
Dockerfile |
Ubuntu 22.04 + Node 22 + Claude Code + scrum4me-mcp + scripts |
docker-compose.yml |
Service-definitie, volumes, env-file, restart-policy, limits |
package.json |
Npm-dependencies van de runner zelf (alleen scrum4me-mcp pin) |
mcp-config.json |
Claude Code MCP-config (verwijst stdio naar scrum4me-mcp) |
CLAUDE.md |
Agent-rol-instructies, auto-geladen door claude -p |
bin/entrypoint.sh |
Container-startup: dirs, health-server, daemon-loop |
bin/install-skills.sh |
Container-startup: publiceert de APPROVED skills-inventory naar de discovery-mappen van Claude/Codex/Agents |
bin/run-agent.sh |
Daemon-loop met backoff, exit-code-routing en state-writes |
bin/check-tokens.sh |
Pre-flight: Scrum4Me-token, Claude OAuth-token, DB-bereikbaarheid, DB-rolcapaciteiten |
bin/check-database-role.cjs |
Pre-flight-helper: weigert verhoogde PostgreSQL-rollen op de aanwezige worker-DB-URL's |
bin/job-prepare.sh |
Per-job: bare-fetch + clone-via-reference naar /tmp/job-<id> |
bin/job-cleanup.sh |
Per-job: logs naar /var/log, working tree weg |
bin/health-server.js |
HTTP-endpoint op 8080 (intern) dat state.json en marker-files leest |
bin/rotate-logs.sh |
Compress/cleanup van oude .log-bestanden |
.env.example |
Alle env-vars met uitleg |
Vereisten op de NAS
- Container Station 2+ (Docker compose v2)
Agentals QTS Shared Folder op een echte volume (bv.CACHEDEV1_DATA). Niet eenmkdir /share/Agent—/sharezelf is een 16 MB tmpfs en handmatige directories overleven geen reboot. Aanmaken via Control Panel → Privilege → Shared Folders → Create. QTS legt dan automatisch de symlink/share/Agent → /share/CACHEDEV1_DATA/Agent.- Drie subdirs onder die share:
/share/Agent/cache,/share/Agent/logs,/share/Agent/state. Aanmaken via File Station of via SSH na share-creatie. - Internet-uitgang naar
api.anthropic.com,git.jp-visser.nl(Forgejo HTTPS-clone/push),cli.github.com(build-time voor de gh CLI), je Neon-host,registry.npmjs.org.
Verifieer vóór je deployt dat
/share/Agentecht op disk staat:ssh admin@<nas> 'ls -la /share/ | grep Agent; df -h /share/Agent'Verwacht een symlink (
l...Agent -> /share/CACHEDEV1_DATA/Agent) en een df-uitvoer met TB-grootte opcachedev1/cachedev2. Als je hiertmpfs 16Mziet, is de share geen geregistreerde QTS Shared Folder en zal elke transfer16 MB falen met
scp: write remote ... Failure.
Deploy
# 1. Op je werkstation: token's regelen
# a. CLAUDE_CODE_OAUTH_TOKEN → draai `claude setup-token` (browser-flow)
# b. SCRUM4ME_TOKEN → log in als de dedicated agent-user in
# Scrum4Me, /settings/tokens, label "NAS-runner"
# c. DATABASE_URL/DIRECT_URL → Neon dashboard
# d. GH_TOKEN → Forgejo → avatar → Settings →
# Applications → Generate New Token; scope
# minimaal `write:repository` op de twee
# repos (janpeter/Scrum4Me + janpeter/
# scrum4me-mcp). Wordt gebruikt voor clone
# en push naar Forgejo. PBI-86 (hybride
# model): `gh pr create` is uit de
# worker-flow verwijderd — de GitHub-PR
# komt via de handmatige promote-Action
# in Forgejo.
# 2. Repo op de NAS plaatsen
ssh admin@nas
cd /share/Agent
git clone https://git.jp-visser.nl/<jij>/scrum4me-agent-runner.git
cd scrum4me-agent-runner
# 3. Env aanmaken
cp .env.example .env
chmod 600 .env
vi .env # vul alle waarden in
# 4. Build + start
docker compose build
docker compose up -d
# 5. Verifiëren
curl http://nas.local:18080/health
docker compose logs -f
QNAP-port: host-poort 8080 is bezet door de QTS-webinterface; daarom mapt deze stack standaard
18080:8080. Override viaAGENT_HEALTH_PORT_HOSTin.envals je een andere host-poort wilt.
Snelle redeploy — bin/deploy-to-nas.sh
Voor een bestaande deploy die je opnieuw wil bouwen + deployen
(bijvoorbeeld na een merge in scrum4me-mcp of een aanpassing aan
CLAUDE.md):
# Eenmalig: NAS-target instellen
cp .env.deploy.example .env.deploy
vi .env.deploy # zet NAS_HOST=admin@<nas>
# Pin in .env exact één bestaande immutable MCP-release
vi .env # zet MCP_GIT_REF=deploy/<immutable-release-tag>
# Daarna: één commando voor de hele cyclus
bin/deploy-to-nas.sh
Het script verifieert eerst de persistente MCP_GIT_REF uit .env zonder het
bestand te sourcen. main, kale SHA's, dubbele assignments en onbekende tags
stoppen vóór de build. Daarna doet het:
docker buildx build --platform linux/amd64 --loaddocker save | gzip → scrum4me-agent-runner-amd64.tar.gzscpvan tarball +docker-compose.yml+ (eerste keer).envnaar NASsshop NAS:docker load+ sanity-check op.env+docker compose up -d --force-recreatedocker compose logs -f— lokaal-volgbaar terwijl pre-flight + eerste batch starten
.env op de NAS wordt niet overschreven als 'ie er al staat. Bij een
verse NAS-installatie wordt 'ie wél geüpload + ge-sed't (NAS_BASE,
AGENT_UID, etc.). Voor de eerste deploy of een schoon volume zie de
volledige procedure hieronder.
Deploy — cross-build vanaf Mac (Apple Silicon → amd64-NAS)
Alternatief voor de in-place build hierboven. Bouw de image op je Mac voor
linux/amd64, schrijf 'm naar een tarball, transfer naar de NAS en laad daar.
Handig als de NAS te langzaam is om te builden (npm install op een N5095 met
NAS-storage is traag) of als je geen git push wilt voor elke iteratie.
Vóór je begint: controleer dat /share/Agent een echte QTS Shared Folder is
(zie Vereisten op de NAS). Dat is de meest voorkomende
val.
1. Tokens en .env op je Mac
Zelfde tokens als in Deploy. Hou er rekening mee dat je twee
.env-bestanden kunt willen: één voor lokaal Mac-testen
(AGENT_PLATFORM=linux/arm64, AGENT_UID=501, paths onder /Users/...) en
één voor NAS-runtime. De NAS-versie wordt bij stap 4 ge-scp'd en met sed
geschikt gemaakt.
2. Image bouwen voor amd64
cd /Users/<jij>/Development/scrum4me-docker
docker buildx build \
--platform linux/amd64 \
--build-arg MCP_GIT_REF=deploy/<immutable-release-tag> \
--build-arg CLAUDE_CODE_VERSION=2.1.287 \
--build-arg AGENT_UID=1000 \
--build-arg AGENT_GID=1000 \
-t scrum4me-agent-runner:local \
--load \
.
AGENT_UID/GID=1000 zijn de NAS-UIDs, niet je Mac-UIDs (vaak 501/20).
Bij verkeerde UIDs kan de container niet schrijven naar de bind-mounts.
Verifieer architectuur:
docker image inspect scrum4me-agent-runner:local --format '{{.Architecture}}'
# verwacht: amd64
3. Image naar tarball
docker save scrum4me-agent-runner:local | gzip > scrum4me-agent-runner-amd64.tar.gz
shasum -a 256 scrum4me-agent-runner-amd64.tar.gz
De tarball is ~580 MB voor de huidige image-size. Hij staat in .gitignore,
dus geen risico op accidenteel committen.
4. NAS voorbereiden + bestanden transferren
# .env naar de NAS (de Mac-versie; we patchen path-velden zo direct)
scp .env admin@<nas>:/share/Agent/scrum4me-agent-runner/.env
# image-tarball + bijgewerkte compose
scp scrum4me-agent-runner-amd64.tar.gz docker-compose.yml package.json README.md \
admin@<nas>:/share/Agent/scrum4me-agent-runner/
Zet de NAS-specifieke waarden in .env:
ssh admin@<nas> "
source /etc/profile
cd /share/Agent/scrum4me-agent-runner
sed -i \
-e 's|^NAS_BASE=.*|NAS_BASE=/share/Agent|' \
-e 's|^AGENT_BASE=.*|AGENT_BASE=/share/Agent|' \
-e 's|^AGENT_PLATFORM=.*|AGENT_PLATFORM=linux/amd64|' \
-e 's|^AGENT_UID=.*|AGENT_UID=1000|' \
-e 's|^AGENT_GID=.*|AGENT_GID=1000|' \
-e 's|^AGENT_HEALTH_PORT_HOST=.*|AGENT_HEALTH_PORT_HOST=18080|' \
.env
chmod 600 .env
"
source /etc/profileis verplicht in non-interactieve ssh op QNAP. Zonder die regel staatdockerniet op$PATH(Container Station's binary zit onder/share/CACHEDEV*_DATA/.qpkg/container-station/...) en faalt elkdocker-commando metcommand not found.
5. Image laden + container recreaten
ssh admin@<nas> "
source /etc/profile
set -e
cd /share/Agent/scrum4me-agent-runner
# checksum-verificatie (optioneel maar aan te raden)
sha256sum scrum4me-agent-runner-amd64.tar.gz
# image laden — overschrijft bestaande :local tag
gunzip -c scrum4me-agent-runner-amd64.tar.gz | docker load
# container recreate; --no-build voorkomt onbedoelde NAS-side build
docker compose up -d --no-build --force-recreate agent
docker compose ps
curl -fsS http://localhost:18080/health | head -c 800
"
Verwacht in de health-output cache_free_bytes met een groot getal
(TB-orde) — dat is je signaal dat /var/cache op echte disk zit.
Een waarde rond 16 MB (16777216) betekent dat je per ongeluk
nog steeds op tmpfs draait.
Updaten (handmatig, bewust)
SCRUM4ME_TOKEN of CLAUDE_CODE_OAUTH_TOKEN rouleer je via een rebuild:
cd /share/Agent/scrum4me-agent-runner
git pull
vi .env # nieuwe waarden
docker compose build # nieuwe scrum4me-mcp-versie als dat veranderd is
docker compose up -d
Dezelfde flow voor schema-drift in scrum4me-mcp: maak eerst een immutable
deploy/*-tag, pin die als MCP_GIT_REF in .env, verifieer hem met
scripts/verify-mcp-release-ref.sh .env, en rebuild. Een kale commit-SHA werkt
niet: de Dockerfile gebruikt git clone --branch.
CLAUDE_CODE_VERSION — pin van de Claude Code CLI
De claude-stage van de Dockerfile installeert @anthropic-ai/claude-code
via npm (sinds 2026-06-08, eerder via claude.ai/install.sh). De ARG is
gepinned op 2.1.287 (npm latest per 2026-10-02; claude-opus-5-5 eist
= 2.1.280, dus
stablemag pas als die daar voorbij is) — bewuste bumps gaan via een PR die de Dockerfile-ARG,.env.exampleén deze paragraaf simultaan updatet zodat verse rebuilds reproducerbaar blijven.
# huidige dist-tags bekijken
npm view @anthropic-ai/claude-code dist-tags
# bump naar een specifieke versie via .env
sed -i 's|^CLAUDE_CODE_VERSION=.*|CLAUDE_CODE_VERSION=2.1.287|' .env
docker compose build agent
docker compose up -d --force-recreate agent
De claude-build verifieert de geïnstalleerde versie: bij een exacte
numerieke pin controleert grep -F dat het binary die versie rapporteert;
bij CLAUDE_CODE_VERSION=latest (of een andere non-numerieke waarde) volstaat
een niet-lege versistring. Een typo of niet-bestaande npm-versie faalt zo
alsnog hard.
PRISMA_CLI_VERSION — pin van de Prisma-CLI in de mcp-clone-laag
De mcp-clone-laag draait na npm ci --omit=dev een npx prisma generate.
--omit=dev verwijdert de prisma-CLI (die is een devDependency van
scrum4me-mcp), waarna npx de CLI ophaalt van npm. Ongepind pakt npx de
dist-tag latest, en die staat sinds 2026-08 op een release candidate
(8.0.0-rc.12) die geen generate-commando meer kent — waardoor élke verse
worker-build hard afbrak met CLI.UNKNOWN_COMMAND. De Dockerfile pint daarom
via ARG PRISMA_CLI_VERSION=7.8.0 en roept npx prisma@${PRISMA_CLI_VERSION} generate aan. 7.8.0 is dezelfde versie als @prisma/client in scrum4me-mcp's
package-lock.json; CLI en client horen dezelfde major/minor te hebben.
Bump deze pin samen met scrum4me-mcp wanneer daar de prisma-versie wijzigt.
Wijzigingen in docker-compose.yml (volumes, tmpfs, env_file, ports)
Let op:
docker compose restartherstart alleen het proces in de bestaande container met de oude config. Wijzigingen in volumes, tmpfs-mounts, env_file of ports worden daarmee niet doorgevoerd.
Gebruik altijd --force-recreate als je docker-compose.yml is veranderd:
docker compose up -d --force-recreate agent
Verifieer daarna dat /var/cache op de NAS-overlay staat en niet op tmpfs:
docker exec scrum4me-agent df -h /var/cache
# Verwacht: Filesystem op /dev/mapper/cachedev* of een NAS-share
# Fout: tmpfs 16M ... (dan is force-recreate niet uitgevoerd)
Health-endpoint
GET http://<nas>:18080/health retourneert:
{
"status": "running", // running | idle | unhealthy | token-expired
"lastBatchAt": "2026-05-01T12:34:56Z",
"lastBatchExit": 0,
"consecutiveFailures": 0,
"tokenStatus": { "anthropic": "ok", "scrum4me": "ok", "db": "ok" }
}
HTTP-status: 200 als running/idle, 503 bij token-expired of als de
laatste heartbeat ouder is dan 5 minuten.
Ook 503 met {"status":"unhealthy","reason":"tmp-low","tmpUsedPct":<n>}
zolang de marker TMP_PRESSURE in de state-dir staat. bin/tmp-sweep.sh
draait vanuit run-agent.sh tussen iteraties (bij de start, na een overload
en na elke iteratie). Het wist alles van de agent-gebruiker direct onder /tmp,
behalve tsx-*, node-compile-cache en verborgen entries. Zit /tmp daarna
nog op of boven AGENT_TMP_WARN_PCT (default 80), dan schrijft het de marker.
Filesystem-grenzen
De agent-user heeft geen SSH-keys en geen toegang tot andere shares dan
/share/Agent/*. Wel een ~/.git-credentials met de GH_TOKEN voor
HTTPS-clone/push (zie volgende sectie) — die token is scoped tot de twee
configured repos en mag worden gerouleerd door rebuild + redeploy.
Repo bootstrap (clone-on-start)
Bij elke container-start runt bin/repo-bootstrap.sh (als de
agent-user, ná drop-privileges) en zet zo'n setup neer:
- Configureert git's credential-helper met
GH_TOKENzodatgit clone/pushnaarhttps://git.jp-visser.nl/...(Forgejo) zonder prompt werkt. - Voor elke repo in
GH_PRECLONE_REPOS(komma-gescheiden owner/name):- Bestaat
~/Projects/<name>/.gital? →git fetch origin --prune - Anders → fresh
git clone
- Bestaat
Daarna vindt scrum4me-mcp's resolveRepoRoot (in wait_for_job) de
clone via z'n convention-fallback ~/Projects/<name>/.git. Worktrees
voor jobs landen vervolgens onder ~/.scrum4me-agent-worktrees/<jobId>/
zodat de hoofd-clone niet wordt aangeraakt.
Push gaat over dezelfde token: git push -u origin feat/story-<id>
slaagt zonder prompt. gh pr create is in PBI-86 (T-1005) verwijderd
uit de worker-flow — de GitHub-PR ontstaat via een handmatig
getriggerde promote-Action in Forgejo (zie de Scrum4Me-repo
docs/runbooks/forgejo-hybrid-flow.md).
Veelvoorkomende issues
Verzameld uit echte deploy-sessies op een TS-664 met QTS 5.x. Eerst checken voor je gaat tweaken.
/share/Agent is tmpfs 16M — scp faalt op grote files
Symptoom:
scp: write remote ".../scrum4me-agent-runner-amd64.tar.gz": Failure
Eerst 100 % bij kleine files, dan blijvend "Failure" zodra een file de tmpfs vol maakt.
Oorzaak: /share/Agent is geen geregistreerde QTS Shared Folder maar een
gewone directory in de root-tmpfs. Na elke reboot is alle inhoud weg en is
/share/Agent slechts een entry in de 16 MB /share-tmpfs.
Fix: maak Agent aan via Control Panel → Privilege → Shared Folders →
Create. Daarna verschijnt automatisch lrwxrwxrwx /share/Agent → /share/CACHEDEV1_DATA/Agent.
Als die symlink ontbreekt omdat er al een tmpfs-directory /share/Agent staat:
ssh admin@<nas> "
source /etc/profile
docker stop scrum4me-agent 2>/dev/null
docker rm scrum4me-agent 2>/dev/null
rm -rf /share/Agent
ln -s /share/CACHEDEV1_DATA/Agent /share/Agent
mkdir -p /share/Agent/{cache,logs,state,scrum4me-agent-runner}
ls -la /share/ | grep Agent # moet nu een symlink tonen
"
rm -rf /share/Agentis veilig zolang/share/Agentnog tmpfs is — alle content stond in RAM en zou de volgende reboot toch verdwenen zijn. Maar verifieer eerst metdf -h /share/Agentdat 't echt tmpfs is.
docker: command not found in non-interactieve ssh
Symptoom:
ssh admin@<nas> 'docker ps'
# sh: line 1: docker: command not found
Oorzaak: QNAP's admin-user heeft docker op $PATH via
/etc/profile.d/*.sh van Container Station. Login-shells laden die scripts;
non-interactieve ssh user@host 'cmd' doet dat niet.
Fix: source /etc/profile aan het begin van je remote command:
ssh admin@<nas> "source /etc/profile && docker ps"
yaml: control characters are not allowed
Symptoom:
docker compose down
# yaml: control characters are not allowed
Oorzaak: een eerdere scp van docker-compose.yml faalde halverwege en
liet een file met NUL-bytes / partial writes achter. De file-grootte klopt
maar de inhoud is corrupt.
Fix: scp opnieuw vanaf je werkstation. Voor het stoppen van een al- draaiende container heb je de yml niet nodig:
docker stop scrum4me-agent && docker rm scrum4me-agent
Daarna scp docker-compose.yml admin@<nas>:/share/Agent/scrum4me-agent-runner/
en pas dan docker compose up -d --no-build --force-recreate.
QTS-console-menu Python-traceback bij ssh-commando
Symptoom: lange Python-traceback over consolemenu_q/prompt_utils.py:31
en Inappropriate ioctl for device, gevolgd door je daadwerkelijke output.
Oorzaak: QTS' admin-shell start een Python-TUI (qts-console-menu) die crasht zonder echte TTY. Niet-fataal — je commando wordt alsnog gerund.
Fix: negeer de traceback. Als je 'm écht weg wilt, gebruik
ssh -tt admin@<nas> '...' voor een geforceerde pseudo-TTY, maar pas op:
sommige scripts hangen daarop omdat ze interactie verwachten.
.env is weg na rm -rf /share/Agent
Symptoom: na het opruimen van een tmpfs-/share/Agent weigert
docker compose up:
env file /share/Agent/scrum4me-agent-runner/.env not found
Oorzaak: .env zit in .gitignore en wordt nooit door git pull of scp . automatisch teruggezet. Bij het wegnuken van een corrupte share verdwijnt-ie
mee.
Fix: scp 'm vanaf je werkstation, en patch path-velden op de NAS-zijde
(zie Deploy — cross-build
stap 4). Zorg dat je .env op je werkstation alle secrets bevat — vooral
SCRUM4ME_TOKEN, dat in een Mac-only .env makkelijk ontbreekt omdat
lokale tests soms zonder Scrum4Me-API draaien.
Environment variables
WORKER_CAPABILITY (optional during rollout, required after Phase C)
Hardware tier of this host. The MCP claim-filter biases jobs to higher-tier idle workers (pure preference race, no waiting). Set per host in .env:
HIGH_P— top performance (current:max2)MEDIUM_P— mid performance (current:mac)LOW_P— entry performance (current:scrum4me-server)
See design.
INTERNAL_PUSH_URL + INTERNAL_PUSH_SECRET (optioneel — web-push, PBI-55)
push-trigger.ts in de MCP-subprocess POST't fire-and-forget naar
<app>/api/internal/push/send bij ask_user_question (o.a. grill-vragen) en
update_job_status naar done/failed. Beide vars leeg of afwezig = feature
stil uit. De URL moet wijzen naar de app-instantie die dezelfde database
gebruikt als deze worker (anders bestaan userId/push-subscriptions daar
niet); het secret moet byte-gelijk zijn aan INTERNAL_PUSH_SECRET in de
app-env. Doorgifte aan het MCP-proces is per runner expliciet geregeld:
mcp-config.json (Claude, ${VAR:-}-expansie) en codex/config.toml
env_vars (Codex) — alleen .env vullen is dus niet genoeg ná wijziging
van die bestanden; rebuild vereist omdat beide in het image gebakken worden.
AGENT_TRANSCRIPT_RETENTION_DAYS (optioneel, default 7)
entrypoint.sh installeert de user-settings van Claude Code en roept daarna
bin/set-transcript-retention.sh aan. Dat script zet cleanupPeriodDays in
/home/agent/.claude/settings.json, zodat Claude Code transcripts onder
~/.claude/projects die ouder zijn dan dit aantal dagen zelf opruimt. Alleen
positieve gehele getallen zijn geldig. Bij ongeldige invoer blijft het bestand
ongewijzigd en logt de entrypoint een WARN; de worker start gewoon.
Queue dispatch rollout (IDEA-213)
This repository supplies the managed runtime side of the central dispatch service: the
operator-owned broker (bin/run-dispatch-runtime-broker.ts), the egress proxy and the supervisor
image built from Dockerfile.dispatch. deploy/queue-dispatch.compose.yml is an opt-in template
behind the profile dispatch-not-activated; read docs/dispatch-runtime-operator.md first, it
holds the trust model, the egress layout and the lifecycle this section only points at.
Nothing here is activated. The supervisor entrypoint no longer refuses outright: it reads its
configuration, registers its own incarnation and polls for managed work over the real dispatch
routes. The service-side surface it needs now exists, and the opt-in cross-repo gate
(docs/dispatch-runtime-operator.md › Cross-repo contract gate) drives one whole read-only attempt
against it. What is still missing is named below under Unmet gates; the broker, the proxy and their
confinement are real and testable today, and no attempt has yet run end to end against a live
service.
A refused attempt: precise failure or uncertain
A refusal that the broker proves never started a container — its stop observation for the
journalled scope reads status:'created' — is closed centrally as a failed result carrying only
the bounded DISPATCH_* code, so the reservation is released instead of held. Everything else
stays uncertain, including a scope that actually ran and any refusal without a bounded code; raw
error text, paths and credentials never become a central failure reason. A refusal that happens
before any scope exists cannot be closed by the supervisor at all: the service accepts no
result without stop evidence, and stop evidence binds to a scope id. That request keeps its slot
until an operator recovers it. docs/dispatch-runtime-operator.md › Unstarted attempts has the
phase-by-phase rule, the measurements behind it and the open gap.
Credentials stay outside the image
The supervisor holds exactly one credential — its own bound API token — and gets it as a file the
operator owns (DISPATCH_SUPERVISOR_TOKEN_FILE → /run/secrets/dispatch_supervisor_token). No
database URL, no Forgejo token and no dispatch key is an environment value, a build argument or a
layer in the image: Dockerfile.dispatch copies only named source paths, installs no Docker CLI
and no model CLI, and both targets drop to uid 10001 before their fixed entrypoint. The broker
owns Docker access; the supervisor sees only the operator's Unix socket, mounted read-only. The
child sees none of it.
Every other value in the compose template comes from the operator's own env file and is required
(${VAR:?}), so a missing one fails the command instead of starting a half-configured slot:
DISPATCH_SUPERVISOR_IMAGE, S4M_DISPATCH_URL, DISPATCH_MODE, DISPATCH_SLOT_ID,
DISPATCH_BOOT_ID, DISPATCH_BROKER_SOCKET_DIR, DISPATCH_JOURNAL_DIR, DISPATCH_PROFILE_DIR,
DISPATCH_ATTEMPT_OUTPUT_DIR, DISPATCH_HOST_SOCKET_DIR and DISPATCH_SUPERVISOR_TOKEN_FILE.
DISPATCH_HOST_SOCKET is the host route's own listener path and stays empty on a job slot.
DISPATCH_WORKER_LOG_DIR is optional and stays unset in the env file: the overlay
deploy/queue-dispatch.worker-logs.compose.yml sets it and mounts DISPATCH_WORKER_LOG_POOL_DIR
(the host's /srv/scrum4me/worker-logs/dispatch, created once with
sudo install -d -o 10001 -g 10001 -m 0755 /srv/scrum4me/worker-logs/dispatch) as the one writable
log directory, so every attempt with a job id gets a run-log in Ops Worker Logs (pool dispatch, one
instance per slot). See docs/dispatch-runtime-operator.md › Run-log for Worker Logs.
SCRUM4ME_WORKER_INSTANCE_ID was removed from this service: the managed-only supervisor never
runs the ordinary worker loop and never sends an instance id, and the slot's pairing with a managed
worker lives server-side in queue_dispatch_slots.config. Keeping it here suggested a binding this
process does not make. The ordinary worker service still uses it, unchanged.
lib/dispatch-config.ts reads exactly these keys and refuses before anything registers, with one
code per cause (DISPATCH_CONFIG_MODE_INVALID, DISPATCH_CONFIG_URL_INVALID,
DISPATCH_CONFIG_SLOT_INVALID, DISPATCH_CONFIG_BOOT_ID_INVALID,
DISPATCH_CONFIG_BROKER_SOCKET_INVALID, DISPATCH_CONFIG_JOURNAL_INVALID,
DISPATCH_CONFIG_OUTPUT_ROOT_INVALID, DISPATCH_CONFIG_HOST_SOCKET_INVALID,
DISPATCH_CONFIG_TOKEN_FILE_INVALID, DISPATCH_CONFIG_TOKEN_INVALID,
DISPATCH_CONFIG_TOKEN_IN_ENV, DISPATCH_CONFIG_PROFILE_FILE_INVALID,
DISPATCH_CONFIG_PROFILE_INVALID, DISPATCH_CONFIG_DURATION_INVALID,
DISPATCH_CONFIG_WORKER_LOG_DIR_INVALID). A refusal is always cheaper
than a claim this slot would have to abandon after the model already ran.
The pinned profile is the digest
DISPATCH_PROFILE_DIR holds dispatch-profile.json: the exact profile revision this slot serves.
The supervisor derives profile_sha256 from those bytes with the same canonical form the service
hashes into queue_dispatch_profiles.sha256, so there is no second configured digest to drift. A
profile the service would accept but this assembly cannot honour is still refused outright. Every
profile, of any action or access, needs the prepared-sources producer — every request carries at
least the service's signed __dispatch_input source — and a slot that has not armed it gives
DISPATCH_PREPARED_SOURCES_PRODUCER_REQUIRED before it registers.
Arming the prepared-sources producer
Every slot arms it. Two operator keys do so, and they are all-or-nothing
(DISPATCH_CONFIG_PREPARED_SOURCES_INCOMPLETE):
DISPATCH_PREPARED_SOURCES_ROOT— the staging area, mounted writable at/run/sourcesand mounted into the broker as itspreparedSourcesRoot. The supervisor writes one immutable, hash-verified directory per attempt here; the broker copies its mounts out of it.DISPATCH_PERMIT_PUBLIC_KEYS— the service's trusted Ed25519 public keyset as comma-separatedkid:base64urlpairs (DER SPKI), the same keys the broker pins. It verifies the signed source manifest bykidand grants nothing, so it rides in the env value likeDISPATCH_CREDENTIAL_KEYS, not a file. A private key is refused (DISPATCH_CONFIG_PERMIT_PUBLIC_KEY_INVALID): this slot verifies, it never signs.
Because Compose profiles gate whole services rather than a single volume, the writable
/run/sources mount lives in the producer overlay deploy/queue-dispatch.producer.compose.yml,
not in the base template. The overlay is mandatory for every slot: it adds the mount and turns
both variables above into required (:?) render inputs. The base file alone still renders, but a
slot started from it refuses to register (DISPATCH_PREPARED_SOURCES_PRODUCER_REQUIRED). See
Validating the template without deploying for the exact command.
With both armed the supervisor fetches POST /dispatch/v1/attempts/sources/manifest under its own
attempt proof, downloads exactly the artifact ids that signed manifest names through
GET /dispatch/v1/artifacts/:id, and re-hashes every byte against its pin. A changed, missing or
substituted source is DISPATCH_PREPARED_SOURCES_REFUSED — there is no Git fetch, no URL from the
model, and no forge or database credential anywhere on that path. A repo_write attempt additionally
requires a signed repository source, which the broker materialises into /work on the pinned base
and the request's own codex/queue-<request-id> branch, from local bytes alone. A read attempt (a review) whose
manifest carries a signed repository gets the same checkout in /work, but broker-owned with no
group/other write bit, so the child can read it and never change it.
Pinnable image digests
DISPATCH_SUPERVISOR_IMAGE is required and belongs in the release manifest as a digest:
docker build -f Dockerfile.dispatch --target supervisor -t <registry>/<repo>:<release> .
docker build -f Dockerfile.dispatch --target egress -t <registry>/<repo>-egress:<release> .
# The manifest pins what the registry returns, never a local build id:
docker image inspect --format '{{index .RepoDigests 0}}' <registry>/<repo>:<release>
Build it locally to verify it, but pin only what the registry returns: on a containerd-backed
daemon a local build already shows a manifest digest, and that digest is neither a registry
reference nor activation approval. Measured on macOS Docker Desktop, both targets build, the image
runs as uid 10001 with the fixed entrypoint, and it contains no Docker CLI and no env or key file. The egress proxy image digest is registered with
the broker config in the same way (see docs/dispatch-runtime-operator.md › Egress layout).
Broker configuration
BrokerConfig in lib/dispatch-runtime-broker.ts is JSON owned by root and read by the operator's
own broker process. maxDurationSeconds is required (integer, 1 to 86400) and is new: the
supervisor enforces the profile's max_duration_seconds only while it is alive, so the broker
keeps its own deadline and stops a scope even when the supervisor is gone. The deadline is
journaled before docker start, armed when the container runs, and re-armed from the journal after
a broker restart. Set it no lower than the longest max_duration_seconds of the profiles this slot
serves, or the broker stops work the profile still allows. A config without it is refused outright.
socketGroup is required too (numeric gid, 1 or higher): the broker socket is mode 0660, so
its group is the whole authorization boundary between the supervisor and the broker. The broker
chowns the socket to that group after it listens, instead of inheriting whatever the parent
directory grants — which would make access depend on a setgid bit nobody configured. A missing,
zero or unusable group refuses with DISPATCH_BROKER_SOCKET_GROUP_REFUSED and the broker does not
serve.
The same group reads attempt output (M41 T-1948). The supervisor collects <attempt>/output
from the broker root (DISPATCH_ATTEMPT_OUTPUT_DIR, read-only). With the root and every attempt
directory at root 0700, uid 10001 got EACCES on Linux (measured on scrum4me-server), so no attempt
could ever be collected. The broker therefore makes the root and each attempt directory
root:<socketGroup> 0710: the supervisor reaches an attempt by its exact id and never lists the
root. At startup the root must be exactly 0700, or 0710 with exactly that group; any other mode (owner,
group, other or special bits) is still DISPATCH_BROKER_JOURNAL_ACCESS_REFUSED.
Release. Besides create/start/stop/inspect the broker serves one cleanup operation,
release({attemptId}). After the supervisor has a terminal answer for the result it asks the broker
to remove exactly that attempt's stopped container (docker rm, never -f, after checking the
exact id and both s4m.dispatch.* labels) and its attempt directory — including the operator sources
copied into it, such as the child's own provider key. It is refused before the broker journalled its
own stop evidence, it keeps the journal and that evidence, and a second call is a no-op. See
docs/dispatch-runtime-operator.md › Release after the terminal answer.
Per-host confinement gate
npm run test:dispatch-runtime builds the real proxy library into a deterministic fixture and runs
the actual broker against local Docker: confinement, egress binding, deadline and stop evidence.
Run it on every host before enabling that host's profile. It proves the broker and confinement
on that host, not model or provider compatibility.
- Pass → record the runtime, image and boot evidence with the release manifest, then enable the profile for that host only.
- Fail, or Docker unavailable → the host's profile stays disabled and the host is ineligible. Report it as an unmet acceptance gate; never select a broader profile to work around it.
Validating the template without deploying
# Every slot: layer the mandatory producer overlay, which adds the writable /run/sources mount and
# requires DISPATCH_PREPARED_SOURCES_ROOT + DISPATCH_PERMIT_PUBLIC_KEYS.
docker compose -f deploy/queue-dispatch.compose.yml -f deploy/queue-dispatch.producer.compose.yml \
--env-file <operator env file> --profile dispatch-not-activated config
config renders the template and fails on any missing required variable. Pass the profile:
without it the rendered service list is empty and the check proves nothing. Render with the overlay:
it is what makes both producer variables required and mounts /run/sources, and
lib/dispatch-config.ts refuses every profile without them. It starts nothing, and on a host without
the operator's broker socket and journal directory this is the only check that should be run.
The child's own output
The child writes result.json (a DispatchResult) in its /output bind mount, and may write
transcript.jsonl next to it; the transcript only feeds the optional run-log (see the
DISPATCH_WORKER_LOG_DIR paragraph above) and is never part of the result. The
supervisor reads result.json from the broker's attempt root (DISPATCH_ATTEMPT_OUTPUT_DIR → DISPATCH_OUTPUT_ROOT=/run/attempts,
read-only) after the stop it observed, validates it with the shared result validator and never
executes or interpolates any of it. A cancelled run without a report still terminalizes, because
cancellation is the supervisor's own observation rather than a model claim; a normal run without one
fails with DISPATCH_RESULT_OUTPUT_MISSING.
Those exact bytes are then staged as an immutable report artifact through
PUT /dispatch/v1/attempts/collected/:key before the result is submitted. That is the only artifact
path open at this point: the stop this supervisor already submitted revoked the attempt that
PUT /attempts/artifacts/:key requires, so collection is bound to the historical start binding
instead. It is idempotent on identical bytes, and a failure there stays a failure — it never becomes
a local completion. A code result goes through the same route. On a repo_write
attempt the supervisor collects the stopped child's committed change from the prepared work area
into the exact JSON the central publisher accepts — a Git bundle of base..HEAD with the diff and
file list — and stages it as code. Base, branch and origin come from the pinned request and its
signed manifest, never from the child: a result naming a different base, branch or head is refused
(DISPATCH_CODE_ARTIFACT_REFUSED), and the artifact_id in the submitted result is the service's
own receipt rather than anything the child wrote. A code result on a read request, or on a slot
with no producer armed, is still refused (DISPATCH_CODE_ARTIFACT_REFUSED /
DISPATCH_CODE_ARTIFACT_UNSUPPORTED). Publication itself stays central: the supervisor pushes
nothing, holds no forge credential, and never merges or deploys.
Unmet gates
- No child ever receives its capability. The service mints an attempt-scoped
agent_tokenbeside the start permit and servesGET /dispatch/v1/agent/sources/:keyandPUT /dispatch/v1/agent/outputs/:key, but nothing here places that token in the container environment, and the broker has no mechanism to. Neither served profile needs it: the child writes its oneresult.jsoninto its own/outputmount, andcodeis collected post-stop by the supervisor under its own historical binding. The child's only credential is its own provider key (M41 decision B1), delivered as an operator source; see "Model image" in the operator doc. Do not setDISPATCH_AGENT_OUTPUT_KEY: the service then addsagent_tokento the start permit and this broker refuses any permit with fields other thanexpiresAt,permitId. checksis still staged by nobody. Collection knowsreportandcode; the child's checks travel insideresult.jsonand no separatechecksartifact is produced.- Non-launch recovery is not driven end to end from here.
POST /attempts/recovery/{lookup,stop,result}exists, the transport andassembleManagedPortsare wired to it, and the service's own tests cover it;recoverNonLaunchAttemptagainst the real routes is not in the cross-repo gate. - No attempt has run against a live service, and no attempt has ever run through this compose
template. The cross-repo gate boots the service in-process against a disposable test cluster with
a fake runtime: no container, no image, no model, no egress. Its repo_write case publishes into a
local bare repository over
file://, so Forgejo remote-head conflicts and draft-PR creation stay unmeasured, and the root broker's own copy/verify/mount step is simulated rather than executed. - Linux confinement is proved on scrum4me-server only (2026-10-02, M41 increment 3:
npm run test:dispatch-runtimeandnpm run test:dispatch-model-imagegreen); max2 has not been measured. - A meaningful Unix-socket liveness healthcheck for the egress proxy is still operator work; file
existence is not readiness. Provider proxy support in the release model image is proved for the
model-codexadapter (T-1869) and, since M41 5b, themodel-claudeadapter (probe child in the same test:npm run test:dispatch-model-image; live review on Docker Desktop:npm run smoke:dispatch-live-review, 2026-10-02, verdict NO-GO in 34 s, key absent from every output, container and attempt directory released).
Bekende grenzen
- Eén actieve job tegelijk. De wrapper-loop is sequentieel. Voor
parallellisme zou je meerdere containers met dezelfde
SCRUM4ME_TOKENkunnen draaien —wait_for_jobgebruiktFOR UPDATE SKIP LOCKEDdus dat is veilig op DB-niveau, maar dan moet je jenode_modules-cache per container scheiden. - OAuth-token: 1 jaar geldig. Bij verloop schrijft de wrapper een
TOKEN_EXPIRED-marker en wordt de containerunhealthy. Geen auto-rotatie. npm installper job kost op een N5095 ~30–60 s per Next.js-clone, óók met de pnpm-store. Voor zeer kleine fixes is dat de dominante factor. Kan later vervangen worden door een persistente warm-node_modulesper repo als dat een knelpunt wordt.