feat(IDEA-213): supervisors, runtimebroker en egress-proxy (ST-1590) #82

Merged
janpeter merged 33 commits from feat/idea-213-dispatch into master 2026-09-21 19:30:00 +02:00
Owner

Draft — niet mergen. Geopend om CI te laten draaien; IP-10 t/m IP-14 van ST-1590 zijn nog niet uitgevoerd. Hoort bij Scrum4Me #251; de zeven feat/idea-213-dispatch-branches horen bij elkaar.

Inhoud

Uitvoeringslaag (IP-07 t/m IP-09): supervisor voor geïsoleerde job- en hostpogingen, root-runtimebroker met journaal, egress-proxy met allowlist en DNS-pinning, bronstaging, receipts en niet-lancerend herstel.

Review-fixes (code-review 2026-09-20)

  • T-1841 549b597 — een child die ophangt tijdens een geweigerde CONNECT kon de proxy laten crashen (gereproduceerd met echte sockets).
  • T-1842 15c6f33 — prepare heeft een eigen budget en hervat; de broker joint een lopende create en gooit het residu van een mislukte weg.
  • T-1843 ba56d0e — de broker handhaaft zelf een gejournaliseerde deadline, ook na herstart. Nieuwe verplichte config-sleutel maxDurationSeconds.

Verificatie (lokaal)

  • npm test: 649 groen, twee keer; npm run typecheck:dispatch 0 fouten.
  • npm run test:dispatch-runtime tegen echte Docker: exit 0, alle 16 isolatie-probes true.
  • De deadline-stop en het weggooien van een PREPARING-residu zijn alleen met het nagebootste docker-commando bewezen.

Commits

  • 50442a1 feat(IDEA-213): supervise isolated job and host attempts
  • 5d0123f fix(IDEA-213): serialize scope stops and recover non-launch receipts
  • f03740f feat(IDEA-213): pin execution sources and preserve code artifacts
  • 14dfc3f fix(IDEA-213): recover source staging and align broker keys
  • a561806 fix(IDEA-213): preserve canonical receipts and unexported recovery evidence
  • 549b597 fix(IDEA-213): keep a disconnecting child from crashing the egress proxy
  • 15c6f33 fix(IDEA-213): let a timed-out or failed prepare be resumed
  • ba56d0e fix(IDEA-213): enforce the run deadline in the broker, not only the supervisor

🤖 Generated with Claude Code

Draft — niet mergen. Geopend om CI te laten draaien; IP-10 t/m IP-14 van ST-1590 zijn nog niet uitgevoerd. Hoort bij [Scrum4Me #251](https://git.jp-visser.nl/janpeter/Scrum4Me/pulls/251); de zeven `feat/idea-213-dispatch`-branches horen bij elkaar. ## Inhoud Uitvoeringslaag (IP-07 t/m IP-09): supervisor voor geïsoleerde job- en hostpogingen, root-runtimebroker met journaal, egress-proxy met allowlist en DNS-pinning, bronstaging, receipts en niet-lancerend herstel. ## Review-fixes (code-review 2026-09-20) - **T-1841** `549b597` — een child die ophangt tijdens een geweigerde CONNECT kon de proxy laten crashen (gereproduceerd met echte sockets). - **T-1842** `15c6f33` — prepare heeft een eigen budget en hervat; de broker joint een lopende create en gooit het residu van een mislukte weg. - **T-1843** `ba56d0e` — de broker handhaaft zelf een gejournaliseerde deadline, ook na herstart. **Nieuwe verplichte config-sleutel `maxDurationSeconds`.** ## Verificatie (lokaal) - `npm test`: 649 groen, twee keer; `npm run typecheck:dispatch` 0 fouten. - `npm run test:dispatch-runtime` tegen echte Docker: exit 0, alle 16 isolatie-probes true. - De deadline-stop en het weggooien van een PREPARING-residu zijn alleen met het nagebootste docker-commando bewezen. ## Commits - `50442a1` feat(IDEA-213): supervise isolated job and host attempts - `5d0123f` fix(IDEA-213): serialize scope stops and recover non-launch receipts - `f03740f` feat(IDEA-213): pin execution sources and preserve code artifacts - `14dfc3f` fix(IDEA-213): recover source staging and align broker keys - `a561806` fix(IDEA-213): preserve canonical receipts and unexported recovery evidence - `549b597` fix(IDEA-213): keep a disconnecting child from crashing the egress proxy - `15c6f33` fix(IDEA-213): let a timed-out or failed prepare be resumed - `ba56d0e` fix(IDEA-213): enforce the run deadline in the broker, not only the supervisor 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Node removes its socket error handler on 'connect' and the proxy only
added its own after the DNS lookup, never on the refusal path. A child
that sent CONNECT and hung up raised an uncaught EPIPE, killing the one
proxy process that carries the slot's egress.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Prepare had the 10 s default budget although it downloads sources,
re-hashes them, checks out git and makes several 30 s Docker calls.
After a timeout the journal kept a null scope and every rerun returned
uncertain without ever preparing again, while the broker finished and
orphaned a CREATED container. A failed or crashed create left the
broker record in PREPARING, which it then refused forever.

Prepare now has its own budget and runs again whenever the journal has
no scope. The broker joins a create that is still in flight and
discards the residue of one that failed; PREPARING precedes any start
permit, so nothing can launch twice.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(IDEA-213): enforce the run deadline in the broker, not only the supervisor
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 2s
CI / Docker build (pull_request) Successful in 1m48s
ba56d0e892
max_duration_seconds was checked in the supervisor's poll loop alone.
When the supervisor died, or a stop failed and the attempt went
uncertain, the container kept running with egress until something
happened to drive it again.

The broker now journals its own deadline before docker start, stops the
scope when it passes, and re-arms it from the journal after a restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
chore(IDEA-213): pin shared 5209199
All checks were successful
CI / Compose config (pull_request) Successful in 10s
CI / Build-arg coverage (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 2m7s
9aab0cab12
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
feat(IDEA-213): finalize the dispatch deployment template and its gates
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 2s
CI / Docker build (pull_request) Successful in 1m13s
c190be7919
The compose template now pins the image it runs: DISPATCH_SUPERVISOR_IMAGE is
required, so a release manifest supplies a digest and the build stanza beside it
only reproduces that same image. Its confinement is asserted by tests instead of
by reading: opt-in profile, uid 10001, read-only rootfs, all capabilities dropped,
no-new-privileges, bounded pids/memory/cpu, no Docker socket and no environment
value that could carry a database URL, a forge token or a dispatch key — the one
credential the supervisor holds arrives as an operator-owned secret file.

Fixes a real defect in that template: `tmpfs: [/tmp:rw,nosuid,nodev,size=64m]`
is a YAML flow sequence, so `docker compose config` rendered four tmpfs mounts,
three of them at paths called `nosuid`, `nodev` and `size=64m`.

README gains the deployment instructions: credentials that stay outside the
image, digest pinning, the now-required broker `maxDurationSeconds`, the per-host
`npm run test:dispatch-runtime` gate with its disabled-on-failure rule, the
template validation command (which needs the profile flag, or it renders nothing)
and the gates that remain unmet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Replace the blanket DISPATCH_RUNTIME_WIRING_REQUIRED_IP13 refusal with a real
assembly: an authenticated HTTP transport for the executor and attempt routes,
a configuration reader that refuses every cause before anything registers, a
child-output producer for stop evidence and results, and a job/host main block.

The transport sends exactly what the service's strict parsers accept and what
the queue CLI client sends where they overlap. Only calls carrying the caller's
own unique key -- the registration key and the claim key -- are replayed after a
transport loss; a start, a stop or a result is replayed from the durable journal
instead, so a lost response can never become a second decision. Error bodies are
reduced to status and code, because the request that produced them carries the
attempt credential.

The configuration reader derives the profile digest from the operator's pinned
profile rather than taking a second configured value that can drift, and refuses
a repo_write or source-mount profile outright: the prepared-sources producer
does not exist here, and refusing keeps the slot unregistered instead of
claiming work it would have to abandon after the model already ran.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
feat(IDEA-213): reconcile the dispatch deployment template and prove the wire
All checks were successful
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Compose config (pull_request) Successful in 10s
CI / Docker build (pull_request) Successful in 53s
7ed0918aa5
The compose template now supplies exactly the keys the configuration reader
refuses to start without, and a test holds both sides to one list so neither can
drift. New: DISPATCH_MODE, DISPATCH_BOOT_ID, the pinned profile directory, the
broker's attempt output root and the host listener directory.

SCRUM4ME_WORKER_INSTANCE_ID is removed from this service. The managed-only
supervisor never runs the ordinary worker loop and never sends an instance id;
the slot's pairing with a managed worker is server-side configuration. Keeping
it suggested a binding this process does not make. The ordinary worker service
is untouched.

The manual cross-repo gate boots the real mcp dispatch service against the
disposable cluster and measures this repository's transport against the actual
routes: the derived profile digest equals the service's own, register, executor
heartbeat and claim succeed through the real strict parsers, the same claim key
twice gives the same answer, and a refused start surfaces the service's code
without echoing the credential. It is opt-in behind DISPATCH_CONTRACT_MCP_ROOT
and a guard test warns when it did not run, so the skip cannot pass for proof.

The docs record what is assembled and what is not: the central result route
returns no canonical result, and neither the agent gateway nor the post-stop
collection path is mounted on any route, so an end-to-end managed attempt is
still blocked service-side.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The transport gains the three non-launch recovery calls and the post-stop
collection call, and the entrypoint assembles both. A claim answering
`authority:'none'` now reaches a real recovery instead of
DISPATCH_RECOVERY_TRANSPORT_REQUIRED.

`createOutputArtifacts` stages the child's exact `result.json` bytes as
`report` through PUT /attempts/collected/:key before the result is
submitted. That is the only artifact path open at this point: the stop
this supervisor already submitted revoked the attempt that
PUT /attempts/artifacts/:key requires. A failure there stays a failure
and never becomes a local completion. Collection is idempotent on
identical bytes, so a replayed attempt stages nothing twice.

`code` stays refused: collection stages `report` only, so a producer for
a repo_write profile genuinely does not exist yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
feat(IDEA-213): drive one full attempt in the cross-repo contract gate
All checks were successful
CI / Compose config (pull_request) Successful in 9s
CI / Build-arg coverage (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 53s
1ce05982e2
The opt-in gate now runs a whole read-only, artifact-publishing attempt
against the real mcp app over HTTP: intake, the service's own selection,
register, claim, start permit, a fake runtime whose child writes
result.json into its own output directory, stop evidence, post-stop
collection, result, canonical receipt and cleanup.

Authoritative state is read from the service's database through the
fixture's own control channel rather than from the receipt: request
terminal SUCCEEDED, exactly one canonical result, no open reservation,
no live attempt, the collected report as the only non-control attempt
artifact, and an empty journal directory.

The fixture serves one service at a time, because its host capacity key
is unique database-wide.

The operator doc drops the statements that are no longer true and keeps
the two that are: Linux confinement is unproven, and no attempt has run
against a LIVE deployed service.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs(IDEA-213): correct the header's standing contract claim
All checks were successful
CI / Build-arg coverage (pull_request) Successful in 2s
CI / Compose config (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 5s
4f930b43c1
The header still pointed at *Unmet service contract*, which no longer
exists. IP13 closed those gaps; what remains untrue is installation,
activation, a production slot and a live deployed service.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs(IDEA-213): drop unmet gates that the service surface now meets
All checks were successful
CI / Build-arg coverage (pull_request) Successful in 2s
CI / Compose config (pull_request) Successful in 10s
CI / Docker build (pull_request) Successful in 5s
845ec5d405
The unmet-gates list still said POST /attempts/result answers no
canonical result and that the transport therefore refuses with
DISPATCH_RESULT_RECEIPT_INCOMPLETE. mcp ee898fe superseded that: the
route answers {status, result_id, reason, canonical_result?}. The
refusal stays in the code as the answer to a service that has not been
updated, which is what the text now says.

Also corrected: the child output section now describes post-stop
collection of the child's exact bytes as `report`, and the recovery port
is described as served rather than as something IP13 must still
implement.

Kept, because they are still true: no child receives its agent_token,
there is no producer for `code`, non-launch recovery is not driven end
to end from here, Linux confinement is unproven, and no attempt has run
against a live service or through this compose template.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The supervisor could not honour a repo_write or source-mount profile at all: the
configuration reader refused it outright, and a result carrying `code` had no
producer behind it. Both halves now exist, and the refusals that remain are the
precise ones.

`lib/dispatch-prepared-sources.ts` materialises exactly the pinned sources of one
attempt. It reads the service's Ed25519-signed manifest under its own attempt
proof, downloads only the artifact ids that manifest names, and hands them to the
IP-08 staging step, which re-hashes every byte against its pin and publishes one
immutable directory. The claim's own source list must match that signed manifest
key for key and hash for hash, and a repo_write request without a signed
repository source is refused rather than started against an empty or newer tree.
No Git fetch, no model-supplied URL or path, and no forge or database credential
is anywhere on that path.

`lib/dispatch-code-artifact.ts` collects the stopped child's committed change into
the exact JSON the central publisher verifies. The child's repository is untrusted
at collection, so its configuration is read raw before any Git command that could
execute a filter, hook or helper it names. Base, branch and origin come from the
pinned request and its signed manifest; a result naming a different base, branch
or head is refused, and the `artifact_id` that goes back is the service's own
collection receipt. Publication stays central: nothing here pushes, merges or
deploys, and the child receives no credential of any kind.

Two operator keys arm the producer, all-or-nothing, and a slot that has not armed
it still refuses the profile before it registers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
feat(IDEA-213): prove one repo_write attempt against the real dispatch service
All checks were successful
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Compose config (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 32s
5d7f76d56c
The cross-repo gate drove a read-only attempt only, so nothing measured the two
producers this task added against the service that has to accept their output.

A second fixture mode submits a repo_write request pinned to a base commit in a
local bare repository, lets the service's own workspace producer prepare the
repository source, and then runs the real prepared-sources producer against the
real signed-manifest and artifact routes. The fake runtime stands in for the root
broker: it stages and materialises exactly as `createRuntimeBroker` does, because
no Docker daemon is involved. The child commits one file, the real collector turns
that commit into the code artifact, and the service's own central publisher pushes
the branch.

What the gate reads back is authoritative database and repository state, not the
receipt: one canonical result, `code` and `report` as the only non-control attempt
artifacts, one CONFIRMED branch publication at the child's commit, that commit
really present in the bare repository, no merge or deploy job, no open reservation,
no live attempt and an empty journal. Weakening the collector so the publisher's
own verification rejects it leaves the attempt at `pending_receipt` instead.

A real forge remote is out of reach here, so remote-head conflict handling and
draft-PR creation stay unmeasured; the docs now say so.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(IDEA-213): deny reserved IPv4 ranges and bound egress tunnels
Some checks failed
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Has been cancelled
411c5fa3be
publicIPv4 admitted 192.0.0.0/24, 192.0.2.0/24, 198.18.0.0/15,
198.51.100.0/24 and 203.0.113.0/24, so a provider name resolving into
special-purpose space became a permitted upstream connect. Deny them all
and export the predicate so the boundary is table-tested.

Tunnels had no idle timeout and no per-slot ceiling: a model child could
hold sockets open indefinitely and open unbounded ones. Both halves now
carry an idle timeout (any byte rearms it) and a CONNECT beyond the slot
limit is refused with 503 until a tunnel ends. Both are proved with real
sockets in the child-process harness, with injected short timeouts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(IDEA-213): schema-check the staged registry and bound the repository origin
All checks were successful
CI / Compose config (pull_request) Successful in 13s
CI / Build-arg coverage (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 2m4s
5424d287c1
loadStagedSources cast manifest.json straight from JSON, so the root
broker dereferenced `pins` exactly as written. Parse it against a strict
schema and require every pin to resolve to <root>/<attempt>/<key>, which
refuses a rewritten registry that points outside the staging root or
reaches out through a symlink.

repository.repoUrl had no pattern and is written verbatim into the
child-readable /work/.git/config, so userinfo in the URL would hand a
credential to the child. Require an exact ordinary https URL without
userinfo, query or fragment, refused in verifyPreparedManifest before
anything is materialised. No host allowlist: no operator or profile data
names repository hosts (provider_egress_hosts is model egress), the
origin is signed metadata that nothing fetches, and inventing a second
config surface would only add drift.

Test fixtures that used a file:// origin now use https, as production does.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(IDEA-213): handle every stop signal, add --init and name the socket group
Some checks failed
CI / Compose config (pull_request) Successful in 10s
CI / Build-arg coverage (pull_request) Successful in 12s
CI / Docker build (pull_request) Has been cancelled
76a5a890f3
runManagedWithSignals only attached to SIGTERM, so SIGINT (a terminal)
and SIGHUP (a reload or closed session) killed the supervisor by default
mid-attempt with the scope still running. All three now take the same
stop route and all three are detached again, in the entrypoint loop too.

Containers were created without --init, so PID 1 was the model itself:
nothing reaped its orphans and nothing forwarded the stop signal, which
made the ten-second docker stop grace period meaningless inside the
container. --init adds no capability, mount or namespace, and all 16
confinement probe checks of test:dispatch-runtime stay true.

The broker socket's group decided who may talk to the broker, yet
nothing set it: access silently depended on a setgid directory the
operator doc never mentioned. Chosen route: explicit chown from operator
config. The broker config gains a required numeric socketGroup (>=1) and
the broker chowns the socket after listen; a missing, zero or unusable
group refuses with DISPATCH_BROKER_SOCKET_GROUP_REFUSED and the broker
does not serve. A setgid requirement would have stayed invisible in the
same way, so it was not the option taken.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
janpeter force-pushed feat/idea-213-dispatch from 76a5a890f3
Some checks failed
CI / Compose config (pull_request) Successful in 10s
CI / Build-arg coverage (pull_request) Successful in 12s
CI / Docker build (pull_request) Has been cancelled
to 4884f6f8c9
All checks were successful
CI / Compose config (pull_request) Successful in 9s
CI / Build-arg coverage (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 2m2s
2026-09-20 22:17:54 +02:00
Compare
fix(IDEA-213): keep the repo_write contract gate green under the bounded origin rule
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Successful in 54s
2387016cc0
Bisected: the gate was 4/4 at 5d7f76d and at 411c5fa (m1) and went red at
5424d28 (m2), with the mcp worktree held constant, so the moving mcp tip
is ruled out. Cause: m2's origin rule was a second, unconfigurable copy
of transport policy. The gate has no forge, so the service's product
repo_url is a local bare repo over file://, which the signed manifest
then carries; verifyPreparedManifest refused it and prepare threw, which
the supervisor's catch-all reports as uncertain.

Split the rule where the design already splits it. The shape rule is a
property of the value and stays unconditional at verification: absolute
URL, no userinfo, no query or fragment, nothing that reads as a git
option. The protocol is transport policy and now sits where the origin
is actually written, on materializePreparedRepository, defaulting to
REPOSITORY_ORIGIN_PROTOCOLS = https only. The operator broker passes no
widening, so production is unchanged and strict; the gate passes
['file:','https:'] for its bare repo, exactly as the central publisher
opts that protocol in through createGitPublicationPort's own
allowedProtocols. Widening the protocols never relaxes the shape rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(IDEA-213): close a provably unstarted attempt as a precise failure
All checks were successful
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Compose config (pull_request) Successful in 10s
CI / Docker build (pull_request) Successful in 53s
026d9ebbaf
The single catch in runDispatchAttempt turned every throw into phase='uncertain',
so an attempt whose container never started kept its reservation until someone
cancelled or recovered it by hand.

The authority on "nothing was launched" is the broker, not the journal: when its
own stop observation for the journalled scope reports status='created', the
container exists and never ran. Such an attempt is now finished centrally through
the path the service actually accepts - stop evidence, then a failed result whose
summary is a bounded DISPATCH_* code - which releases the slot with exactly one
canonical result. Anything else stays uncertain: a scope that ran, a joined
existing scope, an already staged result, and any refusal without a bounded code,
so no raw error text, path or credential can travel centrally.

Measured against the real service: POST /attempts/result refuses a result without
accepted stop evidence (DISPATCH_STATE_CONFLICT), and every stop-evidence shape
binds to a scope id. A refusal before any scope exists therefore still has no
closure path; that gap is documented rather than papered over with invented
evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
chore(IDEA-213): pin shared ee7fe1a
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Successful in 6s
0d4882b3da
Bump vendor/scrum4me-shared 5209199 -> ee7fe1a (feat/idea-213-dispatch).
Shared contract tightening m9a/m9b/m9c/m10; lib-only. Docker has no generated
prisma schema; the vendored canonical prisma/schema.prisma is byte-identical
(72 models / 47 enums).

npm test 794 passed/4 skipped (was 788; +6 are the new vendored shared contract
tests the bump carries: m9a x4, m9b x1, m10 x1). typecheck:dispatch, skills:verify,
test:dispatch-runtime all green. Cross-repo contract gate (m9a signer<->verifier,
m9b) passes 5/5 against scrum4me-mcp at the same pin -- the coherence proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
fix(IDEA-213): close a pre-scope refusal with claim-bound stop evidence
All checks were successful
CI / Compose config (pull_request) Successful in 9s
CI / Build-arg coverage (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 1m47s
158a359603
When source preparation refuses before the broker create (a provably pre-scope
DISPATCH_PREPARED_SOURCES_REFUSED), no runtime scope exists, so closeUnstarted's
`created` evidence does not apply and the attempt fell through to `uncertain`,
holding its reservation and looping the supervisor on retry.

Add a claim-bound close: the supervisor submits stop evidence bound to the claim
(POST /attempts/claim-stop) plus the bounded reason, then the same bounded
failure reaches FAILED. The observedAt is journalled once for byte-identical
retries and the journal is removed on receipt. This fires ONLY for a reason that
provably precedes create; a prepare timeout, a broker-create error, an unbounded
throw, a journalled scope or a reconciliation stays uncertain — never closed
without real stop evidence.

The opt-in contract gate now proves the pre-scope prepared-source refusal ends
FAILED (one canonical result summarizing the bounded code, reservation released,
zero model starts, zero publications, empty journal), with the broker-created-
but-never-started case kept as a separate closeUnstarted regression.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
docs(IDEA-213): the pre-scope refusal is now closed by a claim-bound stop
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 10s
CI / Docker build (pull_request) Successful in 32s
c4e6045087
T-1863 (mcp 379cbe4, docker 158a359) shipped the claim-bound stop path,
so the operator doc's "Known gap, for JP" and the "no closure path" note
are outdated. Describe the shipped closeClaimBound behaviour and keep
uncertain as the answer only for a genuinely ambiguous failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
chore(IDEA-213): pin shared 9c3ce16
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Successful in 6s
83c296c8b3
Bump vendored scrum4me-shared to 9c3ce16 (feat/idea-213-dispatch), which
adds the DispatchState value UNCERTAIN_UNSTARTED and narrows the state
table. Docker consumes the shared lib as types only; no consumer edits
required. Cross-repo dispatch-contract gate stays green (6/6).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
feat(IDEA-213): verify the permit and manifest against a trusted key-id keyset
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 9s
CI / Docker build (pull_request) Successful in 33s
f32472e044
Pins shared 925af45 and turns the broker/supervisor verifiers into keyset
verifiers for zero-downtime Ed25519 rotation:

- The single pin file DISPATCH_PERMIT_PUBLIC_KEY_FILE is removed in favour of
  DISPATCH_PERMIT_PUBLIC_KEYS: comma-separated kid:base64url(DER SPKI) entries,
  parsed like DISPATCH_CREDENTIAL_KEYS. The keyset is the one source of truth and
  holds every currently-trusted key, so the next key can be armed before the
  signer flips to it. A private key in that value is still refused. The broker's
  root-owned config carries the same keyset inline (permitPublicKeys).
- verifyBrokerPermit and verifyPreparedManifest select the public key by the
  token's own kid (algorithm stays Ed25519) and refuse an unknown kid; the broker
  keyset holds both the permit kid and the manifest kid.
- The prepared-sources producer, staging, materialisation and code collection all
  take the keyset; every derived binding schema omits kid.
- Contract gate extended: a real permit and manifest are each verified under an
  explicit kid, and a permit/manifest signed under a kid NOT in the verifier's
  keyset is refused (rotation-safety). Gate is green at shared pin 925af45.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
janpeter changed title from WIP: feat(IDEA-213): supervisors, runtimebroker en egress-proxy (ST-1590) to feat(IDEA-213): supervisors, runtimebroker en egress-proxy (ST-1590) 2026-09-21 18:00:10 +02:00
s4m-codex-reviewer requested changes 2026-09-21 18:02:14 +02:00
Dismissed
s4m-codex-reviewer left a comment

REQUEST_CHANGES

geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden.

Findings

  • error — lib/dispatch-runtime.ts:10: buildChildEnv() wijst HTTP_PROXY en HTTPS_PROXY naar http://127.0.0.1:18081, maar de container die met buildContainerArgs() wordt gestart krijgt alleen de egress Unix-socket gemount op /run/dispatch en er wordt geen bridge/listener op 127.0.0.1:18081 gestart. startEgressBridge() bestaat in lib/dispatch-egress-proxy.ts:81, maar wordt nergens in de child-startup geïnjecteerd. Daardoor hebben normale HTTP(S)-clients in de attempt geen werkende proxy en faalt provider-egress ondanks de gemonteerde socket. De bestaande runtime probe test direct tegen /run/dispatch/egress.sock, waardoor dit pad niet wordt afgedekt.

  • error — deploy/queue-dispatch.compose.yml:66: de compose-template vereist ${DISPATCH_PREPARED_SOURCES_DIR:?prepared sources staging directory required} als host mount, terwijl de supervisorconfig en README de operatorvariabele DISPATCH_PREPARED_SOURCES_ROOT gebruiken voor prepared sources. Met alleen de gedocumenteerde variabelen faalt docker compose ... config; met alleen DISPATCH_PREPARED_SOURCES_DIR blijft DISPATCH_PREPARED_SOURCES_ROOT leeg en refuseert de supervisor repo_write/source-mount profielen. Dit maakt de rollout-template niet conform de eigen configuratiecontracten.

Verificatie

  • Productdocs gelezen: architecture/overview; er is geen specifiek dispatch-productdoc in de actieve docs-index gevonden via full-text search.
  • Lokale checkout van PR-head f32472e044a19e46eafc33385b1bd4d06759e1ca geïnspecteerd.
  • npm ci && npm run typecheck:dispatch geprobeerd, maar npm ci faalde in deze tijdelijke omgeving op node_modules/esbuild/bin/esbuild met EACCES; daarom geen groen testresultaat kunnen bevestigen.
# REQUEST_CHANGES geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden. ## Findings - **error** — `lib/dispatch-runtime.ts:10`: `buildChildEnv()` wijst `HTTP_PROXY` en `HTTPS_PROXY` naar `http://127.0.0.1:18081`, maar de container die met `buildContainerArgs()` wordt gestart krijgt alleen de egress Unix-socket gemount op `/run/dispatch` en er wordt geen bridge/listener op `127.0.0.1:18081` gestart. `startEgressBridge()` bestaat in `lib/dispatch-egress-proxy.ts:81`, maar wordt nergens in de child-startup geïnjecteerd. Daardoor hebben normale HTTP(S)-clients in de attempt geen werkende proxy en faalt provider-egress ondanks de gemonteerde socket. De bestaande runtime probe test direct tegen `/run/dispatch/egress.sock`, waardoor dit pad niet wordt afgedekt. - **error** — `deploy/queue-dispatch.compose.yml:66`: de compose-template vereist `${DISPATCH_PREPARED_SOURCES_DIR:?prepared sources staging directory required}` als host mount, terwijl de supervisorconfig en README de operatorvariabele `DISPATCH_PREPARED_SOURCES_ROOT` gebruiken voor prepared sources. Met alleen de gedocumenteerde variabelen faalt `docker compose ... config`; met alleen `DISPATCH_PREPARED_SOURCES_DIR` blijft `DISPATCH_PREPARED_SOURCES_ROOT` leeg en refuseert de supervisor repo_write/source-mount profielen. Dit maakt de rollout-template niet conform de eigen configuratiecontracten. ## Verificatie - Productdocs gelezen: `architecture/overview`; er is geen specifiek dispatch-productdoc in de actieve docs-index gevonden via full-text search. - Lokale checkout van PR-head `f32472e044a19e46eafc33385b1bd4d06759e1ca` geïnspecteerd. - `npm ci && npm run typecheck:dispatch` geprobeerd, maar `npm ci` faalde in deze tijdelijke omgeving op `node_modules/esbuild/bin/esbuild` met `EACCES`; daarom geen groen testresultaat kunnen bevestigen.
fix(dispatch): align prepared-sources mount var with supervisor env
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Successful in 6s
e8aecf42e0
The /run/sources host mount used DISPATCH_PREPARED_SOURCES_DIR while the
supervisor env, dispatch-config.ts and the README use
DISPATCH_PREPARED_SOURCES_ROOT for the same staging dir. With only the
documented _ROOT set, `docker compose config` failed on the mount's
:? guard; with only _DIR, _ROOT stayed empty and the supervisor
refused repo_write/source-mount profiles. Use _ROOT in both places.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
s4m-codex-reviewer requested changes 2026-09-21 18:32:03 +02:00
Dismissed
s4m-codex-reviewer left a comment

Verdict: REQUEST_CHANGES

Geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden.

Findings

  • error — bin/run-dispatch-attempt.ts:160: de supervisor zet busy=true alleen rond ports.claim(...) en direct daarna weer false. Daardoor rapporteert de executor-heartbeat tijdens een lopende attempt als niet-busy, terwijl de job-supervisor in dezelfde loop wel één attempt bezet houdt tot runManagedWithSignals terugkeert. Dit breekt de liveness/busy-contracten uit de operator-docs en kan de centrale dispatcher een bezette slot als beschikbaar laten zien. Zet de busy-state rond de volledige attempt-run, niet alleen rond de claim-call, en voeg een regressietest toe die runJobSupervisor zelf tijdens een langlopende attempt heartbeat-busy laat rapporteren.

  • error — deploy/queue-dispatch.compose.yml:66: DISPATCH_PREPARED_SOURCES_ROOT wordt in de volume-mount verplicht gemaakt met :?, terwijl lib/dispatch-config.ts en de comments erboven expliciet toestaan dat prepared-sources ontbreekt voor profielen zonder repo_write/source mounts. Een read-only/artifact-only slot zonder producer kan met deze compose-template niet starten. Maak deze mount optioneel via een aparte compose-override/profile, of zorg dat de documentatie/config niet langer belooft dat de producer kan ontbreken.

  • warning — deploy/queue-dispatch.compose.yml:69: dezelfde template vereist altijd DISPATCH_HOST_SOCKET_DIR, ook voor DISPATCH_MODE=job, terwijl de comment zegt dat de host route alleen voor host mode is. Dit maakt eenvoudige job-mode deploys onnodig afhankelijk van een dummy host-socket directory. Overweeg ook deze mount mode-specifiek te maken of de operator-instructie aan te passen.

Samenvatting

De runtime/broker-code is uitgebreid getest en volgt op veel plekken de bestaande runnergrenzen, maar bovenstaande contractbreuken raken scheduling en deploybaarheid. Daarom geen approval in deze vorm.

# Verdict: REQUEST_CHANGES Geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden. ## Findings - **error** — `bin/run-dispatch-attempt.ts:160`: de supervisor zet `busy=true` alleen rond `ports.claim(...)` en direct daarna weer `false`. Daardoor rapporteert de executor-heartbeat tijdens een lopende attempt als niet-busy, terwijl de job-supervisor in dezelfde loop wel één attempt bezet houdt tot `runManagedWithSignals` terugkeert. Dit breekt de liveness/busy-contracten uit de operator-docs en kan de centrale dispatcher een bezette slot als beschikbaar laten zien. Zet de busy-state rond de volledige attempt-run, niet alleen rond de claim-call, en voeg een regressietest toe die `runJobSupervisor` zelf tijdens een langlopende attempt heartbeat-busy laat rapporteren. - **error** — `deploy/queue-dispatch.compose.yml:66`: `DISPATCH_PREPARED_SOURCES_ROOT` wordt in de volume-mount verplicht gemaakt met `:?`, terwijl `lib/dispatch-config.ts` en de comments erboven expliciet toestaan dat prepared-sources ontbreekt voor profielen zonder `repo_write`/source mounts. Een read-only/artifact-only slot zonder producer kan met deze compose-template niet starten. Maak deze mount optioneel via een aparte compose-override/profile, of zorg dat de documentatie/config niet langer belooft dat de producer kan ontbreken. - **warning** — `deploy/queue-dispatch.compose.yml:69`: dezelfde template vereist altijd `DISPATCH_HOST_SOCKET_DIR`, ook voor `DISPATCH_MODE=job`, terwijl de comment zegt dat de host route alleen voor host mode is. Dit maakt eenvoudige job-mode deploys onnodig afhankelijk van een dummy host-socket directory. Overweeg ook deze mount mode-specifiek te maken of de operator-instructie aan te passen. ## Samenvatting De runtime/broker-code is uitgebreid getest en volgt op veel plekken de bestaande runnergrenzen, maar bovenstaande contractbreuken raken scheduling en deploybaarheid. Daarom geen approval in deze vorm.
chore(IDEA-213): re-pin shared c65f5a8 (F-S1)
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Successful in 5s
4baa39a580
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
s4m-codex-reviewer requested changes 2026-09-21 18:36:56 +02:00
Dismissed
s4m-codex-reviewer left a comment

Verdict: REQUEST_CHANGES

geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden.

Findings

  • error — deploy/queue-dispatch.compose.yml:66 — De compose-template maakt DISPATCH_PREPARED_SOURCES_ROOT verplicht met ${DISPATCH_PREPARED_SOURCES_ROOT:?…} in de volume mount, terwijl dezelfde service deze waarde in environment optioneel doorgeeft (${DISPATCH_PREPARED_SOURCES_ROOT:-}) en lib/dispatch-config.ts bewust een no-producer/read-only configuratie toestaat zolang het profiel geen repo_write of source_mount_keys vraagt. Daardoor faalt docker compose ... --profile dispatch-not-activated config al voor een geldige read-only slot zonder prepared-sources producer, in strijd met de README/operator-docs die de producer als all-or-nothing maar optioneel beschrijven. Maak deze mount conditioneel via een apart producer-profiel/servicevariant, of zorg dat read-only slots zonder producer de template kunnen renderen zonder deze variabele.

Review log

De diff is niet plan-gekoppeld. Ik heb vooral de nieuwe dispatch-runtime, supervisor/broker/proxy code, Dockerfile en compose-template beoordeeld op contractconsistentie, secret boundaries en test/deploybaarheid. De blocker zit in de operator-template en verhindert een ondersteunde configuratie voordat de runtime zelf kan weigeren of registreren.

# Verdict: REQUEST_CHANGES geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden. ## Findings - **error** — `deploy/queue-dispatch.compose.yml:66` — De compose-template maakt `DISPATCH_PREPARED_SOURCES_ROOT` verplicht met `${DISPATCH_PREPARED_SOURCES_ROOT:?…}` in de volume mount, terwijl dezelfde service deze waarde in `environment` optioneel doorgeeft (`${DISPATCH_PREPARED_SOURCES_ROOT:-}`) en `lib/dispatch-config.ts` bewust een no-producer/read-only configuratie toestaat zolang het profiel geen `repo_write` of `source_mount_keys` vraagt. Daardoor faalt `docker compose ... --profile dispatch-not-activated config` al voor een geldige read-only slot zonder prepared-sources producer, in strijd met de README/operator-docs die de producer als all-or-nothing maar optioneel beschrijven. Maak deze mount conditioneel via een apart producer-profiel/servicevariant, of zorg dat read-only slots zonder producer de template kunnen renderen zonder deze variabele. ## Review log De diff is niet plan-gekoppeld. Ik heb vooral de nieuwe dispatch-runtime, supervisor/broker/proxy code, Dockerfile en compose-template beoordeeld op contractconsistentie, secret boundaries en test/deploybaarheid. De blocker zit in de operator-template en verhindert een ondersteunde configuratie voordat de runtime zelf kan weigeren of registreren.
Author
Owner

Review afgehandeld — F-D2 gerepareerd; F-D1 is by design.

  • F-D2 (ERROR): deploy/queue-dispatch.compose.yml r66 mountte via DISPATCH_PREPARED_SOURCES_DIR terwijl de supervisor-env, lib/dispatch-config.ts en de README DISPATCH_PREPARED_SOURCES_ROOT gebruiken. Gelijkgetrokken op _ROOT (enige plek in de repo). Commit e8aecf4, npm test 815 passed | 6 skipped.
  • F-D1 (proxy 127.0.0.1:18081 zonder in-container listener): dit is een bewust contractpunt, geen defect. docs/dispatch-runtime-operator.md stelt expliciet dat het model-image startEgressBridge() intern moet starten (binnen de netns) en de proxy-env moet honoreren — en markeert dat óók als "not proved for a release model/provider image". De broker mount alleen de socket + zet de proxy-env; de host-kant (startEgressProxy) is wél gewired en wordt door de broker geverifieerd. Geen codewijziging; het blijft een gedocumenteerd, nog te bewijzen integratiepunt bij een release-model-image.

Herpind op shared c65f5a8 (4baa39a); cross-repo contractgate 7/7 groen.

**Review afgehandeld — F-D2 gerepareerd; F-D1 is by design.** - **F-D2 (ERROR)**: `deploy/queue-dispatch.compose.yml` r66 mountte via `DISPATCH_PREPARED_SOURCES_DIR` terwijl de supervisor-env, `lib/dispatch-config.ts` en de README `DISPATCH_PREPARED_SOURCES_ROOT` gebruiken. Gelijkgetrokken op `_ROOT` (enige plek in de repo). Commit `e8aecf4`, `npm test` 815 passed | 6 skipped. - **F-D1 (proxy 127.0.0.1:18081 zonder in-container listener)**: dit is een bewust contractpunt, geen defect. `docs/dispatch-runtime-operator.md` stelt expliciet dat het **model-image** `startEgressBridge()` intern moet starten (binnen de netns) en de proxy-env moet honoreren — en markeert dat óók als *"not proved for a release model/provider image"*. De broker mount alleen de socket + zet de proxy-env; de host-kant (`startEgressProxy`) is wél gewired en wordt door de broker geverifieerd. Geen codewijziging; het blijft een gedocumenteerd, nog te bewijzen integratiepunt bij een release-model-image. Herpind op shared `c65f5a8` (`4baa39a`); cross-repo contractgate 7/7 groen.
fix(dispatch): make prepared-sources mount conditional via producer overlay
All checks were successful
CI / Compose config (pull_request) Successful in 11s
CI / Build-arg coverage (pull_request) Successful in 8s
CI / Docker build (pull_request) Successful in 1m58s
3301bdc3b4
The base queue-dispatch.compose.yml required DISPATCH_PREPARED_SOURCES_ROOT
(`:?`) for the /run/sources mount, so `docker compose config` failed for a
valid read-only-only slot with no producer — while lib/dispatch-config.ts
explicitly accepts such a slot (it only refuses repo_write / source_mount
profiles without a producer). Compose profiles gate whole services, not a
single volume, so move the writable /run/sources mount into a producer
overlay (deploy/queue-dispatch.producer.compose.yml). The base renders
producer-less; layering the overlay adds the mount and upgrades both
producer variables from optional to required. README arming/validation docs
updated (incl. correcting DISPATCH_PERMIT_PUBLIC_KEY_FILE -> the actual
env-value key DISPATCH_PERMIT_PUBLIC_KEYS), plus deployment tests for both.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Author
Owner

Re-review-blocker afgehandeld (F-D2, dieper). Terecht: de verplichte ${DISPATCH_PREPARED_SOURCES_ROOT:?}-mount botste met lib/dispatch-config.ts, dat een read-only slot zónder producer toestaat → docker compose config faalde voor een geldige read-only-only deployment.

Fix (3301bdc): de /run/sources-mount is uit de base-template gehaald en verplaatst naar een additieve producer-overlay deploy/queue-dispatch.producer.compose.yml (Compose profiles gelden per service, niet per volume, dus een overlay is het schoonste). De base rendert nu read-only zonder de var; de overlay voegt de mount toe en maakt DISPATCH_PREPARED_SOURCES_ROOT + DISPATCH_PERMIT_PUBLIC_KEYS verplicht (:?), zodat de all-or-nothing-invariant uit dispatch-config.ts op render-tijd wordt afgedwongen. Bewezen met echte docker compose config (read-only rendert zonder mount; producer voegt /run/sources toe; overlay-zonder-KEYS faalt). README bijgewerkt (base vs. overlay render-commando's) en een pre-existing doc-drift DISPATCH_PERMIT_PUBLIC_KEY_FILE→DISPATCH_PERMIT_PUBLIC_KEYS gecorrigeerd. 2 tests toegevoegd, npm test 817 passed | 6 skipped.

**Re-review-blocker afgehandeld (F-D2, dieper).** Terecht: de verplichte `${DISPATCH_PREPARED_SOURCES_ROOT:?}`-mount botste met `lib/dispatch-config.ts`, dat een read-only slot zónder producer toestaat → `docker compose config` faalde voor een geldige read-only-only deployment. Fix (`3301bdc`): de `/run/sources`-mount is uit de base-template gehaald en verplaatst naar een additieve producer-overlay `deploy/queue-dispatch.producer.compose.yml` (Compose `profiles` gelden per service, niet per volume, dus een overlay is het schoonste). De base rendert nu read-only zonder de var; de overlay voegt de mount toe en maakt `DISPATCH_PREPARED_SOURCES_ROOT` + `DISPATCH_PERMIT_PUBLIC_KEYS` verplicht (`:?`), zodat de all-or-nothing-invariant uit `dispatch-config.ts` op render-tijd wordt afgedwongen. Bewezen met echte `docker compose config` (read-only rendert zonder mount; producer voegt `/run/sources` toe; overlay-zonder-KEYS faalt). README bijgewerkt (base vs. overlay render-commando's) en een pre-existing doc-drift `DISPATCH_PERMIT_PUBLIC_KEY_FILE`→`DISPATCH_PERMIT_PUBLIC_KEYS` gecorrigeerd. 2 tests toegevoegd, `npm test` 817 passed | 6 skipped.
s4m-codex-reviewer requested changes 2026-09-21 18:51:38 +02:00
Dismissed
s4m-codex-reviewer left a comment

Verdict: REQUEST_CHANGES

geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden.

Findings

  • error — bin/run-dispatch-runtime-broker.ts:39: serveRuntimeBroker accepteert socketGroup=0, omdat de helper alleen <0 weigert. De CLI-main weigert socketGroup<1 en de docs noemen gid 1 of hoger terecht als harde eis: gid 0 zou de runtime-socket aan de root-groep geven en daarmee de expliciete supervisor-groep als autorisatiegrens omzeilen. Omdat serveRuntimeBroker geëxporteerd is en ook in tests/andere assemblages wordt gebruikt, moet dezelfde >=1-validatie in de helper zelf zitten, bij voorkeur met een regressietest.

Opmerkingen

De PR bevat veel relevante testdekking en operator-documentatie voor dispatch-supervisor, runtime broker, prepared sources en egress. Zonder gekoppeld plan heb ik geen acceptatiecriteria kunnen afvinken; de review is daarom beperkt tot diff, productdocs en de zichtbare veiligheids-/architectuurclaims.

# Verdict: REQUEST_CHANGES geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden. ## Findings - **error** — `bin/run-dispatch-runtime-broker.ts:39`: `serveRuntimeBroker` accepteert `socketGroup=0`, omdat de helper alleen `<0` weigert. De CLI-main weigert `socketGroup<1` en de docs noemen gid 1 of hoger terecht als harde eis: gid 0 zou de runtime-socket aan de root-groep geven en daarmee de expliciete supervisor-groep als autorisatiegrens omzeilen. Omdat `serveRuntimeBroker` geëxporteerd is en ook in tests/andere assemblages wordt gebruikt, moet dezelfde `>=1`-validatie in de helper zelf zitten, bij voorkeur met een regressietest. ## Opmerkingen De PR bevat veel relevante testdekking en operator-documentatie voor dispatch-supervisor, runtime broker, prepared sources en egress. Zonder gekoppeld plan heb ik geen acceptatiecriteria kunnen afvinken; de review is daarom beperkt tot diff, productdocs en de zichtbare veiligheids-/architectuurclaims.
fix(dispatch): refuse gid 0 in serveRuntimeBroker socket-group check
All checks were successful
CI / Compose config (pull_request) Successful in 3s
CI / Build-arg coverage (pull_request) Successful in 3s
CI / Docker build (pull_request) Successful in 33s
e014c66ac5
The exported serveRuntimeBroker helper validated socketGroup with
`<0`, so gid 0 (the root group) was accepted while the CLI-main config
check and the operator docs require gid >=1. gid 0 would hand the
runtime socket to root's own group and bypass the explicit supervisor
group as authorization boundary. Since the helper is exported and used
by tests/other assemblies, the >=1 floor must live in the helper itself.

Change the check to `<1` and throw the meaningful
DISPATCH_BROKER_SOCKET_GROUP_REFUSED code (matching the CLI-main and the
surrounding catch) instead of a bare Error(). Add a red-first regression
test that stubs chown so only the validation can refuse gid 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Author
Owner

Re-review-blocker afgehandeld (security). Terecht: serveRuntimeBroker (bin/run-dispatch-runtime-broker.ts:37) weigerde alleen socketGroup<0, dus gid 0 (root-groep) werd geaccepteerd — in tegenstelling tot de CLI-main (r54, <1) en de docs. Helper-check nu socketGroup<1 met DISPATCH_BROKER_SOCKET_GROUP_REFUSED (dezelfde code als CLI-main en de omliggende catch). Regressietest toegevoegd (__tests__/broker-socket-group-floor.test.ts) die chown mockt zodat uitsluitend de >=1-validatie gid 0 kan weigeren — anders geeft chown EPERM al een false-green. Rood-eerst bevestigd. Commit e014c66, npm test 819 passed | 6 skipped.

**Re-review-blocker afgehandeld (security).** Terecht: `serveRuntimeBroker` (`bin/run-dispatch-runtime-broker.ts:37`) weigerde alleen `socketGroup<0`, dus gid 0 (root-groep) werd geaccepteerd — in tegenstelling tot de CLI-main (r54, `<1`) en de docs. Helper-check nu `socketGroup<1` met `DISPATCH_BROKER_SOCKET_GROUP_REFUSED` (dezelfde code als CLI-main en de omliggende catch). Regressietest toegevoegd (`__tests__/broker-socket-group-floor.test.ts`) die `chown` mockt zodat uitsluitend de `>=1`-validatie gid 0 kan weigeren — anders geeft `chown` EPERM al een false-green. Rood-eerst bevestigd. Commit `e014c66`, `npm test` 819 passed | 6 skipped.
s4m-codex-reviewer left a comment

REQUEST_CHANGES

geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden.

Findings

  • BLOCKER — lib/dispatch-runtime.ts:9 / lib/dispatch-runtime.ts:24: de child-container krijgt HTTP_PROXY en HTTPS_PROXY op http://127.0.0.1:18081, maar buildContainerArgs start alleen het child-image met --network none en mount de egress-volume read-only. Er wordt geen startEgressBridge() in dezelfde netwerknamespace gestart om 127.0.0.1:18081 door te verbinden naar /run/dispatch/egress.sock; bin/run-dispatch-egress-proxy.ts:18 start alleen de Unix-socket-proxy in de aparte egress-container. Daardoor falen normale model/tool HTTP(S)-requests vanuit de child met connection refused, terwijl de huidige integration probe dit mist omdat die rechtstreeks op de Unix-socket test. Dit blokkeert activatie van de managed runtime.

Verdict

REQUEST_CHANGES — de PR introduceert de dispatch-runtime, maar de netwerk-egress wiring voor echte child-processen is niet functioneel.

# REQUEST_CHANGES geen gekoppeld plan gevonden — beoordeeld op codekwaliteit + product-standaarden. ## Findings - **BLOCKER** — `lib/dispatch-runtime.ts:9` / `lib/dispatch-runtime.ts:24`: de child-container krijgt `HTTP_PROXY` en `HTTPS_PROXY` op `http://127.0.0.1:18081`, maar `buildContainerArgs` start alleen het child-image met `--network none` en mount de egress-volume read-only. Er wordt geen `startEgressBridge()` in dezelfde netwerknamespace gestart om `127.0.0.1:18081` door te verbinden naar `/run/dispatch/egress.sock`; `bin/run-dispatch-egress-proxy.ts:18` start alleen de Unix-socket-proxy in de aparte egress-container. Daardoor falen normale model/tool HTTP(S)-requests vanuit de child met connection refused, terwijl de huidige integration probe dit mist omdat die rechtstreeks op de Unix-socket test. Dit blokkeert activatie van de managed runtime. ## Verdict REQUEST_CHANGES — de PR introduceert de dispatch-runtime, maar de netwerk-egress wiring voor echte child-processen is niet functioneel.
docs(dispatch): make the model-image egress-bridge boundary watertight (F-D1)
All checks were successful
CI / Compose config (pull_request) Successful in 24s
CI / Docker build (pull_request) Successful in 6s
CI / Build-arg coverage (pull_request) Successful in 18s
0ac04af473
Spell out that the child's HTTP_PROXY/HTTPS_PROXY (127.0.0.1:18081) has no
listener inside the --network none netns until the model image runs
startEgressBridge(); that the broker only mounts the socket and never injects
a bridge; that tests/dispatch-runtime.integration.sh drives the socket
directly and NOT the proxy path; and that this is an activation gate tracked
as T-1869, not a defect in the merged dispatch code.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
janpeter merged commit ac4c106f48 into master 2026-09-21 19:30:00 +02:00
Sign in to join this conversation.
No reviewers
No labels
severity/s3
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
janpeter/scrum4me-docker!82
No description provided.