Files
hq/02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
T
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00

6.2 KiB

topic, status, date, deciders, reconstructed, supersedes
topic status date deciders reconstructed supersedes
building it accepted 2026-09-30 jochen false 02-DECISIONS/0068-the-lab-takes-requests.md

149. The live mesh is the test bed

Context

ADR 0068 proposed that the lab accept queued requests — a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so an agent could start a run and come back to it. It has been proposed since 2026-09-12 and nothing was built.

What happened instead is that the mesh became the thing under test. It runs on four machines; every fault worth finding in the last month was found on them, and none was found in a bed:

  • a container holding an address that had not existed for five days, on the control node (issue 135);
  • a machine reading healthy for eleven hours while no module could reach another (issue 145);
  • one name replacing every container on the hub (issue 151);
  • a consumer assertion that is correct on a mesh being raised and fatal on one that is running (issue 156).

The last is the one that settles it. That change was exercised on the raise path — which is what a bed is — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it does not have a mesh that has been running for weeks, with consumers already bound, containers created against an older roster, and an adopted machine carrying a predecessor's configuration. The faults that cost the most were all faults of a mesh that already exists, and a bed is by construction a mesh that does not.

Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and several minutes, so they were batched, and a batched test is one whose result arrives after the next three changes were already written.

Considered Options

1. Build 0068 as proposed. Rejected. It answers a question nobody is asking: the bottleneck was never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in. Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.

2. Leave 0068 proposed. Rejected, and it is why this record exists rather than nothing. A record that contradicts current practice and sits unresolved is worse than either answer: it reads as intent to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.

3. Record that the live mesh is the test bed, and supersede 0068. Chosen.

Decision

A change is verified against the mesh that is running. Not because a bed would be unwelcome, but because the state that breaks things is state a bed does not have: containers made against an older roster, consumers already bound, an adopted machine, a store with weeks of history.

A change that can only be exercised on the raise path is not verified. If the only test available raises a fresh mesh, the record says so, and says which case was therefore not covered. The words "exercised on a fresh mesh" are a statement about coverage, not a pass.

The lab is not retired, and ADR 0016 stands. It remains the place to raise a mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does better than anything else. What this record removes is the lab as the default answer to "is this change good", and with it 0068's queue, tools and request protocol.

Accuracy over a green run. An honest failure on the live mesh beats a pass in a bed that could not have failed — and a change that is risky on the running mesh is a reason to make the change smaller, not a reason to test it somewhere it cannot break.

What 0068 got right is kept as a rule, not a mechanism: a run reads a copy that is not anybody's working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and the ones that were not produced results about code nobody had written down.

How this is checked

  • A record that says a change was verified says on what. Where it was a fresh mesh, it says which case is uncovered. This is the clause that would have caught 156: its change was verified, honestly, on the only path where it works.
  • The lab is not in the path of a merge. No check, playbook or handoff requires a bed to have run.
  • 0068 is unreachable as intent. Its status is superseded and it names this record, so a reader arriving at the queue design finds out immediately that it was not built and why.

Consequences

  • A fault can be introduced on the machines that serve. This is the cost, it is real, and it was paid twice in one evening — a control plane crash-looping for half an hour, and every container on the hub recreated five times. Both were found in minutes because they were live, and both would have passed a bed.
  • There is no pre-merge gate beyond the repositories' own suites. make check and the three hq checks are what stands between a change and the machines, which raises what those suites are worth and makes a test that cannot fail a genuine defect rather than an untidiness.
  • Raising a mesh from bare is now the lab's whole job, and is exercised deliberately rather than as a side effect of testing something else. The foundation work (issue 146) is that job, and it is also the proof that the mesh can make another of itself.
  • An agent cannot hand a run to a queue and come back, which 0068 would have given. In practice it watches a push and reads the machines, which is what happened anyway.

References

  • ADR 0068 — superseded by this
  • ADR 0016 — the lab, which stands
  • issue 156 — correct on the raise path, fatal on a running mesh