Files
hq/02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
T
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00

107 lines
6.2 KiB
Markdown

---
topic: building it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
---
# 149. The live mesh is the test bed
## Context
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
built.
What happened instead is that the mesh became the thing under test. It runs on four machines; every
fault worth finding in the last month was found on them, and none was found in a bed:
- a container holding an address that had not existed for five days, on the control node
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
- a machine reading healthy for eleven hours while no module could reach another
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
- one name replacing every container on the hub
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
does not have a mesh that has been running for weeks, with consumers already bound, containers created
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
that does not.
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
several minutes, so they were batched, and a batched test is one whose result arrives after the next
three changes were already written.
## Considered Options
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
## Decision
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
because the state that breaks things is state a bed does not have: containers made against an older
roster, consumers already bound, an adopted machine, a store with weeks of history.
**A change that can only be exercised on the raise path is not verified.** If the only test available
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
"exercised on a fresh mesh" are a statement about coverage, not a pass.
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
better than anything else. What this record removes is the lab as the *default* answer to "is this
change good", and with it 0068's queue, tools and request protocol.
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
not a reason to test it somewhere it cannot break.
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
the ones that were not produced results about code nobody had written down.
## How this is checked
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
on the only path where it works.
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
arriving at the queue design finds out immediately that it was not built and why.
## Consequences
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
hub recreated five times. Both were found in minutes because they were live, and both would have
passed a bed.
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
checks are what stands between a change and the machines, which raises what those suites are worth
and makes a test that cannot fail a genuine defect rather than an untidiness.
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
as a side effect of testing something else. The foundation work
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
is that job, and it is also the proof that the mesh can make another of itself.
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
watches a push and reads the machines, which is what happened anyway.
## References
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
- [ADR 0016](0016-the-lab.md) — the lab, which stands
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh