Merge pull request 'Issues 307, 308: the lab's beds kept the bus and runtime the mesh had left; a lab unit test's verdict depended on the machine' (#192) from issues/307-308-the-labs-beds-and-a-host-dependent-test into main

This commit was merged in pull request #192.
This commit is contained in:
2026-10-08 08:34:52 +00:00
4 changed files with 201 additions and 0 deletions
@@ -0,0 +1,93 @@
---
status: located
opened: 2026-10-08
located-in: [mesh-lab test/integration (mesh.test.ts, builds.test.ts), mesh-lab scripts/build-module-runtime.sh, mesh-lab replays/cmd/prove, mesh-lab src/lifecycle/raise.ts, mesh-controller examples/objectstore-provisioner]
fixed-by:
amended-design:
---
# 307. The lab's beds still dialled the bus and the runtime the mesh had left
## Symptom
On 2026-10-08 an agent proving the tunnel join (novox/mesh-lab PR #61) on the laptop could not run the
two-node walk, the lab's oldest end-to-end bed. Four things stood in the way, and none of them was
reported by any check before someone tried to run a bed:
1. **The walk's scenario stocks a packet-filter runtime image that cannot be built.**
`scripts/build-module-runtime.sh` copied the tool runner's compiled TypeScript and its `node_modules`
out of mesh-tools. mesh-tools has had no package at its root since the runtime became `node-tools`, a
Go binary that launches every bundle ([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0193](../../02-DECISIONS/0193-every-bundle-the-runtime-serves-is-launched-and-the-runtime-knows-no-language.md)),
so the script failed at its first step.
2. **The walk's builder dialled the old broker.** It was started with the old bus's guest account on
loopback. The mesh moved its bus to NATS ([ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md),
[0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md),
[0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)), and the builder
now reads only a credential the mesh sealed to the machine. The variable it was given is not read at
all. The single-node build bed (`builds.test.ts`) does the same.
3. **Behind a corporate VPN client that routes 10/8 and 172.16/12, nothing can be raised.** The lab
asks the hypervisor to choose the uplink's range itself, and it finds none free.
4. **The replays prover could not run a test in mesh-tools**, because that repository's Go module lives
in a subdirectory (`node-tools/`), and the prover ran `go test` only at a repository's root.
The same day, while trying to run the walk again, two more gaps surfaced: the object store's
provisioner image, which the walk also stocks, cannot be built, because the client image it copies
from is gone upstream; and the laptop could not reach the public container registry, which the lab's
egress check needs (an outside condition, not the lab's fault).
## Why it is a design issue
The integration beds are the only place the parts of the mesh are proven to meet
([ADR 0017](../../02-DECISIONS/0017-a-test-defends-a-decision.md); the lab's README). The lab's own check
runs the type-check and the unit suite, and says plainly that it does not run the beds
([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)
as amended). That is said, not silent. But nothing ties a mesh-wide move to the beds that depend on what
moved. The bus moved, the tool runner changed language, and the beds kept the old shapes. They
type-checked and their unit tests passed. They broke only when a person ran them. The beds go stale
quietly, and each run starts with an archaeology of what changed under them since the last one.
This is [issue 005](../005-pipeline-test-harness-unbuildable/00-report.md)'s shape, one layer up: the
lab's receipt (`last-run`) shows that a bed has not passed lately. It does not show that a bed can no
longer even start, and it does not show which mesh change stopped it.
## Fix
In novox/mesh-lab PR #63, stacked on #61:
- **The runtime image is built on `node-tools`.** The script compiles the module's declared entrypoints
against the SDK, writes beside each one the launcher the mesh's builder writes, stages the SDK,
builds `node-tools` from mesh-tools' Go module, and runs it as the image's entrypoint. The image serves
on the node's runtime credential: the Go runtime refuses to serve a module on that module's own
credential.
- **The walk no longer stocks a packet-filter runtime.** None of the catalogue's nftables containers
names an image (its verbs are a tools bundle the node's runtime launches), so the image was never run.
- **The walk's builder holds the build seat over the bus.** A lab module claims `node-build-agent` with
an own `broker` secret, as the catalogue's build agent does. Assigning it to the anchor issues the
account ([issue 203](../203-a-fresh-assignment-is-pushed-before-its-credential-exists/00-report.md)).
The bus's user list is placed, the push writes the sealed credential, and the builder starts on it.
It starts once the anchor has joined, in the test that builds.
- **The uplink's range may be named** (`MESH_LAB_UPLINK_V4`), documented in the README and unit-tested,
in novox/mesh-lab PR #61.
- **The prover runs a replay inside a module subdirectory**, through the register's `Module` field, in
novox/mesh-lab PR #62. Nothing separate is needed for mesh-tools: its replays name `node-tools`.
Still open:
- `builds.test.ts` still dials the old broker. Its bed enrols no machine, so the mesh has nowhere to
deliver a builder's credential. It needs the walk's enrolment, or it should be retired into the walk.
- Beds that run a per-module runtime container on the module's own credential (the tools beds for the
wiki and the code forge, and the audit bed) run into the runtime's refusal. They need the node's
runtime credential instead.
- The object store provisioner's image needs a client source that still exists.
- The walk has not been run end to end since these changes: on 2026-10-08 the laptop could not reach the
public container registry.
## How it is checked
**Partly, today:** the lab's `merge-check.sh` type-checks every bed. That would have caught a renamed
import, but not a bed dialling a bus that no longer runs. `last-run` judges a receipt that is old, failed
or against moved commits, and says nothing about a bed that cannot start.
**Not yet checked, and that is the open part of this issue:** that every bed can still start against the
mesh as it is. Until a check exists, this record names the rule as unenforced.
@@ -0,0 +1,38 @@
# 307 — diagnosis
## 2026-10-08
**The runtime image.** `scripts/build-module-runtime.sh` ran `npm run build` in mesh-tools and copied its
`dist` and `node_modules` into the image, then started `node dist/main.js` with the module's entrypoints
in `MESH_TOOL_MODULES`. That was the TypeScript runtime, which imported each bundle. mesh-tools now holds
two modules: the toolchain images at its root and `node-tools/`, whose runtime is the Go program
`cmd/node-tools`, built by the mesh as that module's bundle. It launches each bundle as a child and speaks
MCP over stdio to it, and it takes `MESH_TOOL_MODULES` only as `<module>=<entrypoint>`. A TypeScript bundle
becomes launchable because the builder writes `<entry>.serve.mjs` beside each entrypoint. The launcher
imports the entrypoint and calls the SDK's stdio loop, and the SDK has no runtime dependencies. The
script now does those steps by hand. Verified: the image for the packet filter builds; its launcher
answers `initialize` and `tools/list`; `node-tools` in it, started against a real bus with a plain URL,
launches the bundle and serves nothing until the mesh issues a membership, which is correct.
**Why the walk stocked that image at all.** The tunnel branch passed the held images to
`deriveTheFilterOn`, which turns the catalogue's manifest into the lab's form: build section gone, each
container's artifact replaced by a stocked image. The catalogue's nftables manifest has no container,
so no image is referenced. Removing the build section is what the bed needed; the stocked image was
not.
**The builder.** mesh-controller's `cmd/mesh-builder` reads `MESH_BROKER_FILE` and nothing else for the
bus. Without it, it says it has no credential. The walk set the old broker's variable, which the builder
no longer reads. A builder's account is a module's account (`builder issue` and `module issue` both mint
through the inventory), composed into the bus's user list and sealed to the machine as the module's
`broker` own secret. Assigning a module that declares that secret now issues it in the same act (issue
203). The catalogue's build agent is that module, but it runs the builder as a container from an
artifact the mesh builds, and it needs the artifact store and package registry bound. The walk has
neither before it can build, so it declares a lab module with the same claim and the same secret. In
this bed the bus comes from the bundle, so its user list is placed by hand, the step the walk already
takes at enrolment.
**The uplink and the prover** were each fixed by the agent that met them (mesh-lab #61 and #62), and are
listed here only so the lab's state on that day is in one place.
**Ruled out.** That the walk failed because of the tunnel changes. The builder and runtime faults are in
code that #61 does not touch, and they reproduce on main.
@@ -0,0 +1,53 @@
---
status: located
opened: 2026-10-08
located-in: [mesh-lab src/incus/client.ts, mesh-lab test/enumerate.test.ts]
fixed-by:
amended-design:
---
# 308. A lab unit test passed or failed by the machine it ran on
## Symptom
On 2026-10-08, `npm test` in mesh-lab on main, on the laptop, failed two tests:
- `listing instances fails rather than reporting none`
- `listing networks fails rather than reporting none`
The same commit passed `mesh/repo-check`, whose unit suite ran them green on the build seat. One test
file, one commit, two verdicts. The merge check was the green one, so the red verdict on the machine
that actually runs the lab looked like a local problem.
## Why it is a design issue
The lab's unit suite is offline by its own rules: it checks the declaration layer and the lab's reading
of what the hypervisor says, and the hypervisor is never mocked (the lab's README). A test in that suite
must therefore give the same verdict on every machine. If it does not, either the check that lets a
change merge proves nothing about the workstation, or the workstation fails a commit the check passed.
Here both were true at once.
## Cause
The test set `MESH_LAB_INCUS=false` in its body, meaning to make every incus call fail. The client reads
that variable once, when its module loads. An ES module's imports are evaluated before its body, so the
client had already read the variable, found it unset, and used the real `incus`. Where no incus is
installed (the build seat's container) the spawn fails and the tests pass. Where incus is installed and
reachable (the laptop) the listing succeeds and the tests fail. Confirmed: with the variable set before
the process started, main passes both on the laptop.
## Fix
In novox/mesh-lab PR #63: the client takes its command by injection (`useIncusCommand`), and the test
names the failing program through it. This test has nothing to do with the hypervisor; it is about how
the lab reads a failure, so injecting a failing command is not mocking the hypervisor.
## How it is checked
**The check:** the unit suite, in `mesh/repo-check` and by hand. The two tests now fail the listing
through the injected command, so they pass on every machine, with or without incus. A machine where the
injection broke would ask the real incus again and fail on the workstation.
**Not checked:** that no other unit test reads the environment after its imports. Nothing enforces that
yet. A run of the unit suite with incus present and one without would catch any such test. The lab's
check runs only the second.
@@ -0,0 +1,17 @@
# 308 — diagnosis
## 2026-10-08
- On the laptop, main: `npm test` gave 158 pass and 2 fail, the two enumeration tests. Both failed with
"a failed `incus list` came back as an empty list". The call had succeeded, so there was nothing to
reject.
- The file's first statement sets `MESH_LAB_INCUS`, and its imports come after that statement in the
source. The imports are hoisted, though: they are evaluated before the body, and `src/incus/client.ts`
reads the variable into a module-level constant when it loads.
- The same file on main, run with `MESH_LAB_INCUS=false` in the environment before node started: 3 pass.
So the cause is the order of evaluation, not the client's handling of a failure.
- On the build seat the container holds no incus binary, so the spawn fails, `incus` rejects, and the
tests pass for a reason they were not written for.
**Ruled out.** That the client's `enumerate` swallowed a failure: it rejects on a non-zero exit and on a
spawn error, and `false` makes both tests pass once it is what actually runs.