ADR 0067's own acceptance check said the lab must raise its anchor by running the
program a bare machine runs. It did not: whole-mesh-full applied the substrate bundle
by hand and then looped enrolment over all four machines as one continuous operation.
That gets the order right by accident and models the wrong shape — and an install
procedure that exists only as a test fixture is exercised by whoever writes tests and
never by whoever installs, which is why every bootstrap fault this year was found late.
Two acts now, and the first gates the second.
GENESIS is novox running mesh-bootstrap: the installer is built from source before
the raise (make bootstrap, carrying the control-plane image built in the same run),
placed beside the host binary, given the two manifests it reads, and run. The bed
then asserts a WORKING MESH OF ONE — the control plane answers, the registry replies
on /v2/, the container called mesh-control is running from a registry-pinned digest
rather than an image id, the registry agrees it serves it, temp-mesh-control is gone,
and the mesh has heard from its node. The image-id check is ADR 0067's "the pivot
completed" verbatim: if it is still an id, nothing was published and this mesh can
never roll out its own upgrades.
JOINING is ace, shanks and g14: host binary, token, enrol, run. novox is NOT enrolled
again — the installer already did it, and a second identity is one the mesh does not
know.
If genesis stops, the bed prints which of the installer's ten steps it stopped at and
goes no further. A second machine joining a mesh that is not ready is a different
failure, and running it would bury this one underneath it.
The anchor is no longer handed mesh-control:development. Its absence is the point: the
installer carries that image inside itself, and handing it over as well would make the
load say "already held" and leave the carrying untested — the same class of fiction the
lab's own registry used to hide. A unit test asserts the scenario keeps it out.
The registry is reached at 127.0.0.1:5000, which is a finding rather than a shortcut: a
runtime refuses a plain-HTTP registry at any address but a loopback one, so the digest
the control-plane module is pinned to is one only the anchor can pull. Enough here,
because only the anchor runs a control plane. Written down in the bed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Every scenario that places a container runtime gives each of its machines
`egress: true` — a node that runs modules pulls images from the internet, which is what
a node does. The underlay-only scenarios (bootstrap-single, behind-nat,
segmented-and-unforwardable, the-ordinary-shape, two-on-a-segment) stay sealed on
purpose: an extra NIC would change the very reachability they are asserting about.
`images:` keeps only the mesh's own — 55 third-party entries leave whole-mesh-full
alone, and the machine fetches them itself by the digest its module.json already pins.
The whole-mesh beds also say per machine which of ours they get: novox the substrate
control plane and its own fifteen runtimes, ace its twenty-four, the two workstations
one each. That is not a lab economy. An operator's workstation holds the images its own
modules need, and giving these two the union would put some thirty gigabytes onto a
thirty-gigabyte disk.
bootstrap-with-registry.yml is deleted. It existed only to demonstrate the lab's
registry, nothing referenced it, and there is nothing left for it to demonstrate.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.
Three decisions, each doing work.
**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.
**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.
**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.
Two things this found in itself while being written, both the same shape
as what it exists to catch:
A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.
And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.
Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
Node block-buffers stdout to a file, so a run log can sit unchanged for
minutes while the run is fine. Read that way twice today — the second
time straight after fixing a real stall, which is the worst version of
it, because a buffering artifact then reads as the fix having failed.
The machines are the source of truth and answer immediately. Written
down with the two commands that settle it.
Pointed at the repository root, the binary it writes is 12 MB of tracked
artifact. Said in the example rather than left to be discovered by a
On branch initialization
Your branch is ahead of 'origin/initialization' by 43 commits.
(use "git push" to publish your local commits)
Changes to be committed:
(use "git restore --staged <file>..." to unstage)
modified: README.md that looks wrong.
Reconstructed from the source twice now, which is 04-ISSUES/005 in its
own README: a test whose artifact was not pointed at skips rather than
fails, so an unset variable is a green run that proved nothing. The
first attempt today reported "skipped 24" and left a receipt claiming
zero of everything — working exactly as designed, and indistinguishable
at a glance from a suite that had nothing to do.
Also records the two things that cost time either side of it: `check`
says which variables are missing before a long run rather than skipping
quietly, and a heredoc into `newgrp` runs the suite as a child of a
shell that immediately exits, so it needs `setsid nohup … &` or it dies
with the shell that launched it.
The danger is not that the suite breaks. It is that nobody notices it
stopped running (novox/hq 04-ISSUES/005). The harness this replaces had
not built for two and a half months and nothing said so — and this suite
needs a hypervisor, so it inherits exactly that: it runs when somebody
remembers, and remembering is not a mechanism.
So running, recording, and rebuilding are one act:
- the host binary, control-plane image and builder are rebuilt from
source first. The last two both parse manifests; building one and not
the other left a binary eleven hours old refusing a field the mesh had
just renamed, found by a full run.
- a receipt lands in XDG state — outside git, because the question is
whether *this machine* has run it, and a receipt in git would be a
claim about everybody's machine made by whoever committed last.
- `last-run` judges it and exits non-zero when it no longer counts.
Three faults found by running the thing rather than reading it, each now
held by a test confirmed to fail without it:
- counted() passed every test while parsing nothing. The runner colours
its summary even into a pipe; the fixtures were clean text that had
been imagined rather than captured. A fixture that agrees with the
mistake proves the mistake.
- a receipt for `suite test/lastrun.test.ts` was indistinguishable from
one for the real thing — 005's own symptom, rebuilt inside its remedy.
The receipt now records what ran.
- a tree with uncommitted work reported the bare commit, claiming
coverage of code nobody can check out. Nothing else could tell: the
hash is identical either way.
Proven on real machines: 22/22, against all three repositories.
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
The lab raised an underlay and put nothing on it: correct, and useless, because
the thing it exists to test did not exist. Tier 0 now does, so `place: [host]`
works and a raised scenario finally contains something.
The refusal narrows rather than disappearing. A scenario placing a host and a
substrate is told which half is missing, by name — not that `place:` is
unsupported when half of it now works.
Placement reads back rather than assuming. A file arriving is not a host
working, so the binary is run before it is trusted to answer questions, and what
it reports is read from the machine (ADR 0035). The binary comes from an
explicit path, because the declaration design leaves where artifacts come from
open and a search would harden into the answer by accident.
The integration test that matters is the one asserting the host reports the
MACHINE and not the workstation that placed it. A raised VM and this workstation
differ in every capability — root versus uid 1000, a clean init versus a
degraded one, no docker versus docker, no wireguard versus wg0 — so a host
reporting the wrong machine is obvious here and invisible anywhere else.
And the placed host independently confirms ADR 0031: overlay absent on a freshly
raised machine. The underlay suite already asserted that by looking for
wireguard interfaces; this is a second witness rather than the same check twice.
Two tests failed the moment placement worked, which is what they were for. They
defended "there is nothing to place yet" while that was true; the decision
changed, so they change with it rather than being deleted.
Gate: 75 unit, 20 integration.
`mesh-lab diagram` renders a scenario as draw.io, from either source, through
one layout — so a difference between what was asked for and what exists is a
difference you can see.
The shape says what a resource is and is fixed per kind. The badges say what is
true about that particular one and come entirely from metadata: translation,
forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The
interesting properties of a network are exactly the ones with no visual
consequence — a translated address looks identical to an untranslated one.
For the live picture to be a record rather than a restatement, raise now writes
down what it applied: a segment's kind, ranges and MTU on the link; a gateway's
translation, forwardability and expiry on the gateway; inbound: deny on the
machine. Every behavioural tag is written AFTER the thing works, never at
creation — a failed raise leaves wreckage standing on purpose, and a picture of
that wreckage must not badge translation the router never got.
The pairing earned itself immediately: drawn side by side, every virtual machine
held no addresses. A container's interface carries the device's name and a VM
names its own, so joining them by name silently dropped one whole class of
machine. Fixed by joining on MAC.
Also brings tests under the typecheck gate, which caught integration timeouts
being passed as a 4th argument and therefore ignored entirely.
Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of
it the half that touches the hypervisor, and no gate. The verification I had
done was real — pings across NAT, TTL counts, ruleset comparisons — and none
of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature.
Ten integration tests against a real hypervisor, each named for what it
defends. ADR 0031: a raised machine carries no overlay, no wireguard, no
mesh config — a scenario that pre-built peering would certify its own work.
ADR 0032: exec is the only way in. ADR 0033: routers are containers while
machines are virtual machines. And the design's claims: raise waits for
usable, snapshots are whole-scenario, NAT hides a private address,
published reaches the machine at the gateway's address.
Mocking the hypervisor is forbidden, so they skip with a reason on a
machine that cannot raise scenarios rather than passing green having
checked nothing.
The suite earned itself on its first run. It found that a snapshot of a
running machine could miss a file written seconds earlier — not stale,
absent — because the write was still in the guest's page cache. That is
exactly the question the lifecycle design listed as open: does a scenario
snapshot need the machines stopped? It does not, but it does need them
flushed. snapshot now syncs every machine before capturing, and the design
records the answer.
The fix buys write-durability, not application-consistency: anything
mid-transaction is still captured mid-transaction, and that is now stated
rather than assumed.
npm run check is the gate — typecheck, 40 unit tests, 10 integration tests.
The full topology now raises: four machines, three routers, a transit
router, six segments, in 35 seconds. Everything the declaration model can
express except `place`, which is refused because the node host it would
place does not exist yet.
Transit was a real gap, not a bug. The design says public networks are
unrelated and routed to each other, never bridged — and I built the
segments and never built the thing that routes between them, so three
public networks were islands and nothing crossed. A transit router now
holds an interface on every public segment, forwarding and no translation:
the closest thing the lab has to the internet, deliberately dumb.
Proven rather than asserted, by ping TTL across the raised topology:
within one segment ttl=64 no hops
across two unrelated public networks ttl=62 gateway + transit
multicast between public networks 0 replies
A flat internet would have shown ttl=64 and answered multicast — which
would let a node discover a peer it could never reach in production, and
report success. That is the fault the as-is layer records the mesh already
hitting with multicast name resolution.
inbound: deny is implemented as a host firewall on the machine, read back
after applying. A declared refusal that silently did not load leaves the
machine wide open, which looks exactly like a machine that is working.
Established and related traffic is accepted, so a defended machine can
still dial out rather than being a disconnected one.
Verified by running, all of it:
home -> devices (policy allow) reachable
devices -> home (policy deny) blocked
behind unforwardable NAT -> out reachable
in -> behind unforwardable NAT unreachable
inbound: deny, dialling out reachable
reaching a machine that denies inbound refused
The two routers differ exactly as declared: the forwardable one carries the
policy rule and no inbound drop, the unforwardable one carries `ct state
new drop` and no DNAT.
A gateway is the one implicit machine in a declaration — a scenario says a
segment sits behind one and never names the thing that serves it. This
materialises it.
A router is a container, not a virtual machine, because it is scenery
rather than something under test (hq ADR 0033). Verified before building
that a plain unprivileged container can do all of it: ip_forward and ipv6
forwarding settable, nftables masquerade accepted, and the conntrack
timeouts mapping_ttl depends on both writable. No privileged mode.
Verified by running, on a machine behind a household gateway reached from
one on a routable address:
home-server -> anchor 0% loss, through masquerade
anchor -> 192.168.1.135 (private, direct) unreachable
anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200
The last line is the published-but-behind-NAT case research 004 says only
exists in production. It is now a 32-second scenario on a workstation.
Segments sharing a gateway declaration share ONE router — that is what a
VLAN-capable router is, and two routers sharing an external address would
not work anyway.
mapping_ttl is read back after setting rather than assumed. Those sysctls
are not on every kernel, and a scenario that declared an expiring mapping
and silently got a permanent one would be exactly the fault being built
against.
Four bugs found by running it, three of them the same fault — a failure
made invisible.
The router had no route to a package repository, by design, so installing
nftables at raise time could not work. The image is now built once with
temporary connectivity and cached; every scenario after that needs no
network. That failure was hidden behind `|| true`, which is why it took a
raise to find.
The builder then failed on DNS: exec works before a container has an
address, and I had treated usable as ready. It now waits for the thing
actually needed.
The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time
networking service flushed the static address the scenario set — on eth0
only, so the outside interface came up bare while inside ones were fine.
The image build now neutralises it: a router reconfiguring itself from an
image default is the lab overriding the declaration. `ip addr add … || true`
had hidden this too, and is now `ip addr replace` with no swallow.
And routers were orphaned by destroy, holding their networks open so
destroy reported removing zero segments. They now carry the same machine
tag as everything else, so one query finds an instance's resources.
A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.
The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.
The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.
Three bugs found by review and by running it, all of one family:
The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.
list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.
restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.
Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.
No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
The node host takes over a machine's packages, services and network, so it
cannot be developed against a machine anyone needs. The place to develop it
has to exist before it does — which makes this phase 0, ahead of every tier
it will later test.
Two scenario classes, per ADR 0029. Bootstrap is virtual machines, the host
and a pinned bundle, with the verdict coming from what the host reports
about the state it reconciled. Full is a complete mesh with a pipeline.
Bootstrap is a strict subset, so the full scenario is reached by putting
more inside the machines rather than by building a second thing.
Nothing here requires a forge, a coordinator or a pipeline to be useful.
Decisions live in novox/hq. This repository carries implementation.