Commit Graph
4 Commits
Author SHA1 Message Date
jschoubben c3d5ec1d55 Build each repository from its own ref, and build the provider
lavinmq needs building now, so the step that assigns it builds it first.

And the ref is per repository rather than one value for all of them: a change
under test lives in one repository, and building the others from that branch
would prove it against itself.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:05:28 +02:00
jschoubben f04911a763 Give the module the broker it asks for, rather than a module that asks for nothing
The mesh refused to place amqp-ping: nothing provides amqp. That refusal is
right. The substrate raises a broker, but as a bundle resource — plumbing, not a
module the mesh has a record of — so it offers nothing to anything, and a module
wanting a broker wants one in the graph.

lavinmq is that module and needs no building, its image being upstream, so this
is a register and an assign. The alternative was to pick a module with no
requires, which would have passed by testing less.

Also: the control plane's image has no /tmp to copy a manifest into. Root does.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:56:35 +02:00
jschoubben 4ef0a19053 A manifest the control plane can open, and stop swallowing the failure when it cannot
mesh-control runs in a container, so a manifest pushed to the machine is not a
file it can read; `module add` said so plainly and it was briefly taken for a
missing manifest. It is copied the last step of the way now.

The base's registration was doing this too, and its failure was swallowed by a
bare catch on the reasoning that the module might already be known. The step
passed regardless — a base with nothing to stand on builds whether or not the
mesh holds a record of it — and the fault surfaced one step later, where the
record was needed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:46:11 +02:00
jschoubben babf08b9f8 Raise machines that are somebody, and a bed that hands over nothing
Every machine in a bed is a clone of one base image, so all of them booted with
the same /etc/machine-id. systemd's DHCP client derives its client identifier
from that file and dnsmasq keys leases on the identifier rather than the MAC, so
four machines with four distinct MACs were handed one address and the host kept
one ARP entry for it. Whichever machine last answered an ARP request received
everybody's replies.

This is the fault behind every run lost to "flaky lab DNS": resolution that works
two times in three, pulls that succeed on a retry, and one machine out of four
being fine while the rest have no path at all. It survived an earlier diagnosis
that blamed resolver ordering, because reordering resolvers on a machine that has
just won the ARP race looks exactly like a fix.

Each machine is now given its own machine-id before the uplink lease is asked
for, and a check after addresses are applied refuses to go on if two machines
took the same one — the positive control this never had, since the fault is
invisible where it happens and unrecognisable where it surfaces.

The egress check also now demands five consecutive lookups rather than one. A
single answer is what let a machine resolving one query in three pass and then
die twenty minutes later inside a pull.

And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The
four-machine bed loads thirty-four of the mesh's own images onto its machines
from the workstation because it does not build them, which is a shape no real
installation has and the same fiction the lab removed when it deleted its own
registry. This scenario names no images at all. The machines pull what is public,
the installer builds the control plane, and the mesh builds the rest — including,
last and deliberately, a module on a machine that did not build it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:41:55 +02:00