One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Every machine in a bed is a clone of one base image, so all of them booted with
the same /etc/machine-id. systemd's DHCP client derives its client identifier
from that file and dnsmasq keys leases on the identifier rather than the MAC, so
four machines with four distinct MACs were handed one address and the host kept
one ARP entry for it. Whichever machine last answered an ARP request received
everybody's replies.
This is the fault behind every run lost to "flaky lab DNS": resolution that works
two times in three, pulls that succeed on a retry, and one machine out of four
being fine while the rest have no path at all. It survived an earlier diagnosis
that blamed resolver ordering, because reordering resolvers on a machine that has
just won the ARP race looks exactly like a fix.
Each machine is now given its own machine-id before the uplink lease is asked
for, and a check after addresses are applied refuses to go on if two machines
took the same one — the positive control this never had, since the fault is
invisible where it happens and unrecognisable where it surfaces.
The egress check also now demands five consecutive lookups rather than one. A
single answer is what let a machine resolving one query in three pass and then
die twenty minutes later inside a pull.
And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The
four-machine bed loads thirty-four of the mesh's own images onto its machines
from the workstation because it does not build them, which is a shape no real
installation has and the same fiction the lab removed when it deleted its own
registry. This scenario names no images at all. The machines pull what is public,
the installer builds the control plane, and the mesh builds the rest — including,
last and deliberately, a module on a machine that did not build it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first attempt wrote /etc/resolv.conf. On these images that is a symlink
owned by systemd-resolved, so the file is either reverted or the link is broken
— found by reading a running machine instead of assuming the change had worked.
The real shape shows in resolvectl: the machine has sensible global fallbacks,
and the link carrying the default route has exactly one server, the uplink
gateway. resolved will not reach a global fallback while the link has a server
of its own, so one unanswered packet is one failed lookup. Three runs have died
that way, each long after the egress check passed.
The uplink stays first, so the modelled path is still what is used and still
what the check proves. Verified on a live machine: three servers on the link,
uplink first, resolution intact.
The egress check proves the uplink resolves and reaches the internet, and that
is the right thing to check. What it does not cover is the hours afterwards: a
bed pulls images on four machines at once, the uplink's resolver is one server
under exactly that load, and three runs have now died at a lookup timeout long
after the check passed — the path was never broken, a query just went
unanswered.
The uplink stays first and keeps proving the path. Public resolvers sit behind
it and answer only when it does not, and the retry is tightened so a silent
server costs seconds. A genuinely broken uplink still fails the check, before
any of this applies.
The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and
rewrote every manifest reference — third-party ones included — to point at it. No
production mesh has such a thing. So every bed proved that a machine could fetch an
image from a registry that exists nowhere else, and the bootstrap problems that only
appear when a machine has to fetch for itself went unfound.
What replaces it is the two things that are true in the world:
**Public images come from the public internet.** mesh-lab already created a NAT'd
uplink for exactly this and attached it to any machine declaring `egress`; no scenario
ever declared it. They do now, and third-party references are left exactly as the
catalogue writes them.
**The mesh's own images have no registry and never will.** mesh-control, mesh-builder,
mesh-route-proxy and the per-module runtimes are built from source and exist in no
registry. A machine gets them the way an operator's machine does — they are built here
and loaded onto it — and is then named by the digest of its own image configuration,
which mesh-host now accepts as "an image this machine already holds".
`images:` therefore means only *ours*, and a third-party entry is refused rather than
quietly loaded: otherwise the fiction returns one convenient line at a time. It is
per-machine as well, because "everything, everywhere" was never a description of
anything real — handing whole-mesh-full's union to its two 30GiB workstations would
fill the disk with runtimes nothing on them will start.
**The uplink and the declared gateway would have fought, silently.** A gateway container
and the transit router reach the scenario and nothing else; a default route through
either is a black hole for anything outside, and it beats the uplink's DHCP route on
metric. So a machine with egress states the scenario's ranges explicitly — through the
same gateway or transit it would have defaulted to, so the overlay-across-NAT path is
unchanged — and leaves the default to the uplink. A range with no path inside the
scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an
ordinary private range in fact, and letting it escape would put scenario traffic on
whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested,
because a decision only a full raise could check is one nobody checks.
The registry-reachability check the raise gained earlier is kept, pointed at the real
thing: every machine with egress must resolve a name and reach the internet before the
raise says it finished. Same failure it was written for — a raise that returns, an apply
that dies on its first pull, an instance left a bare shell — now guarding the path that
actually carries.
The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It
was never only for the lab's registry: the mesh has one of its own, the `registry`
module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF