Commit Graph
3 Commits
Author SHA1 Message Date
jschoubben 5d6e8fbe7a Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 7641bb2059 A one-node mesh, and twelve things that have to be true of it
The common case, and the one that was never tested as a whole. What existed
asked whether four machines converged; it never asked whether ONE machine ends
up holding a mesh.

The order was wrong too. Three machines were enrolled second, into a mesh that
could not yet produce a single module, and that was reported as though something
had been shown. 17-raising-a-mesh is explicit: genesis ends with a mesh that
RUNS, and what remains after the core modules are built is "adding machines".
So the core comes first and machines arrive last — here, not at all, because a
second node is only meaningful once the first is complete.

Three things were missing entirely and nothing complained, because nothing asked:
the mesh never built its own catalogue, never had a store of its own for that
catalogue to use, and never rebuilt its own control plane through the module
path.

And four checks that were absent rather than failing:

  - it can describe itself — status, module list, plan --json, and the
    catalogue's five tools ASKED rather than observed. A container being up was
    being read as the catalogue working, which is the same error as matching a
    container by substring and finding the wrong one.
  - its networking is what the modules asked for — default closed, ssh open,
    declared ports open, .internal names written, module networks present. Left
    out altogether, which is hard to defend given the firewall work this week.
  - a change to a module's source reaches the machine on its own. The capability
    the migration depends on.
  - it comes back after a reboot. Never once tested; the lab had no way to
    restart a machine, because nothing had ever needed one.

Machines are named by role now — anchor, home-server, workstation, laptop — not
after the operator's own nodes, which made test output and real state hard to
tell apart.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:52:24 +02:00
jschoubben babf08b9f8 Raise machines that are somebody, and a bed that hands over nothing
Every machine in a bed is a clone of one base image, so all of them booted with
the same /etc/machine-id. systemd's DHCP client derives its client identifier
from that file and dnsmasq keys leases on the identifier rather than the MAC, so
four machines with four distinct MACs were handed one address and the host kept
one ARP entry for it. Whichever machine last answered an ARP request received
everybody's replies.

This is the fault behind every run lost to "flaky lab DNS": resolution that works
two times in three, pulls that succeed on a retry, and one machine out of four
being fine while the rest have no path at all. It survived an earlier diagnosis
that blamed resolver ordering, because reordering resolvers on a machine that has
just won the ARP race looks exactly like a fix.

Each machine is now given its own machine-id before the uplink lease is asked
for, and a check after addresses are applied refuses to go on if two machines
took the same one — the positive control this never had, since the fault is
invisible where it happens and unrecognisable where it surfaces.

The egress check also now demands five consecutive lookups rather than one. A
single answer is what let a machine resolving one query in three pass and then
die twenty minutes later inside a pull.

And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The
four-machine bed loads thirty-four of the mesh's own images onto its machines
from the workstation because it does not build them, which is a shape no real
installation has and the same fiction the lab removed when it deleted its own
registry. This scenario names no images at all. The machines pull what is public,
the installer builds the control plane, and the mesh builds the rest — including,
last and deliberately, a module on a machine that did not build it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:41:55 +02:00