Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben 5d6e8fbe7a Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 94e617915c The home segment moves off 192.168.1.0/24
It is the commonest home LAN range there is, so on an ordinary workstation the
lab's private segment and the machine's own network are the same addresses. The
scenario routes an egress machine explicitly and marks the rest unreachable, so
nothing leaked — but that guard was carrying the whole weight of a collision
nobody chose, and a guard is a bad place for that.

10.99.1.0/24 is still RFC 1918, so the bed still models a home LAN behind an
access point. It is simply far from what this kind of machine already has:
192.168.1 is the LAN, 172.16-31 and 192.168.16-95 are container bridges, and
10.10/10.42/10.208 are a tunnel, the mesh overlay and the virtualisation daemon.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:00:19 +02:00
jschoubben 675facdb0d The beds name images the way a machine would find them
Twenty-eight integration tests each carried their own copy of the same two helpers,
which pointed a manifest and the substrate bundle at whatever the lab's registry had
assigned. They now share two in the harness, and the difference is the point: ours is
rewritten to the ID the machine holds it under, and everything else is left exactly as
written so the machine pulls it.

**The substrate bundle is where the fiction was most load-bearing.** mesh-host's
`examples/substrate-first-node.lock` pins all three of its images at
`192.0.2.250:5000/…`, which is the address the lab's registry served from — it was
written for a target, and the target was the lab. Two of those are ordinary third-party
images and become the digests mesh-catalog's own postgres and lavinmq modules pin, so
the substrate's store and broker are literally the images the mesh runs. mesh-control
exists in no registry at all and becomes the ID the machine was handed. **The bundle
itself should be fixed in mesh-host and this substitution deleted with it.**

Beds that wrote a manifest by hand named an image by repository and let the rewrite
supply a digest. There is nothing to supply one now, so `onTheMachine` refuses an
unpinned reference and hands back the digest the catalogue pins — a bed runs the image
the mesh ships, and a bed that drifts from the catalogue is testing a different
postgres.

Three beds took a third-party image out of the raised list, which no longer contains
one: certificates (pebble), objectstore (minio and its client) and provisioner
(postgres) now name theirs and pull it. builds and mesh publish into the MESH's own
artifact store — the `registry` module's image, on the node, on 5000 — rather than into
scenery the lab raised. That is a different claim, and only one of them exists in
production.

New unit tests cover what a full raise would otherwise be the only way to check: the
routes an egress machine gets (that its gateway is still the path to the rest of the
scenario, that a range with no path is unreachable rather than leaked to the uplink,
that each family gets its own next hop), which machine is handed which of our images,
and the `images:` rule that refuses a third-party entry. The "shipped scenarios are
valid" test now loads every scenario rather than two of them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:41 +02:00
jschoubben 0d286b7cc8 What review found in the lab, fixed
A segment named "uplink" is refused. The lab claims that name for the
NAT bridge behind `egress: true`, and a scenario wearing it first would
have its egress machines silently attached to an isolated bridge — a
declared key doing nothing, which is the fault this repo exists to
refuse, in the repo that refuses it.

settled() parses inside the try. A truncated status from a struggling
machine was the one shape of bad answer that still threw out of the
wait, and the likeliest moment for one is exactly the machine the poll
is watching. Malformed now counts as "could not ask", like the exec
that times out.

And a sentence on the uplink's UseDNS saying its inertness is
load-bearing: it matters only where systemd-resolved runs, and on a
machine whose modules own resolv.conf the uplink must not outvote the
resolver a scenario is testing.
2026-09-01 21:54:59 +02:00
jschoubben 4a343a2652 A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test.

The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.

So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.

`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.

And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
2026-09-01 21:42:18 +02:00
jschoubben a6b7d67e19 Gateways sharing an address are one gateway
Found by asking what gw-devices and gw-home actually were, in a picture that
finally made them easy to see side by side.

planRouters grouped on the exact address list, so `home` declaring a v4 and a v6
address and `devices` declaring only the v4 became two router containers — both
holding 198.51.100.7 on the same segment. The lab raised it without complaint.

Not theoretical. On the raised instance the transit router resolved that one
address to two different MACs across a cache flush:

    198.51.100.7 -> 02:c9:16:70:23:29   (gw0, which HAS the :443 dnat)
    198.51.100.7 -> 02:bd:75:0b:b0:75   (gw1, which has none)

So home-server's published port worked or did not depending on which container
answered ARP last — intermittent, and it would have presented as a flaky test
rather than as a broken scenario.

One public address is one box. Checked against the thing this models rather than
argued from the model: a bridged modem, a single gateway holding the public
address, one network behind it, and every port forward landing on one host at
that address. Two routers on one address is not a topology, it is a collision.

Gateways to the same segment sharing any address are now one router and their
address lists union, so a v6 address declared on only one of the segments it
serves is still carried. Where such declarations disagree on nat, forwardable or
mapping_ttl, validate refuses — one box cannot behave two ways.

the-ordinary-shape now raises 7 machines instead of 8, and gw0 holds the public
address on eth0 while serving home on eth1 and devices on eth2.
2026-08-24 23:43:23 +02:00
jschoubben a27d861d3b Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.

The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.

The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.

Three bugs found by review and by running it, all of one family:

The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.

list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.

restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.

Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.

No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
2026-08-24 01:12:49 +02:00