a4c2a9b90b87bc6ded26f5384c0ce82b908c5d2e
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
675facdb0d |
The beds name images the way a machine would find them
Twenty-eight integration tests each carried their own copy of the same two helpers, which pointed a manifest and the substrate bundle at whatever the lab's registry had assigned. They now share two in the harness, and the difference is the point: ours is rewritten to the ID the machine holds it under, and everything else is left exactly as written so the machine pulls it. **The substrate bundle is where the fiction was most load-bearing.** mesh-host's `examples/substrate-first-node.lock` pins all three of its images at `192.0.2.250:5000/…`, which is the address the lab's registry served from — it was written for a target, and the target was the lab. Two of those are ordinary third-party images and become the digests mesh-catalog's own postgres and lavinmq modules pin, so the substrate's store and broker are literally the images the mesh runs. mesh-control exists in no registry at all and becomes the ID the machine was handed. **The bundle itself should be fixed in mesh-host and this substitution deleted with it.** Beds that wrote a manifest by hand named an image by repository and let the rewrite supply a digest. There is nothing to supply one now, so `onTheMachine` refuses an unpinned reference and hands back the digest the catalogue pins — a bed runs the image the mesh ships, and a bed that drifts from the catalogue is testing a different postgres. Three beds took a third-party image out of the raised list, which no longer contains one: certificates (pebble), objectstore (minio and its client) and provisioner (postgres) now name theirs and pull it. builds and mesh publish into the MESH's own artifact store — the `registry` module's image, on the node, on 5000 — rather than into scenery the lab raised. That is a different claim, and only one of them exists in production. New unit tests cover what a full raise would otherwise be the only way to check: the routes an egress machine gets (that its gateway is still the path to the rest of the scenario, that a range with no path is unreachable rather than leaked to the uplink, that each family gets its own next hop), which machine is handed which of our images, and the `images:` rule that refuses a third-party entry. The "shipped scenarios are valid" test now loads every scenario rather than two of them. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
0d286b7cc8 |
What review found in the lab, fixed
A segment named "uplink" is refused. The lab claims that name for the NAT bridge behind `egress: true`, and a scenario wearing it first would have its egress machines silently attached to an isolated bridge — a declared key doing nothing, which is the fault this repo exists to refuse, in the repo that refuses it. settled() parses inside the try. A truncated status from a struggling machine was the one shape of bad answer that still threw out of the wait, and the likeliest moment for one is exactly the machine the poll is watching. Malformed now counts as "could not ask", like the exec that times out. And a sentence on the uplink's UseDNS saying its inertness is load-bearing: it matters only where systemd-resolved runs, and on a machine whose modules own resolv.conf the uplink must not outvote the resolver a scenario is testing. |
||
|
|
4a343a2652 |
A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test. The forge failed three runs in a row as "status hangs", and it was diagnosed twice as contention — real defects, fixed, and not the cause. The heartbeats told the truth in the end: every exec on anchor crawled from 15s to 105s, because eleven containers plus a database pull were running in a 1GiB machine. Starvation presents as whatever you were doing when the page-outs start, which is why it wore two other bugs' clothes first. So machine size is now the scenario's to declare — memory and cpus per machine, default unchanged. The anchor that carries the whole substrate is bigger than the laptop that joins it, and the comment on the scenario says why in terms of what lands there. `egress: true` gives a machine one extra interface on a lab-supplied NAT network, addressed by DHCP because the one address a scenario has no business choosing is on the host's side of the fence. Declared per machine and off by default: a closed scenario stays the rule (novox/hq ADR 0016), and the exception exists because a first node fetches its images before any mesh can serve them — which is now the tested path (04-ISSUES/029), and a lab that can never reach upstream cannot prove the bootstrap it exists to prove. The uplink route is metric-4096, so it never shadows a route the scenario declared. A detached machine declaring egress is refused, not ignored. And settled() treats a poll that threw as a poll that missed. An exec timeout at minute four of a wait is "could not ask", not a verdict on the machine. |
||
|
|
a6b7d67e19 |
Gateways sharing an address are one gateway
Found by asking what gw-devices and gw-home actually were, in a picture that
finally made them easy to see side by side.
planRouters grouped on the exact address list, so `home` declaring a v4 and a v6
address and `devices` declaring only the v4 became two router containers — both
holding 198.51.100.7 on the same segment. The lab raised it without complaint.
Not theoretical. On the raised instance the transit router resolved that one
address to two different MACs across a cache flush:
198.51.100.7 -> 02:c9:16:70:23:29 (gw0, which HAS the :443 dnat)
198.51.100.7 -> 02:bd:75:0b:b0:75 (gw1, which has none)
So home-server's published port worked or did not depending on which container
answered ARP last — intermittent, and it would have presented as a flaky test
rather than as a broken scenario.
One public address is one box. Checked against the thing this models rather than
argued from the model: a bridged modem, a single gateway holding the public
address, one network behind it, and every port forward landing on one host at
that address. Two routers on one address is not a topology, it is a collision.
Gateways to the same segment sharing any address are now one router and their
address lists union, so a v6 address declared on only one of the segments it
serves is still carried. Where such declarations disagree on nat, forwardable or
mapping_ttl, validate refuses — one box cannot behave two ways.
the-ordinary-shape now raises 7 machines instead of 8, and gw0 holds the public
address on eth0 while serving home on eth1 and devices on eth2.
|
||
|
|
a27d861d3b |
Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a workstation, not asserted: two machines raised and addressed in 14.6s, snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no loss, and the workstation with no route into any of it. The declaration layer implements the model in full — three positions a machine can be in, keyed on forwardability; gateways carrying the address the world sees them as; both address families; multi-homing; MTU; inter-segment policy. It is validated hard because the failures it prevents are silent: a private range on a public segment produces no error, the mesh simply never forms. Public segments are refused unless they use RFC 5737 or RFC 3849 space, and a range wider than the reserved block is refused too. 33 tests, all offline. The runtime implements less than the model, and refuses the difference. A scenario declaring gateways, published ports, policy, inbound deny or place is rejected at raise with every gap named. Raising it would produce a mesh that silently lacks what it declared, which is the fault this lab exists to catch — 04-ISSUES/003, where a firewall key is declared in five manifests and read by no code. Three bugs found by review and by running it, all of one family: The readiness check truthiness-tested incusOk's return. `exec … true` succeeds with EMPTY output, so every machine reported unreachable while incus exec on it worked perfectly. succeeds() now exists so the mistake is not available, and network delete had the same bug — it counted zero segments removed while removing them. list() split instance from machine on the last dash, so a machine called home-server absorbed half the instance id and destroy found nothing. Resources are now found by the metadata they carry, never by name. restore reported success in 0.79s while the machine's agent was still starting, so the next command failed. Both raise and restore now wait for usable and say how long that took — reporting the earlier number is transport reported as effect, which is the fault the lab is being built to find. Two incus behaviours worth recording. Its CLI reads a YAML definition from stdin when stdin is not a terminal, so a spawned command hangs until the timeout kills it and arrives with empty stderr — a failure with no explanation, on a command that works when typed. And it assigns a MAC at runtime without recording it in device config, so MACs are derived and set explicitly, which the guest needs anyway: it names interfaces by bus position, and matching by name configures the wrong one on a multi-homed machine. No build step; Node strips the types. The lifecycle has no unit tests because a fake hypervisor would assert that the fake behaves as expected, which is the shape of test this project exists to stop shipping. |