b0eddd28d8899b2d580f1eb9a9a4e536767230e1
220
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c98ee7a82d |
A machine that fell behind catches up without being named
Broken with a package that does not exist, so the failure is real and fixable. The mesh reports it failed; the resources that could be applied were, because one broken thing no longer blocks the rest; `push --behind` names that machine and not the one that is fine; the module is corrected; and the machine recovers with nobody naming it. And with nothing behind, it says so rather than doing nothing quietly. |
||
|
|
e0793c17f7 |
A declaration waits, and unassigning takes away exactly what it should
Two properties the design claims and neither had been run. A push to a machine that is switched off must not be lost — a machine is disconnected as an ordinary situation, not an exception. The queue is durable and the message persistent, which ought to be enough, but a lost declaration is silent and "ought to be" is not a property. It waits: the machine's host is stopped, the push happens, nothing changes on the machine, and when it listens again it applies what it missed with no second push and nobody saying anything. Getting there found a real fault, now fixed in mesh-host and recorded as 04-ISSUES/011: the machine stopped at the first failing resource, so one broken module blocked every module after it for ever. The evidence was the broker's queues being EMPTY — the declaration had been delivered and read. And removal: two modules assigned, one unassigned, and the machine loses exactly that one's file while keeping the other's — and keeps the store, broker and control plane it raised from its own bundle, which the mesh never declared and must never remove. Two of my own traps recorded in the test, because both cost real time: `pkill -f` matches the shell running it, which kills the connection carrying the command and hangs the caller for ever; and a test that depends on state another test left behind fails for a reason that has nothing to do with what it claims. |
||
|
|
230074665f |
A machine that cannot do what it was told, and the mesh saying so
The status path was demonstrated with inserted rows, which proves the query and not the path. This sends a real machine something it will genuinely fail at — a package that does not exist — and asks the mesh afterwards. A failure of the ordinary kind: the host tries, the package manager says no, some of the declaration is applied and some is not. That is the situation `status` exists to distinguish from a machine that refused everything, and the test asserts the distinction survives the whole way: the machine is listed as failed rather than refused, the failing resource is named in the host's own words, and the machine that did as it was told is not implicated. |
||
|
|
e0127df7ce |
A machine in the mesh builds a module, and the catalogue records it
The chain this closes: a repository exists, the mesh asks for it, a build machine takes the work, publishes what it made, and the catalogue then says what the module is, which commit it came from, and — after the source moves — that it is behind. Three assertions, against a real broker and registry, because what is under test is four processes agreeing over a wire: - the mesh asks, a machine builds, and the artifact is really in the registry at the digest the manifest names - a build that cannot succeed says why and records nothing. A failure that is silent is indistinguishable from a builder that is not running - the source moving makes the catalogue say "behind", and rebuilding catches it up git is now in the base image, with the same reasoning as docker and wireguard-tools: a machine that builds modules clones them, and a sealed scenario cannot install anything. Read back from `git --version` rather than from the package manager — an installed package is not a capability, and a build machine whose clone fails does so three minutes into a scenario with the failure reported as a build problem rather than a lab one. |
||
|
|
99b3444b19 |
Two machines, one mesh, a credential neither end had to be told twice
Everything before this proved a part. This proves the parts meet, which the project keeps saying cannot be checked any other way. A bare machine applies the substrate bundle and becomes a mesh — store, schemas, broker with a certificate it generated itself, control plane serving. Both machines then join it with nothing but a token. A database is declared on one and an application on the other, and after a push: - both ends hold the SAME password, or nothing could authenticate - it is mode 0600 on the machine that uses it - it appears in neither machine's stored declaration, neither machine's reported state, nor the control plane's database — so it was not readable by the broker that carried it or the mesh that sent it - the consumer is also told where its database is, by a name the mesh wrote into that machine's hosts file The bundle's image references are rewritten to the ones this scenario's registry serves. A digest belongs to whatever registry serves it, so a committed bundle names a registry that is not this one — rewriting is what makes it applicable rather than a placeholder to tidy away. Two faults found getting here, both fixed in mesh-host: `apply` could not read a file the bundle could, and the token did not say what the mesh calls the machine. |
||
|
|
e89379fef8 |
Prove a mesh credential becomes a login, against a real database
The mesh generates a password, seals it to the machine that must accept it, and discards the plaintext — so it cannot tell PostgreSQL to start accepting it. Something on that machine reads what the host wrote and makes it true. Everything up to that step is proven elsewhere; this is where a password either becomes a login or does not. A scenario with one machine and a database, and six assertions: the password works, running again reaches the same state and says nothing, rotation makes the new one work and the old one stop, a departed consumer loses its login, a role nobody here made is left alone, and a manifest naming a credential that was never written is refused rather than creating a login with no password. Each was confirmed to fail — and only it to fail — with the behaviour removed from the provisioner: only-creates breaks rotation, no-revoke breaks revocation, revoking everything breaks the bystander role, and ignoring a missing credential breaks the refusal. Two faults in the test itself, both worth recording: - it checked logins from inside the database's own container over 127.0.0.1, which PostgreSQL's default pg_hba trusts. No password was ever verified. Demonstrated directly: over loopback a deliberately wrong password still returns a row. Only the rotation assertion noticed, because it is the one that requires a password to STOP working — which is an argument for writing that assertion every time. - the fix then read .NetworkSettings.IPAddress, which docker 29 no longer populates. It templates to empty, psql falls back to a unix socket that is not there, and every login looks impossible rather than misconfigured. |
||
|
|
88cf89194a |
The lab said nothing was running while two machines were
'incus list' failed because this shell had no permission to reach the daemon, incusOk returned null, and the caller wrote ?? "[]". So 'mesh-lab list' printed 'no scenario instances standing' -- confidently, about a question it had never managed to ask. The comment on incusOk warns about exactly this, in those words: absence and success made indistinguishable. Three of its own callers then did it. Two listings and the live diagram, which would have drawn an empty scenario rather than fail -- a picture that is confidently wrong, which is worse than none. Anything enumerating what exists now goes through enumerate() and throws. incusOk stays right where failure genuinely means no, like instanceExists, and there is a test holding that line so this does not get over-corrected until nothing can be asked at all. Worth noting 'mesh-lab check' already diagnoses this precise cause, down to 'a session that predates it cannot see it'. The diagnosis existed; the listing just never asked for it. |
||
|
|
4097ff92c1 |
Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated record holding their reasoning -- the four lab records are 0016, a test defends a decision is 0017. |
||
|
|
37c6a0ba21 |
A registry inside the scenario: the mechanism, verified
Issue 009's resolution, proven manually end to end before any of it was written. A sealed machine pulled an image BY DIGEST from a registry on its own segment and ran it; then the host applied all four shapes -- package, service with boot, container from that digest, and an action inside it -- idempotently. That is the first time the container shape has worked anywhere but a workstation, and it was the shape blocking the whole substrate bootstrap. The registry's digests are its own, not Docker Hub's, and that is correct rather than a compromise: ADR 0046 requires a reference that is exact and cannot move, and a digest this registry assigned is both. It is also not a lab workaround -- 0048 names an OCI registry as substrate and 0046 says a first node fetches "upstream, wherever the image ordinarily lives". This IS that upstream, scenery in the same sense the transit router is the internet. The base image now trusts the RFC 5737 and RFC 3849 documentation ranges as plain-HTTP registries. Scoped to those rather than an address because they never route on the real internet, so it cannot make a real machine trust a real registry whatever it is copied onto. Three faults found while verifying, two of them mine: My probe script picked an interface with `ls /sys/class/net | head -1`, which returns docker0 once a runtime exists -- so it addressed the wrong interface and then, because that address overlapped the segment, broke routing on the machine entirely. The lab itself is immune: it matches by MAC, for a related reason it already recorded (bus-position naming on multi-homed machines). And a test that proved nothing: I asserted `sha256:tooshort` is rejected, but its letters fall outside a-f, so it failed the character class rather than the length check. Replaced with hex of the wrong length, after which removing the length check bites. |
||
|
|
d6eef25590 |
The lab can give a sealed machine a container runtime
ADR 0046's open consequence: "the lab needs a way to place images, and the
machine it places them into needs a container runtime, which a sealed scenario
cannot install either."
The runtime half is done, and it is research 012's reframing applied literally
-- fetch at build time on a machine with a network, apply on a target that
needs nothing. `mesh-lab base build` launches a machine WITH a network,
installs a runtime, verifies it by asking the runtime rather than the package
manager, and publishes the result. Measured: ~30s to install, ~60s to publish,
~700MiB, paid once per lab rather than per scenario.
A scenario that places `runtime` or an image is then raised from that base
image, chosen rather than declared -- a scenario says what it needs, not which
image provides it. If the base does not exist it says so and how to build it.
Verified in a genuinely sealed machine (no route out, confirmed by ping):
package, service including the new `boot: enabled`, and action all applied,
were idempotent on a second run, and read back correctly. Those three had never
run anywhere but a workstation.
The image half is NOT done, and testing found why: a digest-pinned image cannot
be placed from an archive. `docker save alpine@sha256:...` produces an archive
with no repo tag, because a repo digest only exists for an image a registry
served -- so it loads dangling and a container declaring that digest reaches
for a registry the machine cannot see.
That collides with ADR 0046, which has the host REFUSE an unpinned image. Tag
refused by the host, digest unusable in the lab: there is currently no
declaration the lab can raise that exercises the container shape at all. Filed
as 04-ISSUES/009, whose resolution is a registry inside the scenario -- which is
what the real mesh does rather than a workaround for the lab.
Also fixed a weak check of my own, which is the same fault in miniature: the
load was tested with `includes("Loaded image")`, a prefix of both `Loaded
image:` and `Loaded image ID:`. So an unusable dangling load reported success
and the failure surfaced later as a container that would not start.
|
||
|
|
16c13807a9 |
place: the host — the lab acquires a consumer
The lab raised an underlay and put nothing on it: correct, and useless, because the thing it exists to test did not exist. Tier 0 now does, so `place: [host]` works and a raised scenario finally contains something. The refusal narrows rather than disappearing. A scenario placing a host and a substrate is told which half is missing, by name — not that `place:` is unsupported when half of it now works. Placement reads back rather than assuming. A file arriving is not a host working, so the binary is run before it is trusted to answer questions, and what it reports is read from the machine (ADR 0035). The binary comes from an explicit path, because the declaration design leaves where artifacts come from open and a search would harden into the answer by accident. The integration test that matters is the one asserting the host reports the MACHINE and not the workstation that placed it. A raised VM and this workstation differ in every capability — root versus uid 1000, a clean init versus a degraded one, no docker versus docker, no wireguard versus wg0 — so a host reporting the wrong machine is obvious here and invisible anywhere else. And the placed host independently confirms ADR 0031: overlay absent on a freshly raised machine. The underlay suite already asserted that by looking for wireguard interfaces; this is a second witness rather than the same check twice. Two tests failed the moment placement worked, which is what they were for. They defended "there is nothing to place yet" while that was true; the decision changed, so they change with it rather than being deleted. Gate: 75 unit, 20 integration. |
||
|
|
b015068921 |
Step 2: raise a second scenario, and share the invariants
The suite next door raises one scenario and asks deep questions of it. This one asks shallow questions of every scenario — the half that was missing, since both faults found by hand lived in scenarios nothing ever built. Adds bootstrap-single, the cheapest, and the loop that lets the list grow. Also adds the second universal invariant: every address a scenario declared is one the machine actually holds. A machine that came up bare looks identical to one that came up correctly until something asks it. Verified to bite rather than assumed: against a live instance, the real declaration passes and a declaration claiming an address nothing holds fails with 'anchor declared 192.0.2.99 on hosting but holds 192.0.2.10'. Integration now runs with --test-concurrency=1. Two files raise real instances, node --test runs files in parallel by default, and two concurrent runs of this suite already produced a whole-suite failure once — every test red, from resource contention rather than from any fault in the code. Gate: 45.7s -> 60.2s. |
||
|
|
715f367147 |
Step 1: an invariant that holds of any raised scenario
The address collision was found by eye. This is the mechanical form of it: no two machines hold one address on one segment. Pure over already-collected facts, so the logic is tested without a hypervisor — including the cases that would make it useless if got wrong: the same address on DIFFERENT segments is normal and must not be reported, and one machine holding an address twice is not two machines. Asserted against whatever the integration suite has standing, read from the hypervisor rather than from the declaration. The declaration is what was accepted, and it was accepted. |
||
|
|
a6b7d67e19 |
Gateways sharing an address are one gateway
Found by asking what gw-devices and gw-home actually were, in a picture that
finally made them easy to see side by side.
planRouters grouped on the exact address list, so `home` declaring a v4 and a v6
address and `devices` declaring only the v4 became two router containers — both
holding 198.51.100.7 on the same segment. The lab raised it without complaint.
Not theoretical. On the raised instance the transit router resolved that one
address to two different MACs across a cache flush:
198.51.100.7 -> 02:c9:16:70:23:29 (gw0, which HAS the :443 dnat)
198.51.100.7 -> 02:bd:75:0b:b0:75 (gw1, which has none)
So home-server's published port worked or did not depending on which container
answered ARP last — intermittent, and it would have presented as a flaky test
rather than as a broken scenario.
One public address is one box. Checked against the thing this models rather than
argued from the model: a bridged modem, a single gateway holding the public
address, one network behind it, and every port forward landing on one host at
that address. Two routers on one address is not a topology, it is a collision.
Gateways to the same segment sharing any address are now one router and their
address lists union, so a v6 address declared on only one of the segments it
serves is still carried. Where such declarations disagree on nat, forwardable or
mapping_ttl, validate refuses — one box cannot behave two ways.
the-ordinary-shape now raises 7 machines instead of 8, and gw0 holds the public
address on eth0 while serving home on eth1 and devices on eth2.
|
||
|
|
54d417fdb2 |
Group each public network with what is behind it
The drawings were confusing, and looking at them showed why: a single stack ordered by depth put a private network far from the public one it sits behind, so a gateway's link to the outside ran the full height of the picture through three networks it had nothing to do with — and two such links overlapped, so they read as one wire. Now each public network is followed by everything behind it, depth first. Every gateway is adjacent to the network it serves, every link is a short stub, and "behind" is shown by INDENTATION rather than by a line to follow. Gaps are sized to what they hold, so a gap with no gateway in it takes no room. Transit is not on a boundary — it reaches every public network at once — so it is stated once at the top instead of drawing a line to each. Both sources now order by name rather than by the order the source yielded. The hypervisor cannot know declaration order, and two pictures laid out differently cannot be compared, which is the whole point of having both. Fixed while testing: the gap size and the box placement each decided separately which network a gateway sat above, and disagreed — reserving the gap above one sibling while drawing the box above the other, which landed a gateway on top of a machine in an unrelated network. Both now read one map. Five new tests, run across every scenario: no link crosses a network it does not touch, no box is drawn inside a network it is not on, a network behind another is indented inside it, a public network is not split apart by another group, and both sources lay the same topology out identically. |
||
|
|
2243618f01 |
Draw a scenario, from the declaration and from the hypervisor
`mesh-lab diagram` renders a scenario as draw.io, from either source, through one layout — so a difference between what was asked for and what exists is a difference you can see. The shape says what a resource is and is fixed per kind. The badges say what is true about that particular one and come entirely from metadata: translation, forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The interesting properties of a network are exactly the ones with no visual consequence — a translated address looks identical to an untranslated one. For the live picture to be a record rather than a restatement, raise now writes down what it applied: a segment's kind, ranges and MTU on the link; a gateway's translation, forwardability and expiry on the gateway; inbound: deny on the machine. Every behavioural tag is written AFTER the thing works, never at creation — a failed raise leaves wreckage standing on purpose, and a picture of that wreckage must not badge translation the router never got. The pairing earned itself immediately: drawn side by side, every virtual machine held no addresses. A container's interface carries the device's name and a VM names its own, so joining them by name silently dropped one whole class of machine. Fixed by joining on MAC. Also brings tests under the typecheck gate, which caught integration timeouts being passed as a 4th argument and therefore ignored entirely. |
||
|
|
ca2bbab836 |
Integration tests, each named for the decision it defends
Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of it the half that touches the hypervisor, and no gate. The verification I had done was real — pings across NAT, TTL counts, ruleset comparisons — and none of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature. Ten integration tests against a real hypervisor, each named for what it defends. ADR 0031: a raised machine carries no overlay, no wireguard, no mesh config — a scenario that pre-built peering would certify its own work. ADR 0032: exec is the only way in. ADR 0033: routers are containers while machines are virtual machines. And the design's claims: raise waits for usable, snapshots are whole-scenario, NAT hides a private address, published reaches the machine at the gateway's address. Mocking the hypervisor is forbidden, so they skip with a reason on a machine that cannot raise scenarios rather than passing green having checked nothing. The suite earned itself on its first run. It found that a snapshot of a running machine could miss a file written seconds earlier — not stale, absent — because the write was still in the guest's page cache. That is exactly the question the lifecycle design listed as open: does a scenario snapshot need the machines stopped? It does not, but it does need them flushed. snapshot now syncs every machine before capturing, and the design records the answer. The fix buys write-durability, not application-consistency: anything mid-transaction is still captured mid-transaction, and that is now stated rather than assumed. npm run check is the gate — typecheck, 40 unit tests, 10 integration tests. |
||
|
|
5d01006eab |
Transit, host firewalls, and the whole topology raising
The full topology now raises: four machines, three routers, a transit router, six segments, in 35 seconds. Everything the declaration model can express except `place`, which is refused because the node host it would place does not exist yet. Transit was a real gap, not a bug. The design says public networks are unrelated and routed to each other, never bridged — and I built the segments and never built the thing that routes between them, so three public networks were islands and nothing crossed. A transit router now holds an interface on every public segment, forwarding and no translation: the closest thing the lab has to the internet, deliberately dumb. Proven rather than asserted, by ping TTL across the raised topology: within one segment ttl=64 no hops across two unrelated public networks ttl=62 gateway + transit multicast between public networks 0 replies A flat internet would have shown ttl=64 and answered multicast — which would let a node discover a peer it could never reach in production, and report success. That is the fault the as-is layer records the mesh already hitting with multicast name resolution. inbound: deny is implemented as a host firewall on the machine, read back after applying. A declared refusal that silently did not load leaves the machine wide open, which looks exactly like a machine that is working. Established and related traffic is accepted, so a defended machine can still dial out rather than being a disconnected one. Verified by running, all of it: home -> devices (policy allow) reachable devices -> home (policy deny) blocked behind unforwardable NAT -> out reachable in -> behind unforwardable NAT unreachable inbound: deny, dialling out reachable reaching a machine that denies inbound refused The two routers differ exactly as declared: the forwardable one carries the policy rule and no inbound drop, the unforwardable one carries `ct state new drop` and no DNAT. |
||
|
|
a270cd5b02 |
Routers: NAT, port forwarding, policy and mapping expiry
A gateway is the one implicit machine in a declaration — a scenario says a segment sits behind one and never names the thing that serves it. This materialises it. A router is a container, not a virtual machine, because it is scenery rather than something under test (hq ADR 0033). Verified before building that a plain unprivileged container can do all of it: ip_forward and ipv6 forwarding settable, nftables masquerade accepted, and the conntrack timeouts mapping_ttl depends on both writable. No privileged mode. Verified by running, on a machine behind a household gateway reached from one on a routable address: home-server -> anchor 0% loss, through masquerade anchor -> 192.168.1.135 (private, direct) unreachable anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200 The last line is the published-but-behind-NAT case research 004 says only exists in production. It is now a 32-second scenario on a workstation. Segments sharing a gateway declaration share ONE router — that is what a VLAN-capable router is, and two routers sharing an external address would not work anyway. mapping_ttl is read back after setting rather than assumed. Those sysctls are not on every kernel, and a scenario that declared an expiring mapping and silently got a permanent one would be exactly the fault being built against. Four bugs found by running it, three of them the same fault — a failure made invisible. The router had no route to a package repository, by design, so installing nftables at raise time could not work. The image is now built once with temporary connectivity and cached; every scenario after that needs no network. That failure was hidden behind `|| true`, which is why it took a raise to find. The builder then failed on DNS: exec works before a container has an address, and I had treated usable as ready. It now waits for the thing actually needed. The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time networking service flushed the static address the scenario set — on eth0 only, so the outside interface came up bare while inside ones were fine. The image build now neutralises it: a router reconfiguring itself from an image default is the lab overriding the declaration. `ip addr add … || true` had hidden this too, and is now `ip addr replace` with no swallow. And routers were orphaned by destroy, holding their networks open so destroy reported removing zero segments. They now carry the same machine tag as everything else, so one query finds an instance's resources. |
||
|
|
a27d861d3b |
Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a workstation, not asserted: two machines raised and addressed in 14.6s, snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no loss, and the workstation with no route into any of it. The declaration layer implements the model in full — three positions a machine can be in, keyed on forwardability; gateways carrying the address the world sees them as; both address families; multi-homing; MTU; inter-segment policy. It is validated hard because the failures it prevents are silent: a private range on a public segment produces no error, the mesh simply never forms. Public segments are refused unless they use RFC 5737 or RFC 3849 space, and a range wider than the reserved block is refused too. 33 tests, all offline. The runtime implements less than the model, and refuses the difference. A scenario declaring gateways, published ports, policy, inbound deny or place is rejected at raise with every gap named. Raising it would produce a mesh that silently lacks what it declared, which is the fault this lab exists to catch — 04-ISSUES/003, where a firewall key is declared in five manifests and read by no code. Three bugs found by review and by running it, all of one family: The readiness check truthiness-tested incusOk's return. `exec … true` succeeds with EMPTY output, so every machine reported unreachable while incus exec on it worked perfectly. succeeds() now exists so the mistake is not available, and network delete had the same bug — it counted zero segments removed while removing them. list() split instance from machine on the last dash, so a machine called home-server absorbed half the instance id and destroy found nothing. Resources are now found by the metadata they carry, never by name. restore reported success in 0.79s while the machine's agent was still starting, so the next command failed. Both raise and restore now wait for usable and say how long that took — reporting the earlier number is transport reported as effect, which is the fault the lab is being built to find. Two incus behaviours worth recording. Its CLI reads a YAML definition from stdin when stdin is not a terminal, so a spawned command hangs until the timeout kills it and arrives with empty stderr — a failure with no explanation, on a command that works when typed. And it assigns a MAC at runtime without recording it in device config, so MACs are derived and set explicitly, which the guest needs anyway: it names interfaces by bus position, and matching by name configures the wrong one on a multi-homed machine. No build step; Node strips the types. The lifecycle has no unit tests because a fake hypervisor would assert that the fake behaves as expected, which is the shape of test this project exists to stop shipping. |