Commit Graph
14 Commits
Author SHA1 Message Date
jschoubben 751948f0f9 The lab had a registry that production does not, so it tested a fiction
The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and
rewrote every manifest reference — third-party ones included — to point at it. No
production mesh has such a thing. So every bed proved that a machine could fetch an
image from a registry that exists nowhere else, and the bootstrap problems that only
appear when a machine has to fetch for itself went unfound.

What replaces it is the two things that are true in the world:

**Public images come from the public internet.** mesh-lab already created a NAT'd
uplink for exactly this and attached it to any machine declaring `egress`; no scenario
ever declared it. They do now, and third-party references are left exactly as the
catalogue writes them.

**The mesh's own images have no registry and never will.** mesh-control, mesh-builder,
mesh-route-proxy and the per-module runtimes are built from source and exist in no
registry. A machine gets them the way an operator's machine does — they are built here
and loaded onto it — and is then named by the digest of its own image configuration,
which mesh-host now accepts as "an image this machine already holds".

`images:` therefore means only *ours*, and a third-party entry is refused rather than
quietly loaded: otherwise the fiction returns one convenient line at a time. It is
per-machine as well, because "everything, everywhere" was never a description of
anything real — handing whole-mesh-full's union to its two 30GiB workstations would
fill the disk with runtimes nothing on them will start.

**The uplink and the declared gateway would have fought, silently.** A gateway container
and the transit router reach the scenario and nothing else; a default route through
either is a black hole for anything outside, and it beats the uplink's DHCP route on
metric. So a machine with egress states the scenario's ranges explicitly — through the
same gateway or transit it would have defaulted to, so the overlay-across-NAT path is
unchanged — and leaves the default to the uplink. A range with no path inside the
scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an
ordinary private range in fact, and letting it escape would put scenario traffic on
whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested,
because a decision only a full raise could check is one nobody checks.

The registry-reachability check the raise gained earlier is kept, pointed at the real
thing: every machine with egress must resolve a name and reach the internet before the
raise says it finished. Same failure it was written for — a raise that returns, an apply
that dies on its first pull, an instance left a bare shell — now guarding the path that
actually carries.

The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It
was never only for the lab's registry: the mesh has one of its own, the `registry`
module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:05 +02:00
jschoubben 2f4cb871d9 A raise does not finish until the machines can pull from the registry
The first whole-mesh raise of the ADR 0056 code died on the anchor's substrate
apply: the image pulls failed, the anchor never came up, no node could enrol,
and the instance was left a bare shell — VMs and a registry, no substrate. The
identical apply, run by hand once the registry was warm, succeeded immediately.

`raiseRegistry` proves the wrong thing. It curls `localhost:5000` from inside
the registry's OWN machine, which says the registry process is up and holds the
blobs, and says nothing about the path anybody else uses: across a segment, and
for the home nodes through a NAT gateway whose default route and firewall are
applied two steps LATER. So "serving" was reported on evidence that excluded the
network, and the caller — which pins every image in the substrate bundle to that
registry — was handed a fact it could not rely on.

So the check moves to where it means something. After the routes and the
firewalls, before the minutes spent placing, each machine is asked for `/v2/` and
for one stocked manifest BY DIGEST, at the address it will pin, over the network
it will use. That is the pair of requests a pull begins with, from the same
place. Layers are not fetched: every digest was already read back inside the
registry machine, so what is in question here is the path, not the content.

Verified by typecheck and the unit suite (136 pass), and by confirming against a
standing four-node instance that `curl` exists in the machines and that both
segments — including a home node through the gateway — answer 200 for the
registry's `/v2/`. The ordering itself is unverified in a live raise from cold,
which takes hours.
2026-09-10 21:05:52 +02:00
jschoubben 80b0670ebe whole-mesh-full: the real segmented topology, and the overlay proven across the access point
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.

The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.

Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
  token carries that verbatim as the endpoint a node dials, so with the substrate
  on novox it must be novox's own public address. Rewritten at apply (the cert is
  fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
  5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
  Both provider host publishes are remapped off the substrate's ports.

And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 11:17:00 +02:00
jschoubben a9ecce25cb Two-node DB-consumer bed + a scenario disk field
The GREEN multi-node regression bed that proves the DB-consumer gate: substrate/control on
one node, postgres+redis providers and baserow+letta consumers on another, each consumer
getting its own credential and its own mesh-named database across the overlay. Requires the
mesh-control provider-seal-key fix and the mesh-catalog db-name fix.

Includes a general lab capability: a machine 'disk' field sizing the VM root disk (a broad
install exhausts the pool default and the host fails mid-apply with 'no space left on
device'). The bed sets 60GiB.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 13:48:26 +02:00
jschoubben 4a343a2652 A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test.

The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.

So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.

`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.

And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
2026-09-01 21:42:18 +02:00
jschoubben bb14ecb7e0 Say what the lab is doing, while it is doing it
novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.

Three decisions, each doing work.

**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.

**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.

**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.

Two things this found in itself while being written, both the same shape
as what it exists to catch:

A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.

And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.

Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
2026-09-01 10:43:38 +02:00
jschoubben be176bab2e Automate the lab registry: a sealed machine pulls by digest
Closes 04-ISSUES/009. A scenario declares `images:` by tag; the lab stocks a
registry on this workstation where there is a network, raises it inside the
scenario as scenery, and reports the references a declaration pins -- which are
the digests THIS registry assigned, and are not knowable until it is raised.

Verified in a sealed machine, confirmed by ping to have no route out: package,
service including boot state, a container pinned by digest, and an action
inside that container. Applied, idempotent on re-apply, and read back from the
machine rather than from the apply's own report. That is the first time the
container shape has worked in the lab at all, and it was the shape blocking the
substrate bootstrap.

Four faults found by running it, three of them mine and one worth keeping:

The read-back checked that the catalog endpoint answered, by looking for the
substring "repositories" -- which `{"repositories":[]}` also contains. So it
passed on a registry holding nothing, and the failure surfaced much later as a
container that could not be pulled. It now asks for each image's manifest BY
DIGEST, which is what a machine does.

A recursive push needs its destination to exist, or incus copies the source's
contents rather than the source. The data landed one directory too shallow and
the registry found nothing where it looks.

The registry writes its blobs as root through a bind mount, so the workstation
could not remove its own scratch directory afterwards. Whoever made the files
removes them -- the cleanup now runs in a container too. And a cleanup failure
no longer fails a raise that succeeded: the scenario is standing and usable,
and saying otherwise would be a false report.

The base image build did not verify that the runtime trusts the documentation
ranges as plain-HTTP registries. Writing the file is not the daemon honouring
it, and a base image that looks right fails much later, in a sealed scenario,
a long way from its cause. It is now read back from `docker info`.
2026-08-29 00:04:55 +02:00
jschoubben 4097ff92c1 Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
2026-08-28 23:33:46 +02:00
jschoubben d6eef25590 The lab can give a sealed machine a container runtime
ADR 0046's open consequence: "the lab needs a way to place images, and the
machine it places them into needs a container runtime, which a sealed scenario
cannot install either."

The runtime half is done, and it is research 012's reframing applied literally
-- fetch at build time on a machine with a network, apply on a target that
needs nothing. `mesh-lab base build` launches a machine WITH a network,
installs a runtime, verifies it by asking the runtime rather than the package
manager, and publishes the result. Measured: ~30s to install, ~60s to publish,
~700MiB, paid once per lab rather than per scenario.

A scenario that places `runtime` or an image is then raised from that base
image, chosen rather than declared -- a scenario says what it needs, not which
image provides it. If the base does not exist it says so and how to build it.

Verified in a genuinely sealed machine (no route out, confirmed by ping):
package, service including the new `boot: enabled`, and action all applied,
were idempotent on a second run, and read back correctly. Those three had never
run anywhere but a workstation.

The image half is NOT done, and testing found why: a digest-pinned image cannot
be placed from an archive. `docker save alpine@sha256:...` produces an archive
with no repo tag, because a repo digest only exists for an image a registry
served -- so it loads dangling and a container declaring that digest reaches
for a registry the machine cannot see.

That collides with ADR 0046, which has the host REFUSE an unpinned image. Tag
refused by the host, digest unusable in the lab: there is currently no
declaration the lab can raise that exercises the container shape at all. Filed
as 04-ISSUES/009, whose resolution is a registry inside the scenario -- which is
what the real mesh does rather than a workaround for the lab.

Also fixed a weak check of my own, which is the same fault in miniature: the
load was tested with `includes("Loaded image")`, a prefix of both `Loaded
image:` and `Loaded image ID:`. So an unusable dangling load reported success
and the failure surfaced later as a container that would not start.
2026-08-28 01:56:32 +02:00
jschoubben 16c13807a9 place: the host — the lab acquires a consumer
The lab raised an underlay and put nothing on it: correct, and useless, because
the thing it exists to test did not exist. Tier 0 now does, so `place: [host]`
works and a raised scenario finally contains something.

The refusal narrows rather than disappearing. A scenario placing a host and a
substrate is told which half is missing, by name — not that `place:` is
unsupported when half of it now works.

Placement reads back rather than assuming. A file arriving is not a host
working, so the binary is run before it is trusted to answer questions, and what
it reports is read from the machine (ADR 0035). The binary comes from an
explicit path, because the declaration design leaves where artifacts come from
open and a search would harden into the answer by accident.

The integration test that matters is the one asserting the host reports the
MACHINE and not the workstation that placed it. A raised VM and this workstation
differ in every capability — root versus uid 1000, a clean init versus a
degraded one, no docker versus docker, no wireguard versus wg0 — so a host
reporting the wrong machine is obvious here and invisible anywhere else.

And the placed host independently confirms ADR 0031: overlay absent on a freshly
raised machine. The underlay suite already asserted that by looking for
wireguard interfaces; this is a second witness rather than the same check twice.

Two tests failed the moment placement worked, which is what they were for. They
defended "there is nothing to place yet" while that was true; the decision
changed, so they change with it rather than being deleted.

Gate: 75 unit, 20 integration.
2026-08-26 00:41:25 +02:00
jschoubben 2243618f01 Draw a scenario, from the declaration and from the hypervisor
`mesh-lab diagram` renders a scenario as draw.io, from either source, through
one layout — so a difference between what was asked for and what exists is a
difference you can see.

The shape says what a resource is and is fixed per kind. The badges say what is
true about that particular one and come entirely from metadata: translation,
forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The
interesting properties of a network are exactly the ones with no visual
consequence — a translated address looks identical to an untranslated one.

For the live picture to be a record rather than a restatement, raise now writes
down what it applied: a segment's kind, ranges and MTU on the link; a gateway's
translation, forwardability and expiry on the gateway; inbound: deny on the
machine. Every behavioural tag is written AFTER the thing works, never at
creation — a failed raise leaves wreckage standing on purpose, and a picture of
that wreckage must not badge translation the router never got.

The pairing earned itself immediately: drawn side by side, every virtual machine
held no addresses. A container's interface carries the device's name and a VM
names its own, so joining them by name silently dropped one whole class of
machine. Fixed by joining on MAC.

Also brings tests under the typecheck gate, which caught integration timeouts
being passed as a 4th argument and therefore ignored entirely.
2026-08-24 22:53:00 +02:00
jschoubben 5d01006eab Transit, host firewalls, and the whole topology raising
The full topology now raises: four machines, three routers, a transit
router, six segments, in 35 seconds. Everything the declaration model can
express except `place`, which is refused because the node host it would
place does not exist yet.

Transit was a real gap, not a bug. The design says public networks are
unrelated and routed to each other, never bridged — and I built the
segments and never built the thing that routes between them, so three
public networks were islands and nothing crossed. A transit router now
holds an interface on every public segment, forwarding and no translation:
the closest thing the lab has to the internet, deliberately dumb.

Proven rather than asserted, by ping TTL across the raised topology:

  within one segment                     ttl=64   no hops
  across two unrelated public networks   ttl=62   gateway + transit
  multicast between public networks      0 replies

A flat internet would have shown ttl=64 and answered multicast — which
would let a node discover a peer it could never reach in production, and
report success. That is the fault the as-is layer records the mesh already
hitting with multicast name resolution.

inbound: deny is implemented as a host firewall on the machine, read back
after applying. A declared refusal that silently did not load leaves the
machine wide open, which looks exactly like a machine that is working.
Established and related traffic is accepted, so a defended machine can
still dial out rather than being a disconnected one.

Verified by running, all of it:

  home -> devices (policy allow)               reachable
  devices -> home (policy deny)                blocked
  behind unforwardable NAT -> out              reachable
  in -> behind unforwardable NAT               unreachable
  inbound: deny, dialling out                  reachable
  reaching a machine that denies inbound       refused

The two routers differ exactly as declared: the forwardable one carries the
policy rule and no inbound drop, the unforwardable one carries `ct state
new drop` and no DNAT.
2026-08-24 01:49:30 +02:00
jschoubben a270cd5b02 Routers: NAT, port forwarding, policy and mapping expiry
A gateway is the one implicit machine in a declaration — a scenario says a
segment sits behind one and never names the thing that serves it. This
materialises it.

A router is a container, not a virtual machine, because it is scenery
rather than something under test (hq ADR 0033). Verified before building
that a plain unprivileged container can do all of it: ip_forward and ipv6
forwarding settable, nftables masquerade accepted, and the conntrack
timeouts mapping_ttl depends on both writable. No privileged mode.

Verified by running, on a machine behind a household gateway reached from
one on a routable address:

  home-server -> anchor                      0% loss, through masquerade
  anchor -> 192.168.1.135 (private, direct)  unreachable
  anchor -> 192.0.2.50:8080 (the GATEWAY)    HTTP 200

The last line is the published-but-behind-NAT case research 004 says only
exists in production. It is now a 32-second scenario on a workstation.

Segments sharing a gateway declaration share ONE router — that is what a
VLAN-capable router is, and two routers sharing an external address would
not work anyway.

mapping_ttl is read back after setting rather than assumed. Those sysctls
are not on every kernel, and a scenario that declared an expiring mapping
and silently got a permanent one would be exactly the fault being built
against.

Four bugs found by running it, three of them the same fault — a failure
made invisible.

The router had no route to a package repository, by design, so installing
nftables at raise time could not work. The image is now built once with
temporary connectivity and cached; every scenario after that needs no
network. That failure was hidden behind `|| true`, which is why it took a
raise to find.

The builder then failed on DNS: exec works before a container has an
address, and I had treated usable as ready. It now waits for the thing
actually needed.

The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time
networking service flushed the static address the scenario set — on eth0
only, so the outside interface came up bare while inside ones were fine.
The image build now neutralises it: a router reconfiguring itself from an
image default is the lab overriding the declaration. `ip addr add … || true`
had hidden this too, and is now `ip addr replace` with no swallow.

And routers were orphaned by destroy, holding their networks open so
destroy reported removing zero segments. They now carry the same machine
tag as everything else, so one query finds an instance's resources.
2026-08-24 01:37:19 +02:00
jschoubben a27d861d3b Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.

The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.

The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.

Three bugs found by review and by running it, all of one family:

The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.

list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.

restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.

Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.

No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
2026-08-24 01:12:49 +02:00