Commit Graph
9 Commits
Author SHA1 Message Date
jschoubben 751948f0f9 The lab had a registry that production does not, so it tested a fiction
The lab raised a `registry` VM, pushed ~73 images into it from the workstation, and
rewrote every manifest reference — third-party ones included — to point at it. No
production mesh has such a thing. So every bed proved that a machine could fetch an
image from a registry that exists nowhere else, and the bootstrap problems that only
appear when a machine has to fetch for itself went unfound.

What replaces it is the two things that are true in the world:

**Public images come from the public internet.** mesh-lab already created a NAT'd
uplink for exactly this and attached it to any machine declaring `egress`; no scenario
ever declared it. They do now, and third-party references are left exactly as the
catalogue writes them.

**The mesh's own images have no registry and never will.** mesh-control, mesh-builder,
mesh-route-proxy and the per-module runtimes are built from source and exist in no
registry. A machine gets them the way an operator's machine does — they are built here
and loaded onto it — and is then named by the digest of its own image configuration,
which mesh-host now accepts as "an image this machine already holds".

`images:` therefore means only *ours*, and a third-party entry is refused rather than
quietly loaded: otherwise the fiction returns one convenient line at a time. It is
per-machine as well, because "everything, everywhere" was never a description of
anything real — handing whole-mesh-full's union to its two 30GiB workstations would
fill the disk with runtimes nothing on them will start.

**The uplink and the declared gateway would have fought, silently.** A gateway container
and the transit router reach the scenario and nothing else; a default route through
either is a black hole for anything outside, and it beats the uplink's DHCP route on
metric. So a machine with egress states the scenario's ranges explicitly — through the
same gateway or transit it would have defaulted to, so the overlay-across-NAT path is
unchanged — and leaves the default to the uplink. A range with no path inside the
scenario becomes `unreachable` rather than falling through: 192.168.1.0/24 is an
ordinary private range in fact, and letting it escape would put scenario traffic on
whatever network the workstation is sitting on. `scenarioRoutesFor` is pure and tested,
because a decision only a full raise could check is one nobody checks.

The registry-reachability check the raise gained earlier is kept, pointed at the real
thing: every machine with egress must resolve a name and reach the internet before the
raise says it finished. Same failure it was written for — a raise that returns, an apply
that dies on its first pull, an instance left a bare shell — now guarding the path that
actually carries.

The base image's trust of the documentation ranges as plain-HTTP registries STAYS. It
was never only for the lab's registry: the mesh has one of its own, the `registry`
module, serving artifacts to the whole mesh over plain HTTP from whatever node runs it.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:05 +02:00
jschoubben eba436b6b7 Reach one scenario from the workstation, by name
A scenario is a closed address space: two raised from the same
declaration hold the same addresses and never meet, which is what lets
two run at once and why the lab talks to machines through the
hypervisor rather than over IP. Reaching in from outside breaks that, so
it is opt-in, one scenario at a time, and reversible.

`connect` takes an address on the scenario's public link and writes a
resolver rule answering everything under each machine's name.
`disconnect` gives both back. `connected` says what is true right now,
for somebody who cannot remember.

It refuses rather than guessing when more than one scenario is standing
— the failure being avoided is not an error but one scenario's traffic
arriving in another. It also refuses when a machine's name is already
answered here for something real, because connecting would point that
name at the lab, and the damage would land on the real thing.

Names answer with the segment address rather than the overlay one.
Inside the mesh a name gives a machine's private address; from here that
would need this workstation on the overlay, which is a much larger door.
The segment address reaches the same machine and the same ports, which
is what opening a board in a browser actually needs.

Proven against a live two-node scenario: registry.internal:5000/v2/
answered 200 from this workstation, and so did a wildcard name under the
same machine. Disconnect put the address back, stopped answering, and
left the real mesh's own names alone.

One thing measured rather than assumed: it restarts dnsmasq instead of
reloading it. A reload is SIGHUP, which re-reads the hosts file and
clears the cache but not the configuration — the rule was written, the
reload reported success, and nothing resolved. The daemon's start time
was nine days old afterwards.
2026-09-01 16:42:26 +02:00
jschoubben a516ee847b Adopt a third-party workload, and keep a mesh between runs
**The adoption.** Software nobody here wrote, taking its credentials the
way such software does — from its environment — and needing two
containers that reach each other by name. The first module that could
not have been declared this morning: it needs the network shape and it
needs a sealed value to reach a container's environment.

Its password is accepted rather than generated, which is the whole shape
of an adoption: a service that already exists keeps the credential it
already has. Asserted properly — a wrong password is refused by the same
database, so the passing case means something.

**The warm scenario.** A mesh kept between runs and returned to, which
turned twelve minutes of bootstrap into thirty seconds of restore. Off
unless asked for: a run that is meant to mean something raises from
nothing.

Its guard fired for real during this work, unprompted — a mesh-host
commit landed and it refused the stale base, naming both commits, rather
than passing tests against yesterday's binary. That is 04-ISSUES/005's
rule one level down.

Three things the guard learned the hard way and now handles: a snapshot
captures disk and not memory, so the host is restarted after a restore
and asserted to have come back; the stocked image digests are worked out
while raising and a restored instance never raises, so they are kept;
and comparing only the repositories this run can see clears the ones it
cannot, so both directions are compared.

The one real bug behind five failed attempts was in mesh-host and it
reported itself precisely: a network shape the language had and no host
implemented. Everything else was scaffolding of mine.
2026-09-01 01:20:14 +02:00
jschoubben 033ad7ec69 A run rebuilds what it tests, and leaves a receipt saying what it covered
The danger is not that the suite breaks. It is that nobody notices it
stopped running (novox/hq 04-ISSUES/005). The harness this replaces had
not built for two and a half months and nothing said so — and this suite
needs a hypervisor, so it inherits exactly that: it runs when somebody
remembers, and remembering is not a mechanism.

So running, recording, and rebuilding are one act:

- the host binary, control-plane image and builder are rebuilt from
  source first. The last two both parse manifests; building one and not
  the other left a binary eleven hours old refusing a field the mesh had
  just renamed, found by a full run.
- a receipt lands in XDG state — outside git, because the question is
  whether *this machine* has run it, and a receipt in git would be a
  claim about everybody's machine made by whoever committed last.
- `last-run` judges it and exits non-zero when it no longer counts.

Three faults found by running the thing rather than reading it, each now
held by a test confirmed to fail without it:

- counted() passed every test while parsing nothing. The runner colours
  its summary even into a pipe; the fixtures were clean text that had
  been imagined rather than captured. A fixture that agrees with the
  mistake proves the mistake.
- a receipt for `suite test/lastrun.test.ts` was indistinguishable from
  one for the real thing — 005's own symptom, rebuilt inside its remedy.
  The receipt now records what ran.
- a tree with uncommitted work reported the bare commit, claiming
  coverage of code nobody can check out. Nothing else could tell: the
  hash is identical either way.

Proven on real machines: 22/22, against all three repositories.
2026-08-31 15:02:19 +02:00
jschoubben be176bab2e Automate the lab registry: a sealed machine pulls by digest
Closes 04-ISSUES/009. A scenario declares `images:` by tag; the lab stocks a
registry on this workstation where there is a network, raises it inside the
scenario as scenery, and reports the references a declaration pins -- which are
the digests THIS registry assigned, and are not knowable until it is raised.

Verified in a sealed machine, confirmed by ping to have no route out: package,
service including boot state, a container pinned by digest, and an action
inside that container. Applied, idempotent on re-apply, and read back from the
machine rather than from the apply's own report. That is the first time the
container shape has worked in the lab at all, and it was the shape blocking the
substrate bootstrap.

Four faults found by running it, three of them mine and one worth keeping:

The read-back checked that the catalog endpoint answered, by looking for the
substring "repositories" -- which `{"repositories":[]}` also contains. So it
passed on a registry holding nothing, and the failure surfaced much later as a
container that could not be pulled. It now asks for each image's manifest BY
DIGEST, which is what a machine does.

A recursive push needs its destination to exist, or incus copies the source's
contents rather than the source. The data landed one directory too shallow and
the registry found nothing where it looks.

The registry writes its blobs as root through a bind mount, so the workstation
could not remove its own scratch directory afterwards. Whoever made the files
removes them -- the cleanup now runs in a container too. And a cleanup failure
no longer fails a raise that succeeded: the scenario is standing and usable,
and saying otherwise would be a false report.

The base image build did not verify that the runtime trusts the documentation
ranges as plain-HTTP registries. Writing the file is not the daemon honouring
it, and a base image that looks right fails much later, in a sealed scenario,
a long way from its cause. It is now read back from `docker info`.
2026-08-29 00:04:55 +02:00
jschoubben 4097ff92c1 Repoint ADR references after HQ consolidated 65 records to 23
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
2026-08-28 23:33:46 +02:00
jschoubben d6eef25590 The lab can give a sealed machine a container runtime
ADR 0046's open consequence: "the lab needs a way to place images, and the
machine it places them into needs a container runtime, which a sealed scenario
cannot install either."

The runtime half is done, and it is research 012's reframing applied literally
-- fetch at build time on a machine with a network, apply on a target that
needs nothing. `mesh-lab base build` launches a machine WITH a network,
installs a runtime, verifies it by asking the runtime rather than the package
manager, and publishes the result. Measured: ~30s to install, ~60s to publish,
~700MiB, paid once per lab rather than per scenario.

A scenario that places `runtime` or an image is then raised from that base
image, chosen rather than declared -- a scenario says what it needs, not which
image provides it. If the base does not exist it says so and how to build it.

Verified in a genuinely sealed machine (no route out, confirmed by ping):
package, service including the new `boot: enabled`, and action all applied,
were idempotent on a second run, and read back correctly. Those three had never
run anywhere but a workstation.

The image half is NOT done, and testing found why: a digest-pinned image cannot
be placed from an archive. `docker save alpine@sha256:...` produces an archive
with no repo tag, because a repo digest only exists for an image a registry
served -- so it loads dangling and a container declaring that digest reaches
for a registry the machine cannot see.

That collides with ADR 0046, which has the host REFUSE an unpinned image. Tag
refused by the host, digest unusable in the lab: there is currently no
declaration the lab can raise that exercises the container shape at all. Filed
as 04-ISSUES/009, whose resolution is a registry inside the scenario -- which is
what the real mesh does rather than a workaround for the lab.

Also fixed a weak check of my own, which is the same fault in miniature: the
load was tested with `includes("Loaded image")`, a prefix of both `Loaded
image:` and `Loaded image ID:`. So an unusable dangling load reported success
and the failure surfaced later as a container that would not start.
2026-08-28 01:56:32 +02:00
jschoubben 2243618f01 Draw a scenario, from the declaration and from the hypervisor
`mesh-lab diagram` renders a scenario as draw.io, from either source, through
one layout — so a difference between what was asked for and what exists is a
difference you can see.

The shape says what a resource is and is fixed per kind. The badges say what is
true about that particular one and come entirely from metadata: translation,
forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The
interesting properties of a network are exactly the ones with no visual
consequence — a translated address looks identical to an untranslated one.

For the live picture to be a record rather than a restatement, raise now writes
down what it applied: a segment's kind, ranges and MTU on the link; a gateway's
translation, forwardability and expiry on the gateway; inbound: deny on the
machine. Every behavioural tag is written AFTER the thing works, never at
creation — a failed raise leaves wreckage standing on purpose, and a picture of
that wreckage must not badge translation the router never got.

The pairing earned itself immediately: drawn side by side, every virtual machine
held no addresses. A container's interface carries the device's name and a VM
names its own, so joining them by name silently dropped one whole class of
machine. Fixed by joining on MAC.

Also brings tests under the typecheck gate, which caught integration timeouts
being passed as a 4th argument and therefore ignored entirely.
2026-08-24 22:53:00 +02:00
jschoubben a27d861d3b Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.

The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.

The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.

Three bugs found by review and by running it, all of one family:

The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.

list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.

restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.

Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.

No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
2026-08-24 01:12:49 +02:00