Commit Graph
29 Commits
Author SHA1 Message Date
jschoubben fc36fb5b51 catalogue-media: two modules accessing one operator-owned directory co-resolve (ADR 0051)
sonarr and radarr both access /services/media/downloads — the exact duplicate path
the resolver refused before novox/hq ADR 0051 (04-ISSUES/036, 012). Each now declares
it as an `access`, not a `directory` resource, so the pair co-resolves and one push
configures both. The operator provides the shared media dirs before apply (the host
refuses an absent access); the bed creates them after enrol and before the push.

Proves: the push is not refused, the node converges once, both modules' server and
runtime containers are up, and both server containers mount the same operator-owned
spool.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:02:41 +02:00
jschoubben 571a6cd6fd WIP: catalogue-apps install bed (mongodb, unifi, marrytts)
Adds the mongodb runtime CLI (mongosh) to build-module-runtime.sh, a
catalogue-apps scenario, and its install test. Proven so far: the ADR-0054 slug
applies and the mesh accepts the push (mongodb consumer identity mesh_anchor_mongo
fits). NOT green: the node applies but never reaches 'current' within 1200s — a
persistent reconcile divergence (applied-but-never-current, no crash), likely a
module declaring a resource its container mutates (issue-011 class). Needs live
VM inspection to name the module. Not merged.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 15:30:50 +02:00
jschoubben ab04dd814f Add catalogue-small co-residence bed (four modules, one push)
Raises the first-node substrate and assigns postgres, minio, redis and
plex to one anchor in a single push, proving they resolve and come up
together on one node. postgres and minio each get a consumer that
connects with a real granted credential.

redis follows the corrected provider contract (ADR 0048, issue 032): its
runtime reconciles the contributions the mesh delivers at MESH_RECEIVES
and creates each consumer's ACL user with the mesh-minted password,
sealing nothing — no MESH_SEAL_KEY, no *.grant.json/*.credential path.
The provisioning proof authenticates as the consumer with the mesh's
password (PONG), matching the green provider-uses-mesh-credential bed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 14:14:56 +02:00
jschoubben b7ba7534af e2e: the whole grant for an S3 bucket (skipped — blocked on hq issue 010)
Mirrors the postgres bed for minio: a provider (runtime carries mc) + a consumer
requiring s3-bucket, proving the consumer reaches its bucket with the access key and
secret the mesh delivered. It surfaced a real limit: the mesh derives `as` =
mesh_<node>_<module> (22 chars), and an S3 access key is capped at 20, so minio refuses
the service account. The test is correct and skipped pending 04-ISSUES/010, not worked
around.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:58:01 +02:00
jschoubben de7f3b9c05 e2e: the whole grant for a database — a consumer connects with what the mesh delivered
Assigns a postgres provider (its runtime carries psql) and a module that requires
postgres-database; the mesh mints one password, postgres's provisioner creates a role
and database under the mesh's login with it, and the consumer connects to its database
with the delivered credential (a password-checked connection) — select 1. Nothing placed
by the test. The postgres half of the per-backend provider proof (ADR 0052/0053).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 01:47:57 +02:00
jschoubben 0a414d576a e2e: assigned sonarr + grafana prove the two runtime config paths (ADR 0051/0052)
assigned-sonarr proves the Servarr detection path: the runtime discovers its API
key from the app's config.xml and serves its tools. assigned-grafana proves the
settings path: the operator states URL and token as settings, the mesh merges them
into the module's config file, and the runtime serves from that with nothing in the
manifest. Both green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:08:54 +02:00
jschoubben f093769354 e2e: assigned plex + redis prove the module runtime (ADR 0052)
build-module-runtime.sh generalises the audit-logger runtime image to any module
(mesh-tools + sdk + the module's dist, entrypoints for tools/events/provisioner).
Two scenarios and two tests: assigned-plex proves a tools+events module serves its
tools over a mesh-issued scoped account; assigned-redis proves a provider's runtime
serves tools AND runs its provisioner in the same broker-bound process, provisioning
a grant and emitting its lifecycle event. Both green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 22:29:53 +02:00
jschoubben 37b16cd0e9 events: the assigned audit-logger, proven in the lab (ADR 0048)
test/integration/assigned-audit.test.ts raises a node into a mesh, assigns it
the audit-logger through the control plane, and asserts the mesh delivered a
scoped amqps account (not the broker's own), the host ran the container, and an
emitted event reached the trail — the delivered credential authenticating is
the proof. scenarios/audit-node.yml is the lean single-node bed that stocks the
runtime image. Passes 1/1 against the real lab.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 02:13:01 +02:00
jschoubben 2f2f1d931e The cache edge, proven end to end
A consumer contributes a key prefix and gets an ACL user; the test is
that the grant means exactly what the manifest said, in both
directions: its own keys usable, anyone else's refused by the store
itself, and the flush a tenant must never have refused with them.

Waited for through the store rather than through logs: the user list,
asked with the password the host wrote into the server's own conf file
on the machine — nothing invented, both ends reading what the mesh
delivered.

And the forge is asked on the port the mesh assigned, not the one the
module declared. The old curl aimed at 3000, which was right until
ADR 0038 moved the machine side — a latent break that would have fired
on the first run to get past the settling that used to fail first.

The scenario stocks redis and its provisioner, and the rebuild builds
the provisioner image with the others.
2026-09-01 22:45:05 +02:00
jschoubben 4a343a2652 A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test.

The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.

So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.

`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.

And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
2026-09-01 21:42:18 +02:00
jschoubben 44d3e53bcb Three failures, all mine, all worth having
**A password beginning with a dash broke the search for it.** The
credential test greps the machine's own files for the delivered
password; this run's password started `-S`, so grep read it as an option
and refused the whole invocation. The test compared the usage message
against "0" and reported the password as leaked. That is the worst way
for a search to fail — it says it found something. Fixed with `-e` and
`--`, which is what those exist for.

**The planning test could not redirect what the scenario does not
serve.** Rewriting an image reference only works for repositories the
scenario's registry actually holds, and the object store's provisioner
was not stocked — so that module kept its placeholder and the refusal
fired, correctly. It is stocked now, so all five are planned again. The
skip path stays for anything genuinely unserved, and says which module
and why: a planning test quietly covering four instead of five is the
false coverage this suite exists to prevent.

**The forge failed because of the one above it.** The planning test
threw before its cleanup could run, leaving a module assigned that
refused the next push, so no database container was ever created.
Yesterday's fix moved that cleanup where a failure cannot skip it — but
`after` still only unassigns what was assigned, so it now tracks what
actually got added rather than what was intended.
2026-09-01 16:16:57 +02:00
jschoubben 55022edc1e Run the forge, on a database the mesh gave it
Everything up to now stopped at composing a declaration. That proves the
control plane and the host agree, and proves nothing about whether the
thing described works — which is how five modules sat pinned to images
that did not exist while parsing and resolving perfectly.

The forge is the right one to run first. It needs a database from
another module, a password it did not choose, and a connection string it
could not have written itself: the address and port come from what the
database serves, the user name from what the mesh decided both ends
would call it. If any of that is wrong it cannot start, and nothing else
in this suite would notice.

The test checks the chain in the order it has to happen — the login
exists, the database it owns exists, the forge answers, and its log does
not say authentication failed. That last one matters: a forge that
started and could not reach its database would still answer on its port.

Rewriting image references is now shared rather than copied from the
bundle, which had the same problem first. A digest is not knowable until
something is built, and when it is, it belongs to whichever registry
served it — so the text says which image and the scenario says which
copy. Matching is on the repository, with a test that a repository
ending in another one is not half-replaced.

Also makes the planning test put the machine back. Tests here share one
mesh, so the five modules it assigned were inherited by whatever ran
next; harmless while nothing pushed, and not harmless now.
2026-09-01 15:39:30 +02:00
jschoubben a516ee847b Adopt a third-party workload, and keep a mesh between runs
**The adoption.** Software nobody here wrote, taking its credentials the
way such software does — from its environment — and needing two
containers that reach each other by name. The first module that could
not have been declared this morning: it needs the network shape and it
needs a sealed value to reach a container's environment.

Its password is accepted rather than generated, which is the whole shape
of an adoption: a service that already exists keeps the credential it
already has. Asserted properly — a wrong password is refused by the same
database, so the passing case means something.

**The warm scenario.** A mesh kept between runs and returned to, which
turned twelve minutes of bootstrap into thirty seconds of restore. Off
unless asked for: a run that is meant to mean something raises from
nothing.

Its guard fired for real during this work, unprompted — a mesh-host
commit landed and it refused the stale base, naming both commits, rather
than passing tests against yesterday's binary. That is 04-ISSUES/005's
rule one level down.

Three things the guard learned the hard way and now handles: a snapshot
captures disk and not memory, so the host is restarted after a restore
and asserted to have come back; the stocked image digests are worked out
while raising and a restored instance never raises, so they are kept;
and comparing only the repositories this run can see clears the ones it
cannot, so both directions are compared.

The one real bug behind five failed attempts was in mesh-host and it
reported itself precisely: a network shape the language had and no host
implemented. Everything else was scaffolding of mine.
2026-09-01 01:20:14 +02:00
jschoubben bfcbee49e9 A public name, against a real ACME server
The other half of the certificate split: the mesh's own authority
certifies internal names, and a name reachable from outside needs one
the world already trusts. Against a real server rather than a stub,
because what is under test is whether an order, a challenge and a
handshake agree, and a stub would be told to agree.

One assertion passes and one fails, and the failure is filed as
novox/hq 04-ISSUES/020: the authority issues a certificate and the
client never collects it. Kept as a failing test rather than deleted or
skipped — it is the reproduction, and it proves everything up to the
last hop.

The passing one is the guard that matters day to day: no certificate is
ordered for a name the mesh does not route, so a scan cannot spend an
account's rate limit.

The failure output gathers both sides before asserting. The first
version reported only what the proxy said, which made a server-side
question unanswerable — "the client never spoke to it" and "it refused
what the client said" are different faults with nothing in common.
2026-08-31 21:40:16 +02:00
jschoubben 2096d0b2a1 Prove a bucket is provisioned the way a database is
Seven assertions against a real store, the important one being that a
consumer cannot reach another consumer's bucket — isolation here is a
policy somebody wrote rather than a boundary the product has.

The revocation test stages its own precondition. The first version
asserted a key left by an earlier test, and the rotation test had
already revoked it two tests early: the behaviour was correct and the
test was measuring residue. Its precondition assertion is what caught
that, rather than it passing green having verified nothing.
2026-08-31 17:51:12 +02:00
jschoubben 47d990b33a Prove a route reaches the workload, and does not outlive it
The request goes to the name, across the private network, and returns the
workload's own answer. Then the module is unassigned and the same request must
stop working — a stale public name pointing at nothing fails more visibly than
a stale grant.

The workload declares its port as well as its route, because they are different
questions and the earlier test leaves this machine filtering: a module that
asked for a route and not for the port would be unreachable by the proxy it
just asked for.
2026-08-31 02:43:19 +02:00
jschoubben 21a1e85d32 Prove rotation against a real database, with a real login
Two ends holding a matching string proves they agree, not that either is right.
So the check is three logins over the private network from the consumer's own
machine: the delivered credential works, the rotated one works, and the one
that was rotated away does not. Without the last, the test passes against a
provider that added a password without replacing one.

Not over loopback: pg_hba trusts anything there, and a deliberately wrong
password returned a row for a whole afternoon once.
2026-08-31 02:39:00 +02:00
jschoubben 0100c39845 Build the builder before replacing the hand-started one, and listen from a file
Two setup faults, each of which looked like the thing being tested failing.

The builder module was assigned without its artifact ever being built, so
nothing could start — and the build has to happen while the hand-started
builder is still alive. Same chicken-and-egg as the registry, resolved the same
way: the builder that exists builds the one that replaces it.

The firewall test's listeners were squeezed through three levels of shell
quoting and never started, so the test failed on its own setup — which reads
exactly like the firewall working.
2026-08-31 00:41:24 +02:00
jschoubben 29cdaa4de3 Prove a machine filters what it was told to and nothing else
Written and loaded are different things, and loaded and enforcing are different
again. The test opens two ports on a machine, declares one of them, and checks
from the other machine that the declared one answers and the undeclared one
does not — then removes the module and checks the port closes with nobody
editing a rule.

The base image gains nftables, read back through `nft --version` like the other
three: a machine that cannot load a rule set applies the mesh's filtering,
reports success and filters nothing, which is the exact fault the derivation
exists to remove.

Two earlier tests were asking for things that are not there. The lab's registry
drops tags when it stocks, so `registry:2` is not served and the mirror test
failed with "not found" — it now uses the pinned digest, which is what a
declaration carries anyway.
2026-08-31 00:25:19 +02:00
jschoubben e89379fef8 Prove a mesh credential becomes a login, against a real database
The mesh generates a password, seals it to the machine that must accept
it, and discards the plaintext — so it cannot tell PostgreSQL to start
accepting it. Something on that machine reads what the host wrote and
makes it true. Everything up to that step is proven elsewhere; this is
where a password either becomes a login or does not.

A scenario with one machine and a database, and six assertions: the
password works, running again reaches the same state and says nothing,
rotation makes the new one work and the old one stop, a departed consumer
loses its login, a role nobody here made is left alone, and a manifest
naming a credential that was never written is refused rather than
creating a login with no password.

Each was confirmed to fail — and only it to fail — with the behaviour
removed from the provisioner: only-creates breaks rotation, no-revoke
breaks revocation, revoking everything breaks the bystander role, and
ignoring a missing credential breaks the refusal.

Two faults in the test itself, both worth recording:

- it checked logins from inside the database's own container over
  127.0.0.1, which PostgreSQL's default pg_hba trusts. No password was
  ever verified. Demonstrated directly: over loopback a deliberately
  wrong password still returns a row. Only the rotation assertion
  noticed, because it is the one that requires a password to STOP
  working — which is an argument for writing that assertion every time.
- the fix then read .NetworkSettings.IPAddress, which docker 29 no
  longer populates. It templates to empty, psql falls back to a unix
  socket that is not there, and every login looks impossible rather than
  misconfigured.
2026-08-30 01:29:46 +02:00
jschoubben c79b83c1bc The base image carries the network tools, and a scenario that grows
A sealed scenario cannot install wireguard-tools any more than it can install a
container runtime, so a lab without them cannot test connectivity at all --
which is most of what the mesh does between machines. Installed and not
started: what a node runs is the mesh's decision, and a lab that brought the
interface up itself would be testing its own setup.

growing-mesh exists to be grown. The point is not the third machine, it is that
adding one changes every other node's peer list -- so each has to be told
again, or the newcomer is on a network nobody else can see.
2026-08-29 18:04:17 +02:00
jschoubben 185c414884 A scenario with two machines
The first raises everything from the bundle its host carries and joins the mesh
it made; the second has a host and nothing else, and a person carries it a
token. This is the first scenario where the mesh is a mesh -- everything before
it proved a machine could talk to a control plane on its own loopback, which
proves less than it looks.
2026-08-29 16:51:55 +02:00
jschoubben 88cf89194a The lab said nothing was running while two machines were
'incus list' failed because this shell had no permission to reach the daemon,
incusOk returned null, and the caller wrote ?? "[]". So 'mesh-lab list'
printed 'no scenario instances standing' -- confidently, about a question it
had never managed to ask.

The comment on incusOk warns about exactly this, in those words: absence and
success made indistinguishable. Three of its own callers then did it. Two
listings and the live diagram, which would have drawn an empty scenario rather
than fail -- a picture that is confidently wrong, which is worse than none.

Anything enumerating what exists now goes through enumerate() and throws.
incusOk stays right where failure genuinely means no, like instanceExists,
and there is a test holding that line so this does not get over-corrected
until nothing can be asked at all.

Worth noting 'mesh-lab check' already diagnoses this precise cause, down to
'a session that predates it cannot see it'. The diagnosis existed; the
listing just never asked for it.
2026-08-29 11:54:33 +02:00
jschoubben f88dbcc51e Add the first-node scenario
One machine with PostgreSQL in its registry, for developing the bootstrap.
Used to verify that a sealed machine can raise a store and a database from the
bundle its host carries.
2026-08-29 01:54:18 +02:00
jschoubben be176bab2e Automate the lab registry: a sealed machine pulls by digest
Closes 04-ISSUES/009. A scenario declares `images:` by tag; the lab stocks a
registry on this workstation where there is a network, raises it inside the
scenario as scenery, and reports the references a declaration pins -- which are
the digests THIS registry assigned, and are not knowable until it is raised.

Verified in a sealed machine, confirmed by ping to have no route out: package,
service including boot state, a container pinned by digest, and an action
inside that container. Applied, idempotent on re-apply, and read back from the
machine rather than from the apply's own report. That is the first time the
container shape has worked in the lab at all, and it was the shape blocking the
substrate bootstrap.

Four faults found by running it, three of them mine and one worth keeping:

The read-back checked that the catalog endpoint answered, by looking for the
substring "repositories" -- which `{"repositories":[]}` also contains. So it
passed on a registry holding nothing, and the failure surfaced much later as a
container that could not be pulled. It now asks for each image's manifest BY
DIGEST, which is what a machine does.

A recursive push needs its destination to exist, or incus copies the source's
contents rather than the source. The data landed one directory too shallow and
the registry found nothing where it looks.

The registry writes its blobs as root through a bind mount, so the workstation
could not remove its own scratch directory afterwards. Whoever made the files
removes them -- the cleanup now runs in a container too. And a cleanup failure
no longer fails a raise that succeeded: the scenario is standing and usable,
and saying otherwise would be a false report.

The base image build did not verify that the runtime trusts the documentation
ranges as plain-HTTP registries. Writing the file is not the daemon honouring
it, and a base image that looks right fails much later, in a sealed scenario,
a long way from its cause. It is now read back from `docker info`.
2026-08-29 00:04:55 +02:00
jschoubben 16c13807a9 place: the host — the lab acquires a consumer
The lab raised an underlay and put nothing on it: correct, and useless, because
the thing it exists to test did not exist. Tier 0 now does, so `place: [host]`
works and a raised scenario finally contains something.

The refusal narrows rather than disappearing. A scenario placing a host and a
substrate is told which half is missing, by name — not that `place:` is
unsupported when half of it now works.

Placement reads back rather than assuming. A file arriving is not a host
working, so the binary is run before it is trusted to answer questions, and what
it reports is read from the machine (ADR 0035). The binary comes from an
explicit path, because the declaration design leaves where artifacts come from
open and a search would harden into the answer by accident.

The integration test that matters is the one asserting the host reports the
MACHINE and not the workstation that placed it. A raised VM and this workstation
differ in every capability — root versus uid 1000, a clean init versus a
degraded one, no docker versus docker, no wireguard versus wg0 — so a host
reporting the wrong machine is obvious here and invisible anywhere else.

And the placed host independently confirms ADR 0031: overlay absent on a freshly
raised machine. The underlay suite already asserted that by looking for
wireguard interfaces; this is a second witness rather than the same check twice.

Two tests failed the moment placement worked, which is what they were for. They
defended "there is nothing to place yet" while that was true; the decision
changed, so they change with it rather than being deleted.

Gate: 75 unit, 20 integration.
2026-08-26 00:41:25 +02:00
jschoubben 5d01006eab Transit, host firewalls, and the whole topology raising
The full topology now raises: four machines, three routers, a transit
router, six segments, in 35 seconds. Everything the declaration model can
express except `place`, which is refused because the node host it would
place does not exist yet.

Transit was a real gap, not a bug. The design says public networks are
unrelated and routed to each other, never bridged — and I built the
segments and never built the thing that routes between them, so three
public networks were islands and nothing crossed. A transit router now
holds an interface on every public segment, forwarding and no translation:
the closest thing the lab has to the internet, deliberately dumb.

Proven rather than asserted, by ping TTL across the raised topology:

  within one segment                     ttl=64   no hops
  across two unrelated public networks   ttl=62   gateway + transit
  multicast between public networks      0 replies

A flat internet would have shown ttl=64 and answered multicast — which
would let a node discover a peer it could never reach in production, and
report success. That is the fault the as-is layer records the mesh already
hitting with multicast name resolution.

inbound: deny is implemented as a host firewall on the machine, read back
after applying. A declared refusal that silently did not load leaves the
machine wide open, which looks exactly like a machine that is working.
Established and related traffic is accepted, so a defended machine can
still dial out rather than being a disconnected one.

Verified by running, all of it:

  home -> devices (policy allow)               reachable
  devices -> home (policy deny)                blocked
  behind unforwardable NAT -> out              reachable
  in -> behind unforwardable NAT               unreachable
  inbound: deny, dialling out                  reachable
  reaching a machine that denies inbound       refused

The two routers differ exactly as declared: the forwardable one carries the
policy rule and no inbound drop, the unforwardable one carries `ct state
new drop` and no DNAT.
2026-08-24 01:49:30 +02:00
jschoubben a270cd5b02 Routers: NAT, port forwarding, policy and mapping expiry
A gateway is the one implicit machine in a declaration — a scenario says a
segment sits behind one and never names the thing that serves it. This
materialises it.

A router is a container, not a virtual machine, because it is scenery
rather than something under test (hq ADR 0033). Verified before building
that a plain unprivileged container can do all of it: ip_forward and ipv6
forwarding settable, nftables masquerade accepted, and the conntrack
timeouts mapping_ttl depends on both writable. No privileged mode.

Verified by running, on a machine behind a household gateway reached from
one on a routable address:

  home-server -> anchor                      0% loss, through masquerade
  anchor -> 192.168.1.135 (private, direct)  unreachable
  anchor -> 192.0.2.50:8080 (the GATEWAY)    HTTP 200

The last line is the published-but-behind-NAT case research 004 says only
exists in production. It is now a 32-second scenario on a workstation.

Segments sharing a gateway declaration share ONE router — that is what a
VLAN-capable router is, and two routers sharing an external address would
not work anyway.

mapping_ttl is read back after setting rather than assumed. Those sysctls
are not on every kernel, and a scenario that declared an expiring mapping
and silently got a permanent one would be exactly the fault being built
against.

Four bugs found by running it, three of them the same fault — a failure
made invisible.

The router had no route to a package repository, by design, so installing
nftables at raise time could not work. The image is now built once with
temporary connectivity and cached; every scenario after that needs no
network. That failure was hidden behind `|| true`, which is why it took a
raise to find.

The builder then failed on DNS: exec works before a container has an
address, and I had treated usable as ready. It now waits for the thing
actually needed.

The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time
networking service flushed the static address the scenario set — on eth0
only, so the outside interface came up bare while inside ones were fine.
The image build now neutralises it: a router reconfiguring itself from an
image default is the lab overriding the declaration. `ip addr add … || true`
had hidden this too, and is now `ip addr replace` with no swallow.

And routers were orphaned by destroy, holding their networks open so
destroy reported removing zero segments. They now carry the same machine
tag as everything else, so one query finds an instance's resources.
2026-08-24 01:37:19 +02:00
jschoubben a27d861d3b Scenario lifecycle: raise, exec, snapshot, restore, destroy
A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.

The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.

The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.

Three bugs found by review and by running it, all of one family:

The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.

list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.

restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.

Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.

No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
2026-08-24 01:12:49 +02:00