A consumer contributes a key prefix and gets an ACL user; the test is
that the grant means exactly what the manifest said, in both
directions: its own keys usable, anyone else's refused by the store
itself, and the flush a tenant must never have refused with them.
Waited for through the store rather than through logs: the user list,
asked with the password the host wrote into the server's own conf file
on the machine — nothing invented, both ends reading what the mesh
delivered.
And the forge is asked on the port the mesh assigned, not the one the
module declared. The old curl aimed at 3000, which was right until
ADR 0038 moved the machine side — a latent break that would have fired
on the first run to get past the settling that used to fail first.
The scenario stocks redis and its provisioner, and the rebuild builds
the provisioner image with the others.
whatWasTested read the repositories when the run ended, so a commit
landing during the twenty minutes a suite takes was recorded as tested
without ever being in the binaries. It happened: one receipt named a
commit made mid-run, and the verdict it carried belonged to an older
tree.
The heads are read once, right after the build, and carried to the
receipt. A verdict is only worth something attributed to one exact
state, which is the receipt's whole reason to exist.
A segment named "uplink" is refused. The lab claims that name for the
NAT bridge behind `egress: true`, and a scenario wearing it first would
have its egress machines silently attached to an isolated bridge — a
declared key doing nothing, which is the fault this repo exists to
refuse, in the repo that refuses it.
settled() parses inside the try. A truncated status from a struggling
machine was the one shape of bad answer that still threw out of the
wait, and the likeliest moment for one is exactly the machine the poll
is watching. Malformed now counts as "could not ask", like the exec
that times out.
And a sentence on the uplink's UseDNS saying its inertness is
load-bearing: it matters only where systemd-resolved runs, and on a
machine whose modules own resolv.conf the uplink must not outvote the
resolver a scenario is testing.
Three changes, found by one failing test.
The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.
So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.
`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.
And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
A mesh that has just bootstrapped cannot build the module that gives it
an artifact store: building publishes to the store, and the builder will
not start without one (novox/hq 04-ISSUES/029).
This test built it and passed, because the scenario's registry was
already standing to receive the push — which is precisely why a real
first mesh would have hit this and the lab never did. A stand-in for
Docker Hub was quietly also standing in for the thing under test.
So the module now names its image by digest, the way the bundle names
the three a first node starts from, and is added as a manifest rather
than built. That is the only path open to a real first mesh, so it is
the path this walks.
Its skip on MESH_LAB_BUILDER goes with it. Nothing in the test needs a
builder any more, and a skip that names a thing the test does not use
sends the next person to look in the wrong place.
The canary walked one path on one machine — a mesh comes up, a module
lands, a consumer gets a credential — and stopped the run if it broke.
That path is exactly what the first three tests of the long run walk,
and the long run finishes them about 160 seconds in.
So the gate cost a whole scenario on every passing run to save roughly
45 seconds on a failing one. A scenario is three machines, one of them a
registry that boots a kernel in order to serve files, which is where the
two minutes went.
The test file stays and still runs when it is named. What is gone is
raising it on the way to everything else.
Measured rather than argued: the canary's scenario took 116s of which
60s was standing up a registry, and the run reached the same assertions
without it.
`settled` used `must`, so a failed exec ended the wait as though the
machine had reported a failure. It had reported nothing: the control
plane is a container on the node being polled, and while that node
applies a declaration an exec into it can lose its stdout fifo to
containerd. The run then blamed the mesh for a question that missed.
Could not ask and asked, and the answer was bad are different facts, and
only the second is the machine's. A failed poll now keeps the reason and
tries again; the timeout reports whichever came last, so a control plane
that is genuinely unreachable still fails the test — with the reason
rather than with a stack trace.
Every five seconds rather than every two. Each poll is an exec into a
container on a machine that is busy applying, and thirty times a minute
was competing with the apply rather than observing it.
The forge test read `status` the instant `push` returned and concluded
the machine was fine. It was describing the apply before this one.
`push` sends and returns — it prints "sent N resource(s)" and the
machine applies afterwards. So every assertion made immediately after
one is racing it, and this race lost quietly: no failure reported, and
a container that did not exist yet read as a container that would never
exist.
`settled` asks the mesh, in its own terms: a node is caught up when it
is neither waiting for what it was sent nor wrong about what it applied
— the two questions `status` already answers, read as JSON so a test is
not parsing a report written for a person. A machine reporting a failure
ends the wait immediately rather than at the timeout, because it will
not become right by being waited for.
The container assertion now also prints the plan. A container missing
because the mesh never asked for it and one missing because the machine
could not make it are one sentence and two entirely different faults,
and the plan is what separates them.
The forge test pushed and then waited for a database login. When the
containers were never created at all, it reported "no login was created"
— true, and silent about why. Two hundred and thirty seconds spent
proving something downstream of the actual failure.
A push being accepted and an apply having worked are different facts,
and this test depends on the second. It now reads what the machine says
about itself, and whether a container exists, before it starts waiting —
and carries the host's own log into the failure either way.
The suite otherwise passed 24 of 25 on this run, which is the first time
the forge reached a clean attempt with nothing upstream blocking it.
A scenario is a closed address space: two raised from the same
declaration hold the same addresses and never meet, which is what lets
two run at once and why the lab talks to machines through the
hypervisor rather than over IP. Reaching in from outside breaks that, so
it is opt-in, one scenario at a time, and reversible.
`connect` takes an address on the scenario's public link and writes a
resolver rule answering everything under each machine's name.
`disconnect` gives both back. `connected` says what is true right now,
for somebody who cannot remember.
It refuses rather than guessing when more than one scenario is standing
— the failure being avoided is not an error but one scenario's traffic
arriving in another. It also refuses when a machine's name is already
answered here for something real, because connecting would point that
name at the lab, and the damage would land on the real thing.
Names answer with the segment address rather than the overlay one.
Inside the mesh a name gives a machine's private address; from here that
would need this workstation on the overlay, which is a much larger door.
The segment address reaches the same machine and the same ports, which
is what opening a board in a browser actually needs.
Proven against a live two-node scenario: registry.internal:5000/v2/
answered 200 from this workstation, and so did a wildcard name under the
same machine. Disconnect put the address back, stopped answering, and
left the real mesh's own names alone.
One thing measured rather than assumed: it restarts dnsmasq instead of
reloading it. A reload is SIGHUP, which re-reads the hosts file and
clears the cache but not the configuration — the rule was written, the
reload reported success, and nothing resolved. The daemon's start time
was nine days old afterwards.
Suggested by Jochen, and it paid for itself on its first run.
A suite that takes forty minutes is a suite you hear from once a day.
Every fault found today would have shown up in the first three minutes
of it — a module pinned to an image that does not exist, a consumer
given a password and no name to present with it, a credential file
nothing could read, a search for a password that read the password as an
option. The other thirty-seven minutes proved things that were already
working.
So this runs first, on one machine, with the three images the mesh needs
for itself. It walks one path: a mesh comes up, a module lands, and a
consumer gets a credential it can actually use — the name to present,
the address, the port, and a password only the host could put there.
Deliberately not a smaller copy of the full suite: that path is where
everything went wrong, and a canary checking many things shallowly is a
canary whose failure nobody can read.
`suite` runs it and stops if it dies, saying why rather than leaving
somebody to wonder what the missing thirty-seven minutes would have
said. Skipped when the caller named its own files.
It measured 164 seconds against forty-odd minutes, and failed three
times on its first run for one reason: applying the bundle raises a
control plane but does not tell it a machine exists. I had left out
enrolment, and the long suite would have taken forty minutes to say so.
**A password beginning with a dash broke the search for it.** The
credential test greps the machine's own files for the delivered
password; this run's password started `-S`, so grep read it as an option
and refused the whole invocation. The test compared the usage message
against "0" and reported the password as leaked. That is the worst way
for a search to fail — it says it found something. Fixed with `-e` and
`--`, which is what those exist for.
**The planning test could not redirect what the scenario does not
serve.** Rewriting an image reference only works for repositories the
scenario's registry actually holds, and the object store's provisioner
was not stocked — so that module kept its placeholder and the refusal
fired, correctly. It is stocked now, so all five are planned again. The
skip path stays for anything genuinely unserved, and says which module
and why: a planning test quietly covering four instead of five is the
false coverage this suite exists to prevent.
**The forge failed because of the one above it.** The planning test
threw before its cleanup could run, leaving a module assigned that
refused the next push, so no database container was ever created.
Yesterday's fix moved that cleanup where a failure cannot skip it — but
`after` still only unassigns what was assigned, so it now tracks what
actually got added rather than what was intended.
Two faults, both found by the guard that now refuses a placeholder
digest on its way to a machine.
The planning test added the manifests exactly as they sit on disk, which
includes an image the mesh builds — and that image has no digest until
it is built, so the file legitimately carries a placeholder. The test
was therefore planning something that could never run, which is the
whole complaint. It now points the references at this scenario's
registry first, exactly as the forge test does.
The second is worse and more ordinary. Its cleanup was the last
statement in the test body, so the first failure skipped it and left
five modules assigned. The next test's push was then refused by a module
this one had abandoned — a failure that reads as a fault in the test
that was working. Cleanup that only runs on success is not cleanup, so
it moved to `after`, where a failure cannot skip it.
Everything up to now stopped at composing a declaration. That proves the
control plane and the host agree, and proves nothing about whether the
thing described works — which is how five modules sat pinned to images
that did not exist while parsing and resolving perfectly.
The forge is the right one to run first. It needs a database from
another module, a password it did not choose, and a connection string it
could not have written itself: the address and port come from what the
database serves, the user name from what the mesh decided both ends
would call it. If any of that is wrong it cannot start, and nothing else
in this suite would notice.
The test checks the chain in the order it has to happen — the login
exists, the database it owns exists, the forge answers, and its log does
not say authentication failed. That last one matters: a forge that
started and could not reach its database would still answer on its port.
Rewriting image references is now shared rather than copied from the
bundle, which had the same problem first. A digest is not knowable until
something is built, and when it is, it belongs to whichever registry
served it — so the text says which image and the scenario says which
copy. Matching is on the repository, with a test that a repository
ending in another one is not half-replaced.
Also makes the planning test put the machine back. Tests here share one
mesh, so the five modules it assigned were inherited by whatever ran
next; harmless while nothing pushed, and not harmless now.
novox/hq 04-ISSUES/024. A run stalled for thirty-five minutes and said
nothing. The cause was a link systemd was still configuring, three
layers down inside a `docker load` blocked on a socket — and every one
of those layers knew what it was waiting for. None of them said so.
Three decisions, each doing work.
**Every external command is logged, at the three places that run one.**
Ninety-seven call sites reach a hypervisor or a container runtime
through three wrappers, so instrumenting the wrappers covers all of them
and nothing has to remember to log.
**A command still running says so while it runs.** A line before and a
line after tells you nothing until the after arrives, which is exactly
the case that matters. Anything outstanding past fifteen seconds reports
itself with how long it has been going. It is reported as still running,
not as stuck — which it is is not knowable from there, and a log that
calls a slow step a hang teaches people to ignore it.
**It goes to a file, written synchronously.** Node block-buffers stdout
when redirected and a test runner buffers it again, so a console log can
sit minutes behind. `appendFileSync` cannot lag.
Two things this found in itself while being written, both the same shape
as what it exists to catch:
A question that answers no is not a fault. Half the lab's commands are
questions — does this network exist, is the agent up yet — and they fail
constantly while a scenario comes up. Logging those as faults filled a
healthy run with ✗, which is how you end up ignoring ✗ when one is real.
They are recorded quietly now, and still recorded.
And `around` skipped its own wrapper when a step's level was below the
configured one — taking the failure line and the heartbeat with it. The
two things worth having at a low level were the two that vanished at
exactly the level somebody would use. The gate belongs in `write`.
Also unsilences the four call sites that passed a callback throwing
everything away, including the one the stall sat in, and tees `raise`'s
progress into the file whether or not a caller asked to see it — the
end-to-end test passed no callback, so the one run that mattered
reported not a single step.
The last failure of the run: `MESH_BROKER_FILE=/var/lib/mesh/builder/broker`
reported as "something secret-shaped, which the broker would see".
`/` is in the base64 alphabet, so any absolute path of 24 characters or
more matched the pattern meant to catch a sealed value. An absolute path
is a *reference* to a secret and naming one is the whole design — the
mesh delivers a credential as a file and a module says where.
Excluded explicitly rather than by loosening the pattern, and checked
both ways: a real sealed value and a base64 blob are still flagged, a
relative path still is, only an absolute path is passed over.
Worth the words in the comment. A check that fires on the right shape
for the wrong reason is worse than none — it is the one that gets
suppressed, and then it is not there when it is right.
It surfaced now because tests in this file share one mesh: the builder
was assigned by an earlier test and appears in this one's declaration.
Node block-buffers stdout to a file, so a run log can sit unchanged for
minutes while the run is fine. Read that way twice today — the second
time straight after fixing a real stall, which is the worst version of
it, because a buffering artifact then reads as the fix having failed.
The machines are the source of truth and answer immediately. Written
down with the two commands that settle it.
novox/hq 04-ISSUES/024. The registry machine had its address set with
`ip addr add`; every other machine gets a systemd-networkd unit. That
one difference stalled the lab indefinitely.
An address set by hand leaves networkd waiting to configure a link it
was never told about, so the link sits at `configuring` for ever.
`systemd-networkd-wait-online` has TimeoutStartUSec=infinity, so
`network-online.target` is never reached — and Docker is ordered after
it. `docker load` then blocked on a socket whose daemon was queued
behind a target that would never come.
Measured before and after on the same scenario: stuck with five pending
systemd jobs and `docker` inactive; now `enp5s0 configured`, `docker`
active, no jobs, and the whole raise completes in 87.5s.
The guess in the issue was wrong, and it was wrong in the usual way —
stocking had just been changed, so stocking looked guilty. Stocking
takes 34s and always did.
Two things that made this cost hours rather than minutes are fixed with
it. Placing an image now waits for the container runtime to answer and
refuses after 120s naming what systemd is waiting on, so a stall becomes
a failure that says why instead of three stacked timeouts totalling 35
minutes. And the end-to-end test passes `onProgress`, so a raise says
what step it is on — it printed nothing at all until it finished, which
is why 35 minutes of nothing read as a slow test.
Pointed at the repository root, the binary it writes is 12 MB of tracked
artifact. Said in the example rather than left to be discovered by a
On branch initialization
Your branch is ahead of 'origin/initialization' by 43 commits.
(use "git push" to publish your local commits)
Changes to be committed:
(use "git restore --staged <file>..." to unstage)
modified: README.md that looks wrong.
Nothing did. The sealed-placeholder substitution and the bound-value
substitution were each covered by unit tests in the repository that
performs them, and the two expressions that find the holes live in
different repositories — so both sides could agree with themselves and
disagree with each other, and the first thing to notice would be a
program connecting to a host called "${bound:postgres-database:at}".
So the consumer in the credential test now ships a configuration file
with four holes in it: the address and port from what the provider
serves, the name to present from what the mesh decided, and the password
sealed. The mesh fills the first three before sending, the host opens
the credential and fills the last on the machine, and the test reads the
file off the machine and checks that the password in it is the same one
the credential file holds — and that no ${ survived.
This is the only place those two mechanisms meet a real host.
The credential moved: the sealed password is a password alone, at
`.secret`, and `database.env` is now the connection keycloak could not
have written — address and port from what the provider serves, user name
from what the mesh decided both ends would call this consumer.
So the test asks for both, and for the seam between them: the password
is still a hole, the sealed value travels beside the file that needs it,
and no ${bound:...} survives as a value. That last one matters most —
a placeholder written through would be read as a hostname, and the
failure would name neither the module nor the mesh.
A run today had a control-plane image built that minute and a
provisioner image built the day before. The rotation test failed against
a real database and it looked exactly like the change under test being
wrong — the provisioner was creating logins by a naming rule that had
been replaced hours earlier.
It was the rebuild. It covered `make image` and the builder binary and
none of the three other image targets, all of which the suite runs.
This is the same fault the builder line was added for, one target along,
and the comment there already names the precedent: building one and not
the other is the eleven-hour-old binary. A rebuild that covers most of
what a run uses is worse than one that covers none, because the run that
follows it is believed.
The test names each target rather than counting them, because what goes
wrong is a target that exists and is not run, and a count would not
notice.
Two assertions in the full-mesh test encoded the old naming: the grant
file read back from the provider, and the PostgreSQL role the real
application logs in as. Both are named after the consumer now, and a
consumer is a module on a machine.
These are the two that matter most in this file — it is the only place
where a real application authenticates against a real database with a
password the mesh delivered and cannot read, so they are what would have
caught the naming going wrong end to end.
novox/hq 04-ISSUES/022: a consumer is a module on a machine, not a
machine. The mesh now writes <node>.<module>.secret and the provisioners
name the role and the access key after both.
The fixtures here write what the mesh writes, so they move with it —
that is the whole point of them, and a fixture that kept the old shape
would agree with the bug rather than catch it.
The object-store assertions are the ones that mattered most: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property. One access key per machine meant every module on a node
shared it, and the policy confining each consumer to its own bucket
confined none of them.
Reconstructed from the source twice now, which is 04-ISSUES/005 in its
own README: a test whose artifact was not pointed at skips rather than
fails, so an unset variable is a green run that proved nothing. The
first attempt today reported "skipped 24" and left a receipt claiming
zero of everything — working exactly as designed, and indistinguishable
at a glance from a suite that had nothing to do.
Also records the two things that cost time either side of it: `check`
says which variables are missing before a long run rather than skipping
quietly, and a heredoc into `newgrp` runs the suite as a child of a
shell that immediately exits, so it needs `setsid nohup … &` or it dies
with the shell that launched it.
**The adoption.** Software nobody here wrote, taking its credentials the
way such software does — from its environment — and needing two
containers that reach each other by name. The first module that could
not have been declared this morning: it needs the network shape and it
needs a sealed value to reach a container's environment.
Its password is accepted rather than generated, which is the whole shape
of an adoption: a service that already exists keeps the credential it
already has. Asserted properly — a wrong password is refused by the same
database, so the passing case means something.
**The warm scenario.** A mesh kept between runs and returned to, which
turned twelve minutes of bootstrap into thirty seconds of restore. Off
unless asked for: a run that is meant to mean something raises from
nothing.
Its guard fired for real during this work, unprompted — a mesh-host
commit landed and it refused the stale base, naming both commits, rather
than passing tests against yesterday's binary. That is 04-ISSUES/005's
rule one level down.
Three things the guard learned the hard way and now handles: a snapshot
captures disk and not memory, so the host is restarted after a restore
and asserted to have come back; the stocked image digests are worked out
while raising and a restored instance never raises, so they are kept;
and comparing only the repositories this run can see clears the ones it
cannot, so both directions are compared.
The one real bug behind five failed attempts was in mesh-host and it
reported itself precisely: a network shape the language had and no host
implemented. Everything else was scaffolding of mine.
The other half of the certificate split: the mesh's own authority
certifies internal names, and a name reachable from outside needs one
the world already trusts. Against a real server rather than a stub,
because what is under test is whether an order, a challenge and a
handshake agree, and a stub would be told to agree.
One assertion passes and one fails, and the failure is filed as
novox/hq 04-ISSUES/020: the authority issues a certificate and the
client never collects it. Kept as a failing test rather than deleted or
skipped — it is the reproduction, and it proves everything up to the
last hop.
The passing one is the guard that matters day to day: no certificate is
ordered for a name the mesh does not route, so a scan cannot spend an
account's rate limit.
The failure output gathers both sides before asserting. The first
version reported only what the proxy said, which made a server-side
question unanswerable — "the client never spoke to it" and "it refused
what the client said" are different faults with nothing in common.
Seven assertions against a real store, the important one being that a
consumer cannot reach another consumer's bucket — isolation here is a
policy somebody wrote rather than a boundary the product has.
The revocation test stages its own precondition. The first version
asserted a key left by an earlier test, and the rotation test had
already revoked it two tests early: the behaviour was correct and the
test was measuring residue. Its precondition assertion is what caught
that, rather than it passing green having verified nothing.
The danger is not that the suite breaks. It is that nobody notices it
stopped running (novox/hq 04-ISSUES/005). The harness this replaces had
not built for two and a half months and nothing said so — and this suite
needs a hypervisor, so it inherits exactly that: it runs when somebody
remembers, and remembering is not a mechanism.
So running, recording, and rebuilding are one act:
- the host binary, control-plane image and builder are rebuilt from
source first. The last two both parse manifests; building one and not
the other left a binary eleven hours old refusing a field the mesh had
just renamed, found by a full run.
- a receipt lands in XDG state — outside git, because the question is
whether *this machine* has run it, and a receipt in git would be a
claim about everybody's machine made by whoever committed last.
- `last-run` judges it and exits non-zero when it no longer counts.
Three faults found by running the thing rather than reading it, each now
held by a test confirmed to fail without it:
- counted() passed every test while parsing nothing. The runner colours
its summary even into a pipe; the fixtures were clean text that had
been imagined rather than captured. A fixture that agrees with the
mistake proves the mistake.
- a receipt for `suite test/lastrun.test.ts` was indistinguishable from
one for the real thing — 005's own symptom, rebuilt inside its remedy.
The receipt now records what ran.
- a tree with uncommitted work reported the bare commit, claiming
coverage of code nobody can check out. Nothing else could tell: the
hash is identical either way.
Proven on real machines: 22/22, against all three repositories.
Three times a diagnostic has not run because the thing before it threw: `must`
on a command that exits non-zero, and then a query that hung long enough to
take the harness's own timeout with it — which arrives as an error with no
evidence attached rather than as a failed assertion.
A thirty-second test costs fifteen minutes to re-run, so the evidence has to be
gathered whichever way it fails. One helper, used by every assertion here, and
the query is bounded on the machine rather than by the harness: a query that
hangs is a result, not an accident.
These machines run systemd-resolved, so resolv-conf would fight it over the
file. Both claim the-resolver-configuration so that assigning the wrong one is
refused rather than fought over — and the test was picking the wrong one.
dnsmasq cannot bind 127.0.0.54 and the module's own comment claims that address
is free. Which of those is wrong is the question, and naming the holder answers
it — assuming would be issue 012's mistake, where two things changed and the
plausible one was blamed.
`systemctl is-active` exits non-zero for a unit that failed, so `must` threw
before the assertion carrying every diagnostic — and the run said only
"failed". The journal, the config, what the mesh wrote and resolv.conf are all
things the next run should not have to be re-run to see.
Through the path an application actually takes — nsswitch, files, then DNS —
because the resolv.conf module is half of what is being tested and only that
path goes through it. Asking a server directly would prove less.
Both machines resolve, from their own copy: a mesh where one machine answers
for all of them stops resolving when that machine does, which is the
arrangement this design refuses everywhere else.
The manifests are read from mesh-control's examples rather than written here,
so what is proven is what ships. And dnsmasq joins the base image, read back
through --version like the others: a machine that cannot answer names applies
the resolver data, reports success, and resolves nothing.
mesh-resolver requires name resolution, which requires the network — so
unassigning the domain module alone leaves the machine on the network, pulled
back by its own requirement. The mesh was right and the test was wrong.
Which is the requirement graph doing its job: a module cannot quietly lose
something it depends on because somebody removed the thing that first brought
it in.
The mesh's half: the data is right, complete on every machine, and agrees with
the hosts file — two accounts of where a machine is, disagreeing, would be
worse than either alone, and this is the one place they could drift because
they are generated separately.
And it follows the machines: a node that leaves the private network stops being
answered for, because a wildcard pointing at nothing resolves and then hangs,
where an unresolvable name fails at once and says which name it was.
The mesh gives its names to the containers it declares. This one is started by
the test with `docker run` — nothing declared it, so nothing configured it, and
removing the workaround here was claiming a reach the change does not have.
The boundary is the right one: a container somebody runs by hand is not the
mesh's to configure. Reaching into every container on a machine, declared or
not, is what a resolver in resolv.conf would be for — and that remains the
case for wanting one.
The proof that names work inside containers is its own test, against a
container the mesh declared, and it passes.
The rotation test resolved the address on the machine and passed it in, because
the name failed inside the container. That workaround is gone, and a test that
checks the name from inside a container is added — on the machine it has always
worked, which is what made this easy to miss.
`networking` is the requirement a module offers; `mesh-wireguard` is the
module, and what caused a rule is the module itself. The rule set was right and
the assertion was looking for the wrong name.
The failure guarded against is not subtle and is very hard to recover from: a
rule set that closes the hub's own port takes the private network down, and the
mesh's way of fixing anything is to send a declaration over it.
So the assertion that matters is not the rule file — it is that a declaration
still reaches the other machine afterwards, and that the other machine still
reaches the hub. A rule file that looks right and a mesh that has stopped are
exactly what this is for.
A file in a missing directory is not impossible: the host creates the parents,
which is correct and meant the first version of this test broke nothing at all
— and then reported that the board could not name a failure that never
happened. A unit that does not exist fails immediately and in the host's own
words, which is also what the board is being asked to show.
A machine is given a declaration it cannot apply, and the board names it,
says 'failed' rather than 'error', and shows the host's own words about what it
could not do — a board that said only 'failed' would send a person to ask the
thing they opened the board to avoid asking.
And it agrees with the command, from the same read: two answers to 'which
machine is broken' would be worse than either alone. Reading it changes
nothing.
This assertion passed twice and failed once on nothing but timing, which is the
worst kind of green: it says the mechanism works when what it measured was the
clock.
`status` cannot stand in for the wait either. A machine that has not applied
yet is not a machine that failed — "not yet" and "never" look identical there,
and only one of them is worth failing over. So it waits for the thing itself,
and says on failure that the machine did apply, which is what separates "the
mesh still thinks this contributes" from "nothing was sent".
The mesh withdrawing a route and the machine acting on it are different things,
and a test that reads the file without checking the second reports the first
wrongly whenever the machine is behind for any unrelated reason. On failure it
now also prints what the mesh would send now, which is what separates 'the mesh
still thinks this contributes' from 'the machine never applied'.
A container does not inherit its host's /etc/hosts, so a name the mesh wrote
there resolves for the machine and not for anything it runs. It fails as "could
not translate host name", which reads like a mesh that never wrote the name —
so the address is resolved on the machine and the container is given that.
And `grep -c` prints 0 and exits non-zero when it finds nothing, so the obvious
`|| echo 0` prints a second one and the count is never what it looks like. The
question was always whether, not how many.
The rotation check now also reports what psql said, not only what the
provisioner said: the failure was on the client side and the diagnostics were
all from the server.
Two setup faults, each of which read as the mesh failing.
The provisioner names a role after the machine and a database after what the
module asked for. The rotation test logged in as the module's name into the
wrong database, so a provisioner that had done its job exactly looked like one
that had not.
And the licence test assigned before recording any licence, so the refusal it
got was "nothing provides model-access" — correct, and a different refusal from
the one being tested. Assigning is also the earliest point a person meets it,
so that is where it is now checked.
Commit, build, catalogue, push — and the machine ends up running what the
source says. With the two halves that make the answer trustworthy: it is still
running the old one until it is told, because the mesh changing its mind is not
a machine acting on it; and it stops being reported as behind once it has
caught up, because a status that says "behind" for ever is one nobody reads.
Refused until somebody says which, with both candidates and the command named.
Refused again while it has no key. Then the key is given on standard input, the
public half arrives saying it came from a record rather than a machine, the key
itself arrives readable only by that machine — and it is nowhere in the control
plane's own database, nor in what crossed the broker.
And the rotation test asked for grants without receives, so the provisioner
found a directory of unexplained secrets and said nothing had been granted:
true, and indistinguishable from a credential never delivered.
The request goes to the name, across the private network, and returns the
workload's own answer. Then the module is unassigned and the same request must
stop working — a stale public name pointing at nothing fails more visibly than
a stale grant.
The workload declares its port as well as its route, because they are different
questions and the earlier test leaves this machine filtering: a module that
asked for a route and not for the port would be unreachable by the proxy it
just asked for.
Two ends holding a matching string proves they agree, not that either is right.
So the check is three logins over the private network from the consumer's own
machine: the delivered credential works, the rotated one works, and the one
that was rotated away does not. Without the last, the test passes against a
provider that added a password without replacing one.
Not over loopback: pg_hba trusts anything there, and a deliberately wrong
password returned a row for a whole afternoon once.
It runs in a container, so a path like /root exists for the machine and not for
it. The first build in this test works because the hand-started builder runs on
the host; the second is done by the module, and asked it to clone a path it has
no way to reach.
A real module is cloned from the forge over a URL. The lab has no forge, so the
repository goes in the directory the module already mounts — the same fact
wearing different clothes.
The distribution's nftables.service is Type=oneshot with no RemainAfterExit: it
loads the rules and goes inactive. A host asked for a service that is "running"
then reports, quite correctly, that it is stopped — every packet filtered as
declared, and the machine marked as not doing what it was told.
There is no state in the vocabulary for "ran and exited having done its job",
so a module that needs one brings a unit that stays. That is also the right
shape: how a machine enforces rules is a fact about the machine, and the mesh
has no business depending on what a distribution happens to package.
And when the builder's build times out, dump the builder's own account of
itself. "Nothing consumed the queue" names no cause and is the same sentence
whether the credential was refused, the queue was never declared, or the
process died three seconds in.
Both assertions were wrong and the mesh was right, which the output made
plain: the rule set named its source, dropped by default, restricted the
declared port and omitted the undeclared one.
"From the mesh" resolves to the addresses on the private network — the whole
point — and the assertion was looking for the segment the two machines happen
to share. So the test now reaches the same machine both ways, and asserts the
declared port answers over the private network and does NOT answer off it.
A test with only one path could not tell "open to the mesh" from "open".
And the fingerprint is delivered with a sha256: prefix, which the regex did not
allow.
A builder that cannot connect sits there, and every outward sign — container
up, credential on disk — says it is working. The failure surfaced five minutes
later as nothing consuming the build queue, which names no cause at all.
Two faults in one line of the harness, and the second is the serious one.
Every command was wrapped as `<cmd> 2>&1; echo "__exit=$?"` on a single line,
so any command containing a heredoc broke: the terminator line became
`MARKER 2>&1; echo ...`, matched nothing, and the heredoc swallowed the rest of
the script — the echo with it. `exec 2>&1` on its own first line fixes that: a
heredoc then terminates where it says it does.
And when the marker was gone, `Number("")` is 0, so the missing exit status
read as exit 0. A command whose output was swallowed reported that it worked,
which is the one answer a test harness must never give. It is now a failure,
with whatever was said returned so the reason is visible.
Found because the firewall test's listener is written with a heredoc and never
started, and the test failed on its own setup — which reads exactly like the
firewall working.
Two setup faults, each of which looked like the thing being tested failing.
The builder module was assigned without its artifact ever being built, so
nothing could start — and the build has to happen while the hand-started
builder is still alive. Same chicken-and-egg as the registry, resolved the same
way: the builder that exists builds the one that replaces it.
The firewall test's listeners were squeezed through three levels of shell
quoting and never started, so the test failed on its own setup — which reads
exactly like the firewall working.
Written and loaded are different things, and loaded and enforcing are different
again. The test opens two ports on a machine, declares one of them, and checks
from the other machine that the declared one answers and the undeclared one
does not — then removes the module and checks the port closes with nobody
editing a rule.
The base image gains nftables, read back through `nft --version` like the other
three: a machine that cannot load a rule set applies the mesh's filtering,
reports success and filters nothing, which is the exact fault the derivation
exists to remove.
Two earlier tests were asking for things that are not there. The lab's registry
drops tags when it stocks, so `registry:2` is not served and the mirror test
failed with "not found" — it now uses the pinned digest, which is what a
declaration carries anyway.
went
`exec` waited two minutes always. A build, or anything that waits on
another machine, needs longer — and a caller that cannot say so has to
split the work to fit, which is a test shaped by its harness rather than
by what it is testing.
The scenario also places a build machine when one is given, so anything
in it can ask the mesh to build something. Nothing else here would start
one.
And the mesh-runs-its-own-artifact-store test is not here. It needs a
fourth image so the module has a registry to mirror, and that is caught
behind 04-ISSUES/012 — left as a note saying where it went and why,
rather than silently deleted, because what it asserted is worth
asserting.
Broken with a package that does not exist, so the failure is real and
fixable. The mesh reports it failed; the resources that could be applied
were, because one broken thing no longer blocks the rest; `push --behind`
names that machine and not the one that is fine; the module is corrected;
and the machine recovers with nobody naming it.
And with nothing behind, it says so rather than doing nothing quietly.
Two properties the design claims and neither had been run.
A push to a machine that is switched off must not be lost — a machine is
disconnected as an ordinary situation, not an exception. The queue is
durable and the message persistent, which ought to be enough, but a lost
declaration is silent and "ought to be" is not a property. It waits: the
machine's host is stopped, the push happens, nothing changes on the
machine, and when it listens again it applies what it missed with no
second push and nobody saying anything.
Getting there found a real fault, now fixed in mesh-host and recorded as
04-ISSUES/011: the machine stopped at the first failing resource, so one
broken module blocked every module after it for ever. The evidence was
the broker's queues being EMPTY — the declaration had been delivered and
read.
And removal: two modules assigned, one unassigned, and the machine loses
exactly that one's file while keeping the other's — and keeps the store,
broker and control plane it raised from its own bundle, which the mesh
never declared and must never remove.
Two of my own traps recorded in the test, because both cost real time:
`pkill -f` matches the shell running it, which kills the connection
carrying the command and hangs the caller for ever; and a test that
depends on state another test left behind fails for a reason that has
nothing to do with what it claims.
The status path was demonstrated with inserted rows, which proves the
query and not the path. This sends a real machine something it will
genuinely fail at — a package that does not exist — and asks the mesh
afterwards.
A failure of the ordinary kind: the host tries, the package manager says
no, some of the declaration is applied and some is not. That is the
situation `status` exists to distinguish from a machine that refused
everything, and the test asserts the distinction survives the whole way:
the machine is listed as failed rather than refused, the failing resource
is named in the host's own words, and the machine that did as it was told
is not implicated.
The chain this closes: a repository exists, the mesh asks for it, a build
machine takes the work, publishes what it made, and the catalogue then
says what the module is, which commit it came from, and — after the
source moves — that it is behind.
Three assertions, against a real broker and registry, because what is
under test is four processes agreeing over a wire:
- the mesh asks, a machine builds, and the artifact is really in the
registry at the digest the manifest names
- a build that cannot succeed says why and records nothing. A failure
that is silent is indistinguishable from a builder that is not running
- the source moving makes the catalogue say "behind", and rebuilding
catches it up
git is now in the base image, with the same reasoning as docker and
wireguard-tools: a machine that builds modules clones them, and a sealed
scenario cannot install anything. Read back from `git --version` rather
than from the package manager — an installed package is not a
capability, and a build machine whose clone fails does so three minutes
into a scenario with the failure reported as a build problem rather than
a lab one.
Everything before this proved a part. This proves the parts meet, which
the project keeps saying cannot be checked any other way.
A bare machine applies the substrate bundle and becomes a mesh — store,
schemas, broker with a certificate it generated itself, control plane
serving. Both machines then join it with nothing but a token. A database
is declared on one and an application on the other, and after a push:
- both ends hold the SAME password, or nothing could authenticate
- it is mode 0600 on the machine that uses it
- it appears in neither machine's stored declaration, neither machine's
reported state, nor the control plane's database — so it was not
readable by the broker that carried it or the mesh that sent it
- the consumer is also told where its database is, by a name the mesh
wrote into that machine's hosts file
The bundle's image references are rewritten to the ones this scenario's
registry serves. A digest belongs to whatever registry serves it, so a
committed bundle names a registry that is not this one — rewriting is
what makes it applicable rather than a placeholder to tidy away.
Two faults found getting here, both fixed in mesh-host: `apply` could not
read a file the bundle could, and the token did not say what the mesh
calls the machine.
The mesh generates a password, seals it to the machine that must accept
it, and discards the plaintext — so it cannot tell PostgreSQL to start
accepting it. Something on that machine reads what the host wrote and
makes it true. Everything up to that step is proven elsewhere; this is
where a password either becomes a login or does not.
A scenario with one machine and a database, and six assertions: the
password works, running again reaches the same state and says nothing,
rotation makes the new one work and the old one stop, a departed consumer
loses its login, a role nobody here made is left alone, and a manifest
naming a credential that was never written is refused rather than
creating a login with no password.
Each was confirmed to fail — and only it to fail — with the behaviour
removed from the provisioner: only-creates breaks rotation, no-revoke
breaks revocation, revoking everything breaks the bystander role, and
ignoring a missing credential breaks the refusal.
Two faults in the test itself, both worth recording:
- it checked logins from inside the database's own container over
127.0.0.1, which PostgreSQL's default pg_hba trusts. No password was
ever verified. Demonstrated directly: over loopback a deliberately
wrong password still returns a row. Only the rotation assertion
noticed, because it is the one that requires a password to STOP
working — which is an argument for writing that assertion every time.
- the fix then read .NetworkSettings.IPAddress, which docker 29 no
longer populates. It templates to empty, psql falls back to a unix
socket that is not there, and every login looks impossible rather than
misconfigured.
A sealed scenario cannot install wireguard-tools any more than it can install a
container runtime, so a lab without them cannot test connectivity at all --
which is most of what the mesh does between machines. Installed and not
started: what a node runs is the mesh's decision, and a lab that brought the
interface up itself would be testing its own setup.
growing-mesh exists to be grown. The point is not the third machine, it is that
adding one changes every other node's peer list -- so each has to be told
again, or the newcomer is on a network nobody else can see.
The first raises everything from the bundle its host carries and joins the mesh
it made; the second has a host and nothing else, and a person carries it a
token. This is the first scenario where the mesh is a mesh -- everything before
it proved a machine could talk to a control plane on its own loopback, which
proves less than it looks.
'incus list' failed because this shell had no permission to reach the daemon,
incusOk returned null, and the caller wrote ?? "[]". So 'mesh-lab list'
printed 'no scenario instances standing' -- confidently, about a question it
had never managed to ask.
The comment on incusOk warns about exactly this, in those words: absence and
success made indistinguishable. Three of its own callers then did it. Two
listings and the live diagram, which would have drawn an empty scenario rather
than fail -- a picture that is confidently wrong, which is worse than none.
Anything enumerating what exists now goes through enumerate() and throws.
incusOk stays right where failure genuinely means no, like instanceExists,
and there is a test holding that line so this does not get over-corrected
until nothing can be asked at all.
Worth noting 'mesh-lab check' already diagnoses this precise cause, down to
'a session that predates it cannot see it'. The diagnosis existed; the
listing just never asked for it.
One machine with PostgreSQL in its registry, for developing the bootstrap.
Used to verify that a sealed machine can raise a store and a database from the
bundle its host carries.
Closes 04-ISSUES/009. A scenario declares `images:` by tag; the lab stocks a
registry on this workstation where there is a network, raises it inside the
scenario as scenery, and reports the references a declaration pins -- which are
the digests THIS registry assigned, and are not knowable until it is raised.
Verified in a sealed machine, confirmed by ping to have no route out: package,
service including boot state, a container pinned by digest, and an action
inside that container. Applied, idempotent on re-apply, and read back from the
machine rather than from the apply's own report. That is the first time the
container shape has worked in the lab at all, and it was the shape blocking the
substrate bootstrap.
Four faults found by running it, three of them mine and one worth keeping:
The read-back checked that the catalog endpoint answered, by looking for the
substring "repositories" -- which `{"repositories":[]}` also contains. So it
passed on a registry holding nothing, and the failure surfaced much later as a
container that could not be pulled. It now asks for each image's manifest BY
DIGEST, which is what a machine does.
A recursive push needs its destination to exist, or incus copies the source's
contents rather than the source. The data landed one directory too shallow and
the registry found nothing where it looks.
The registry writes its blobs as root through a bind mount, so the workstation
could not remove its own scratch directory afterwards. Whoever made the files
removes them -- the cleanup now runs in a container too. And a cleanup failure
no longer fails a raise that succeeded: the scenario is standing and usable,
and saying otherwise would be a false report.
The base image build did not verify that the runtime trusts the documentation
ranges as plain-HTTP registries. Writing the file is not the daemon honouring
it, and a base image that looks right fails much later, in a sealed scenario,
a long way from its cause. It is now read back from `docker info`.
Comments naming records that no longer exist now point at the consolidated
record holding their reasoning -- the four lab records are 0016, a test defends
a decision is 0017.
Issue 009's resolution, proven manually end to end before any of it was
written.
A sealed machine pulled an image BY DIGEST from a registry on its own segment
and ran it; then the host applied all four shapes -- package, service with
boot, container from that digest, and an action inside it -- idempotently. That
is the first time the container shape has worked anywhere but a workstation,
and it was the shape blocking the whole substrate bootstrap.
The registry's digests are its own, not Docker Hub's, and that is correct
rather than a compromise: ADR 0046 requires a reference that is exact and
cannot move, and a digest this registry assigned is both. It is also not a
lab workaround -- 0048 names an OCI registry as substrate and 0046 says a first
node fetches "upstream, wherever the image ordinarily lives". This IS that
upstream, scenery in the same sense the transit router is the internet.
The base image now trusts the RFC 5737 and RFC 3849 documentation ranges as
plain-HTTP registries. Scoped to those rather than an address because they
never route on the real internet, so it cannot make a real machine trust a real
registry whatever it is copied onto.
Three faults found while verifying, two of them mine:
My probe script picked an interface with `ls /sys/class/net | head -1`, which
returns docker0 once a runtime exists -- so it addressed the wrong interface and
then, because that address overlapped the segment, broke routing on the machine
entirely. The lab itself is immune: it matches by MAC, for a related reason it
already recorded (bus-position naming on multi-homed machines).
And a test that proved nothing: I asserted `sha256:tooshort` is rejected, but
its letters fall outside a-f, so it failed the character class rather than the
length check. Replaced with hex of the wrong length, after which removing the
length check bites.
ADR 0046's open consequence: "the lab needs a way to place images, and the
machine it places them into needs a container runtime, which a sealed scenario
cannot install either."
The runtime half is done, and it is research 012's reframing applied literally
-- fetch at build time on a machine with a network, apply on a target that
needs nothing. `mesh-lab base build` launches a machine WITH a network,
installs a runtime, verifies it by asking the runtime rather than the package
manager, and publishes the result. Measured: ~30s to install, ~60s to publish,
~700MiB, paid once per lab rather than per scenario.
A scenario that places `runtime` or an image is then raised from that base
image, chosen rather than declared -- a scenario says what it needs, not which
image provides it. If the base does not exist it says so and how to build it.
Verified in a genuinely sealed machine (no route out, confirmed by ping):
package, service including the new `boot: enabled`, and action all applied,
were idempotent on a second run, and read back correctly. Those three had never
run anywhere but a workstation.
The image half is NOT done, and testing found why: a digest-pinned image cannot
be placed from an archive. `docker save alpine@sha256:...` produces an archive
with no repo tag, because a repo digest only exists for an image a registry
served -- so it loads dangling and a container declaring that digest reaches
for a registry the machine cannot see.
That collides with ADR 0046, which has the host REFUSE an unpinned image. Tag
refused by the host, digest unusable in the lab: there is currently no
declaration the lab can raise that exercises the container shape at all. Filed
as 04-ISSUES/009, whose resolution is a registry inside the scenario -- which is
what the real mesh does rather than a workaround for the lab.
Also fixed a weak check of my own, which is the same fault in miniature: the
load was tested with `includes("Loaded image")`, a prefix of both `Loaded
image:` and `Loaded image ID:`. So an unusable dangling load reported success
and the failure surfaced later as a container that would not start.
package.json points mesh-lab at src/cli.ts; the lockfile still said
dist/cli.js. npm install corrected it while installing dependencies to run the
suite. Committing so the worktree is not permanently dirty.
The lab raised an underlay and put nothing on it: correct, and useless, because
the thing it exists to test did not exist. Tier 0 now does, so `place: [host]`
works and a raised scenario finally contains something.
The refusal narrows rather than disappearing. A scenario placing a host and a
substrate is told which half is missing, by name — not that `place:` is
unsupported when half of it now works.
Placement reads back rather than assuming. A file arriving is not a host
working, so the binary is run before it is trusted to answer questions, and what
it reports is read from the machine (ADR 0035). The binary comes from an
explicit path, because the declaration design leaves where artifacts come from
open and a search would harden into the answer by accident.
The integration test that matters is the one asserting the host reports the
MACHINE and not the workstation that placed it. A raised VM and this workstation
differ in every capability — root versus uid 1000, a clean init versus a
degraded one, no docker versus docker, no wireguard versus wg0 — so a host
reporting the wrong machine is obvious here and invisible anywhere else.
And the placed host independently confirms ADR 0031: overlay absent on a freshly
raised machine. The underlay suite already asserted that by looking for
wireguard interfaces; this is a second witness rather than the same check twice.
Two tests failed the moment placement worked, which is what they were for. They
defended "there is nothing to place yet" while that was true; the decision
changed, so they change with it rather than being deleted.
Gate: 75 unit, 20 integration.
The suite next door raises one scenario and asks deep questions of it. This one
asks shallow questions of every scenario — the half that was missing, since both
faults found by hand lived in scenarios nothing ever built.
Adds bootstrap-single, the cheapest, and the loop that lets the list grow. Also
adds the second universal invariant: every address a scenario declared is one
the machine actually holds. A machine that came up bare looks identical to one
that came up correctly until something asks it.
Verified to bite rather than assumed: against a live instance, the real
declaration passes and a declaration claiming an address nothing holds fails
with 'anchor declared 192.0.2.99 on hosting but holds 192.0.2.10'.
Integration now runs with --test-concurrency=1. Two files raise real instances,
node --test runs files in parallel by default, and two concurrent runs of this
suite already produced a whole-suite failure once — every test red, from
resource contention rather than from any fault in the code.
Gate: 45.7s -> 60.2s.
The address collision was found by eye. This is the mechanical form of it: no
two machines hold one address on one segment.
Pure over already-collected facts, so the logic is tested without a hypervisor
— including the cases that would make it useless if got wrong: the same address
on DIFFERENT segments is normal and must not be reported, and one machine
holding an address twice is not two machines.
Asserted against whatever the integration suite has standing, read from the
hypervisor rather than from the declaration. The declaration is what was
accepted, and it was accepted.
Found by asking what gw-devices and gw-home actually were, in a picture that
finally made them easy to see side by side.
planRouters grouped on the exact address list, so `home` declaring a v4 and a v6
address and `devices` declaring only the v4 became two router containers — both
holding 198.51.100.7 on the same segment. The lab raised it without complaint.
Not theoretical. On the raised instance the transit router resolved that one
address to two different MACs across a cache flush:
198.51.100.7 -> 02:c9:16:70:23:29 (gw0, which HAS the :443 dnat)
198.51.100.7 -> 02:bd:75:0b:b0:75 (gw1, which has none)
So home-server's published port worked or did not depending on which container
answered ARP last — intermittent, and it would have presented as a flaky test
rather than as a broken scenario.
One public address is one box. Checked against the thing this models rather than
argued from the model: a bridged modem, a single gateway holding the public
address, one network behind it, and every port forward landing on one host at
that address. Two routers on one address is not a topology, it is a collision.
Gateways to the same segment sharing any address are now one router and their
address lists union, so a v6 address declared on only one of the segments it
serves is still carried. Where such declarations disagree on nat, forwardable or
mapping_ttl, validate refuses — one box cannot behave two ways.
the-ordinary-shape now raises 7 machines instead of 8, and gw0 holds the public
address on eth0 while serving home on eth1 and devices on eth2.
The drawings were confusing, and looking at them showed why: a single stack
ordered by depth put a private network far from the public one it sits behind,
so a gateway's link to the outside ran the full height of the picture through
three networks it had nothing to do with — and two such links overlapped, so
they read as one wire.
Now each public network is followed by everything behind it, depth first. Every
gateway is adjacent to the network it serves, every link is a short stub, and
"behind" is shown by INDENTATION rather than by a line to follow. Gaps are sized
to what they hold, so a gap with no gateway in it takes no room. Transit is not
on a boundary — it reaches every public network at once — so it is stated once
at the top instead of drawing a line to each.
Both sources now order by name rather than by the order the source yielded. The
hypervisor cannot know declaration order, and two pictures laid out differently
cannot be compared, which is the whole point of having both.
Fixed while testing: the gap size and the box placement each decided separately
which network a gateway sat above, and disagreed — reserving the gap above one
sibling while drawing the box above the other, which landed a gateway on top of
a machine in an unrelated network. Both now read one map.
Five new tests, run across every scenario: no link crosses a network it does not
touch, no box is drawn inside a network it is not on, a network behind another
is indented inside it, a public network is not split apart by another group, and
both sources lay the same topology out identically.
Found by rendering the pictures and looking at them, which is the only way
a layout fault shows up.
A gateway was placed below its OUTWARD lane, so one serving `home` and
`devices` was drawn straddling `hosting` and an unrelated `cafe`, with its
connection crossing a network it has nothing to do with. Its first attachment
is the segment it faces; the rest are the ones it serves, and it belongs above
the topmost of those. Transit faces every lane and serves none, so it keeps the
old rule.
Also: the live picture kept its attachments sorted alphabetically, which threw
away the outside-first order the placement now depends on. A segment holding
only gateways-in-the-gaps was counted as occupied and drawn full height with
nothing in it. Badges read left to right, in the order the facts are stated.
The gap between lanes is wide enough that a straddling node no longer covers
the lane's own name and ranges.
`mesh-lab diagram` renders a scenario as draw.io, from either source, through
one layout — so a difference between what was asked for and what exists is a
difference you can see.
The shape says what a resource is and is fixed per kind. The badges say what is
true about that particular one and come entirely from metadata: translation,
forwardability, mapping expiry, refuses-inbound, container-or-VM, running. The
interesting properties of a network are exactly the ones with no visual
consequence — a translated address looks identical to an untranslated one.
For the live picture to be a record rather than a restatement, raise now writes
down what it applied: a segment's kind, ranges and MTU on the link; a gateway's
translation, forwardability and expiry on the gateway; inbound: deny on the
machine. Every behavioural tag is written AFTER the thing works, never at
creation — a failed raise leaves wreckage standing on purpose, and a picture of
that wreckage must not badge translation the router never got.
The pairing earned itself immediately: drawn side by side, every virtual machine
held no addresses. A container's interface carries the device's name and a VM
names its own, so joining them by name silently dropped one whole class of
machine. Fixed by joining on MAC.
Also brings tests under the typecheck gate, which caught integration timeouts
being passed as a 4th argument and therefore ignored entirely.
Reviewed and the criticism was right: 1,072 of 2,128 lines untested, all of
it the half that touches the hypervisor, and no gate. The verification I had
done was real — pings across NAT, TTL counts, ruleset comparisons — and none
of it survived the terminal it ran in, which is 04-ISSUES/005 in miniature.
Ten integration tests against a real hypervisor, each named for what it
defends. ADR 0031: a raised machine carries no overlay, no wireguard, no
mesh config — a scenario that pre-built peering would certify its own work.
ADR 0032: exec is the only way in. ADR 0033: routers are containers while
machines are virtual machines. And the design's claims: raise waits for
usable, snapshots are whole-scenario, NAT hides a private address,
published reaches the machine at the gateway's address.
Mocking the hypervisor is forbidden, so they skip with a reason on a
machine that cannot raise scenarios rather than passing green having
checked nothing.
The suite earned itself on its first run. It found that a snapshot of a
running machine could miss a file written seconds earlier — not stale,
absent — because the write was still in the guest's page cache. That is
exactly the question the lifecycle design listed as open: does a scenario
snapshot need the machines stopped? It does not, but it does need them
flushed. snapshot now syncs every machine before capturing, and the design
records the answer.
The fix buys write-durability, not application-consistency: anything
mid-transaction is still captured mid-transaction, and that is now stated
rather than assumed.
npm run check is the gate — typecheck, 40 unit tests, 10 integration tests.
The full topology now raises: four machines, three routers, a transit
router, six segments, in 35 seconds. Everything the declaration model can
express except `place`, which is refused because the node host it would
place does not exist yet.
Transit was a real gap, not a bug. The design says public networks are
unrelated and routed to each other, never bridged — and I built the
segments and never built the thing that routes between them, so three
public networks were islands and nothing crossed. A transit router now
holds an interface on every public segment, forwarding and no translation:
the closest thing the lab has to the internet, deliberately dumb.
Proven rather than asserted, by ping TTL across the raised topology:
within one segment ttl=64 no hops
across two unrelated public networks ttl=62 gateway + transit
multicast between public networks 0 replies
A flat internet would have shown ttl=64 and answered multicast — which
would let a node discover a peer it could never reach in production, and
report success. That is the fault the as-is layer records the mesh already
hitting with multicast name resolution.
inbound: deny is implemented as a host firewall on the machine, read back
after applying. A declared refusal that silently did not load leaves the
machine wide open, which looks exactly like a machine that is working.
Established and related traffic is accepted, so a defended machine can
still dial out rather than being a disconnected one.
Verified by running, all of it:
home -> devices (policy allow) reachable
devices -> home (policy deny) blocked
behind unforwardable NAT -> out reachable
in -> behind unforwardable NAT unreachable
inbound: deny, dialling out reachable
reaching a machine that denies inbound refused
The two routers differ exactly as declared: the forwardable one carries the
policy rule and no inbound drop, the unforwardable one carries `ct state
new drop` and no DNAT.
A gateway is the one implicit machine in a declaration — a scenario says a
segment sits behind one and never names the thing that serves it. This
materialises it.
A router is a container, not a virtual machine, because it is scenery
rather than something under test (hq ADR 0033). Verified before building
that a plain unprivileged container can do all of it: ip_forward and ipv6
forwarding settable, nftables masquerade accepted, and the conntrack
timeouts mapping_ttl depends on both writable. No privileged mode.
Verified by running, on a machine behind a household gateway reached from
one on a routable address:
home-server -> anchor 0% loss, through masquerade
anchor -> 192.168.1.135 (private, direct) unreachable
anchor -> 192.0.2.50:8080 (the GATEWAY) HTTP 200
The last line is the published-but-behind-NAT case research 004 says only
exists in production. It is now a 32-second scenario on a workstation.
Segments sharing a gateway declaration share ONE router — that is what a
VLAN-capable router is, and two routers sharing an external address would
not work anyway.
mapping_ttl is read back after setting rather than assumed. Those sysctls
are not on every kernel, and a scenario that declared an expiring mapping
and silently got a permanent one would be exactly the fault being built
against.
Four bugs found by running it, three of them the same fault — a failure
made invisible.
The router had no route to a package repository, by design, so installing
nftables at raise time could not work. The image is now built once with
temporary connectivity and cached; every scenario after that needs no
network. That failure was hidden behind `|| true`, which is why it took a
raise to find.
The builder then failed on DNS: exec works before a container has an
address, and I had treated usable as ready. It now waits for the thing
actually needed.
The stock Alpine image ships `auto eth0 / inet dhcp` and its boot-time
networking service flushed the static address the scenario set — on eth0
only, so the outside interface came up bare while inside ones were fine.
The image build now neutralises it: a router reconfiguring itself from an
image default is the lab overriding the declaration. `ip addr add … || true`
had hidden this too, and is now `ip addr replace` with no swallow.
And routers were orphaned by destroy, holding their networks open so
destroy reported removing zero segments. They now carry the same machine
tag as everything else, so one query finds an instance's resources.
A declaration goes in and a disposable mesh comes out. Verified on a
workstation, not asserted: two machines raised and addressed in 14.6s,
snapshot 0.28s, restore-to-usable 11.6s, both families pinging with no
loss, and the workstation with no route into any of it.
The declaration layer implements the model in full — three positions a
machine can be in, keyed on forwardability; gateways carrying the address
the world sees them as; both address families; multi-homing; MTU;
inter-segment policy. It is validated hard because the failures it prevents
are silent: a private range on a public segment produces no error, the mesh
simply never forms. Public segments are refused unless they use RFC 5737 or
RFC 3849 space, and a range wider than the reserved block is refused too.
33 tests, all offline.
The runtime implements less than the model, and refuses the difference.
A scenario declaring gateways, published ports, policy, inbound deny or
place is rejected at raise with every gap named. Raising it would produce a
mesh that silently lacks what it declared, which is the fault this lab
exists to catch — 04-ISSUES/003, where a firewall key is declared in five
manifests and read by no code.
Three bugs found by review and by running it, all of one family:
The readiness check truthiness-tested incusOk's return. `exec … true`
succeeds with EMPTY output, so every machine reported unreachable while
incus exec on it worked perfectly. succeeds() now exists so the mistake is
not available, and network delete had the same bug — it counted zero
segments removed while removing them.
list() split instance from machine on the last dash, so a machine called
home-server absorbed half the instance id and destroy found nothing.
Resources are now found by the metadata they carry, never by name.
restore reported success in 0.79s while the machine's agent was still
starting, so the next command failed. Both raise and restore now wait for
usable and say how long that took — reporting the earlier number is
transport reported as effect, which is the fault the lab is being built to
find.
Two incus behaviours worth recording. Its CLI reads a YAML definition from
stdin when stdin is not a terminal, so a spawned command hangs until the
timeout kills it and arrives with empty stderr — a failure with no
explanation, on a command that works when typed. And it assigns a MAC at
runtime without recording it in device config, so MACs are derived and set
explicitly, which the guest needs anyway: it names interfaces by bus
position, and matching by name configures the wrong one on a multi-homed
machine.
No build step; Node strips the types. The lifecycle has no unit tests
because a fake hypervisor would assert that the fake behaves as expected,
which is the shape of test this project exists to stop shipping.
The node host takes over a machine's packages, services and network, so it
cannot be developed against a machine anyone needs. The place to develop it
has to exist before it does — which makes this phase 0, ahead of every tier
it will later test.
Two scenario classes, per ADR 0029. Bootstrap is virtual machines, the host
and a pinned bundle, with the verdict coming from what the host reports
about the state it reconciled. Full is a complete mesh with a pipeline.
Bootstrap is a strict subset, so the full scenario is reached by putting
more inside the machines rather than by building a second thing.
Nothing here requires a forge, a coordinator or a pipeline to be useful.
Decisions live in novox/hq. This repository carries implementation.