The module that provides artifact-store runs Distribution, the OCI reference
implementation. It was called registry, which named neither the software nor the
provision.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Found by being asked whether processes and containers handle environment the
same way. They do not, and the difference is not cosmetic.
Docker passes --env through literally. A unit file reads three things out of a
value that nothing else does, and a module's environment routinely contains all
three because a generated password is arbitrary bytes:
- % begins a specifier. %H is the hostname. A password containing one is
silently replaced, and it fails later as an authentication error nobody can
explain by reading the declaration.
- whitespace separates assignments. Unquoted, K=a b sets K to "a" and reads
"b" as another assignment.
- a newline ends the line, and what follows is read as a unit DIRECTIVE.
The first two are escaped: quoted, with quotes and backslashes escaped and
percent doubled. The third cannot be — a unit's environment has no way to carry
a line break — so it is refused in validation, near whoever wrote it. Without
that, an environment value could write ExecStart= and have the machine run
something nobody declared.
Ordinary awkward values stay accepted, because refusing those too would leave a
module unable to hold a generated password.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first cut of this added a `daemon` for the long-running case alone. That
would have meant a new vocabulary entry for each of the others — a scheduled
task, a run-once migration, a health check — when they are one thing run at
different cadences. That is a field, not four entries in a vocabulary where every
entry widens what a compromised control plane can express.
So it mirrors a container exactly, because it IS a container's twin: the same
intent, hosted by the machine's own supervisor instead of a runtime. Stays up,
runs once, or runs on a schedule.
Tools, hooks and event consumers are not further modes. They are loaded by a tool
host, which is itself a process that stays up — so the generic case already
covers them, which is the test of whether it is generic.
A scheduled process gets a timer and a unit that finishes; a long-running one
gets a unit that is restarted when it exits. Getting that wrong either way is a
second copy running continuously between fires, or a schedule that never fires.
The modes are exclusive and validation says so near the author: something that
runs once does not run on a schedule, and something not running between fires
cannot be restarted when a file changes.
A missed fire happens when the machine comes back rather than being skipped,
which is the difference between a machine that was down and a schedule that
quietly stopped.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mechanism was leaking into every module. Code of one's own meant a container
and therefore an image; a script meant a service and a unit somebody else had to
install. One intent — run this and keep it running — expressed two unrelated
ways, with the hosting chosen before anything could be declared.
A daemon names a bundle and a command. The host fetches it, refuses it unless it
hashes to what was declared, unpacks it where the mesh keeps such things, writes
the unit and puts it in the state asked for. The unit is the mesh's, generated
whole and saying so, because an edit that survives until the next declaration and
then vanishes is worse than one that is refused.
Its identity is the bytes AND how it is run: two daemons from one bundle
differing only in their command are different daemons, and tracking the digest
alone would call the second unchanged and leave the first running. The unit is
rendered deterministically for the same reason — environment from a map would be
written in Go's iteration order, so every apply would see a different unit and
restart an unchanged daemon for ever.
restart-on is honoured as a service's is: a running process does not re-read its
configuration, so replacing a file and finding the daemon already up leaves the
machine behaving as before while every check passes.
A full-host shape, not a portable one: it needs a process supervisor to install
into. It does NOT need a container runtime, which is the point.
Two guards caught this properly and both were updated deliberately rather than
silenced: the vocabulary count, which exists because every addition widens what a
compromised control plane can express, and the shape test that catches a kind the
language has and a host cannot apply — added after `network` did exactly that.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two faults that both reported success while being wrong, found while proving
the firewall module actually delivers.
A unit whose job is to apply something and exit — load a rule set, set a
sysctl — is inactive the instant it succeeds. Reading that as stopped made it
permanently unsatisfiable: the host started it, it worked, the host read back
stopped and reported the machine as not doing what it was told, on every apply,
for ever, with the rules correctly in place the whole time. That is what the
firewall has been doing on every machine it was assigned to, and why the
four-machine bed was red.
And a container took its identity from its own fields, not from the files it
reads. A file written in an earlier apply — or before the container declared it
as a dependency — left a process holding a credential the mesh had already
replaced, with everything reporting success (novox/hq 04-ISSUES/045). What a
container reads is now part of what it is, so the comparison is a standing one
rather than a tripwire that fires during one apply and never again.
It defaulted to mesh-host.service. What packaging/ ships is nox-mesh-host.service,
so on any machine with the packaged unit installed the installer looked for
something absent and told the operator a real machine needs it installed — when
it was installed, under the name the installer was not using.
Only the lab missed it, because the lab passes --host-in-background and never
names a unit at all.
Genesis ended with a mesh that runs and cannot make anything: every module in
the catalogue names artifacts and nothing had built them, so the first thing
anybody had to do was install a builder by hand.
The installer already carries one — it is what built the control plane — so
this is the same two acts the control plane goes through, in the same order:
publish it, so the mesh names it by a digest its own registry assigned rather
than a local identity nothing else can fetch, then install it as an ordinary
module pinned to that. And then the part only it needs, a broker account, issued
before the push so it arrives with the declaration rather than after it.
Verified on a bare machine: the install ends with a builder running, and that
mesh then built the shared base images and a module on top of them with nobody
helping it.
The carried image is the builder now. The publish step still pushed it, so the
registry got a builder under the control plane's name and the mesh installed it
as the control plane — which presented as a control plane that started, printed
a builder's usage, exited cleanly, and did it again. Caught by the lab on the
first genesis run, at the step that waits for it to answer.
It carried the thing it was going to run; it now carries the thing that makes
it. One artifact either way — but a mesh raised this way holds a control plane
it built from a repository and a commit it can name, and can therefore build
again. A mesh handed a finished image could not, and had no way to find that
out until somebody needed it to.
A build step sits between load and bundle, because the bundle must name an
image and that image no longer arrives finished. Everything after it is
unchanged: a locally built image is named by the digest of its own
configuration, which is exactly what the carried one was named by.
Refused in preflight when nothing says what to build, so a run that cannot
finish says so before it has changed anything.
Enrolling IS the machine speaking to the mesh, so straight after it the mesh has
always heard from this node — and the step took that as proof an agent was
running and skipped starting one.
The cost is silent and total. Everything after is the control plane being told
things, and nothing it is told reaches a machine with no agent to collect it: the
registry push at step 7 was accepted, the module recorded, and no container ever
created. It surfaced three minutes later as 'the registry is not there at all',
one step from its cause and looking nothing like it.
Both halves are asked now. A process may be wedged and collect nothing, which is
why the mesh is asked at all; and the mesh may have heard once from a machine
running nothing, which is why the machine is asked too.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The first real run of mesh-bootstrap stopped in preflight, dialling 192.0.2.250:5000
for ninety seconds on a machine whose network was fine. That address is the registry
the lab used to raise; the substrate template still names the control plane by it,
and step 3 replaces that reference with the id of the image this installer carries.
Nothing ever pulls it.
So preflight excludes the control plane's resource by identity, rather than by the
happy accident of the template filling its slot with something that needs no registry.
Every other container's registry is still dialled, because those are somebody else's
images at somebody else's registry and a machine that cannot reach one fails inside a
pull, which says the wrong thing.
Also: `make bootstrap` takes BOOTSTRAP_OUT. The lab now builds the installer from
source before every raise, into a path it chooses, and a caller that could not say
where the output goes would have to copy it afterwards.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An image id does not survive `docker save` -> transfer -> `docker load`. The id is
the digest of the image's *configuration*, and a runtime rewrites that
configuration as it loads: a newer Docker saves in one format, an older one stores
it in another. Same layers, same program, different name. Measured on a live raise:
saved on the workstation sha256:b86bb81ca2f9691f24f4725f50962d1e49c98c5ffe211113241243d42d18ceea
loaded on the machine sha256:2dc219046c73702fc640317f0342a28ec962ef1e9ef547b2f02861c508ca78fb
`internal/image`.ID read the id out of the carried tar and its comment said that
was the id the runtime would assign. That is true on the machine the image was
built on and false on every machine it is carried to — which is every machine this
program exists for. The installer then either stopped at step 2 refusing the
runtime's answer, or would have written a bundle naming an image the machine does
not hold; and nothing serves an image named by the digest of its own configuration,
which is the whole point of naming one that way, so the apply would have died
inside a pull that cannot succeed. The lab hit this.
So the image is identified by its TAG, which is ordinary metadata the tar carries
through unchanged. The runtime is asked what that tag resolves to before the load
(already held, nothing to do) and again after (this is what the bundle names). The
tag never reaches the bundle — a pinned bundle may not rely on one, ADR 0006 — it
is how the id is obtained, not what is written down.
- image.ID becomes image.ArchiveID, and says plainly that it is a fact about the
file and not a prediction about any machine. It is kept for reports, and printed
beside the runtime's answer whenever the two differ.
- Idempotence is decided from what the runtime holds under the tag, not from a
predicted id, which cannot answer the question at all here.
- An untagged archive is refused, in preflight and again at the load: there would
be no portable name to ask about, and the only thing left is scraping a sentence
`docker load` writes for a person. `make bootstrap` refuses an id or an untagged
image, so it is caught in front of whoever can fix it.
- A dry run cannot know the id and says so rather than pretending. Run refuses to
write a bundle carrying an unconfirmed id at all.
Tests: the injected Runner now answers with an id DIFFERING from the tar's, and the
runtime's answer is what must be used. The test that refused a differing id encoded
the mistake and is replaced by one refusing an answer that is not an id at all.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The catalogue's mesh-control manifest landed while this was being written, and it
does what the ordinary case does: it keeps its secrets under /var/lib/mesh and
mounts them into the container at /run/secrets, so MESH_STORE_INVENTORY_FILE names
a path that no own-secret writes. Matching on the path alone found nothing and
would have refused a correct manifest.
So the lookup follows the volumes. It also reads the other shape the manifest uses
— `VAR=${secret:name}` inside the environment file a container reads — which is
how a value that is not a path gets in at all, and which is where the broker's two
credentials live.
That generalises what is delivered: every variable the module fills from a secret
is looked up in the substrate's control plane. What the substrate names is accepted
through `secret accept`; what it does not is left for the mesh to generate, and
said so. A store connection the substrate does not name stays an error — a control
plane that cannot open a context is not one.
Checked against the real manifest (mesh-catalog feat/control-plane-module): five
variables resolve, the placeholder pins in one place, and the container it waits
for is `mesh-control`.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Steps 6 to 10, which turn a substrate into a mesh that can maintain itself
(novox/hq ADR 0067).
6 enrol a node record, a token, `mesh-host enrol`, and the host agent
running. Proved by the mesh having HEARD from the node, not by a
process existing: a host that cannot reach the broker looks exactly
like a successful install until the first push applies nothing.
7 registry the module that gives this mesh an image store, registered from a
--catalog checkout, assigned and pushed. Its image is upstream and
never built (04-ISSUES/029) — a placeholder digest there is refused.
Verified by asking `/v2/`, because a container that is up is not a
registry that serves.
8 publish the carried image pushed into that registry, which assigns it the
first manifest digest it has ever had. This is the hinge: without
it the mesh works and can never upgrade itself.
9 control the control plane registered as an ordinary module pinned to that
digest, with the substrate's own store connections delivered
through `secret accept` — read out of the bundle that made them,
because the mesh cannot invent a credential that predates it.
10 retire the temporary control plane dropped from the bundle and removed by
the host's ordinary removal pass.
Every step asks before it acts and reports "already done". No step leaves the
machine without a control plane: steps 9 and 10 overlap deliberately, and two
stateless control planes are untidy rather than broken.
mesh-control's `internal/builder`.PublishImage is mirrored rather than imported —
tier 0 depends on nothing that must be installed first — with one correction: the
digest is chosen from RepoDigests by repository instead of taken as element zero,
so an image pushed to two registries cannot silently pin this mesh to the wrong
one.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The substrate raises a control plane and a module will later declare one. If both
are called `mesh-control` then for one moment two owners hold one container, and
the host — which tracks what it owns — has no way to stop owning something without
destroying it. That looked like a missing mechanism.
It is a naming problem. The substrate's container becomes `temp-mesh-control` and
the module's keeps the plain name: two containers, two owners, nothing to hand
over. Dropping the temporary one from the bundle at the end is then destruction by
omission, which is what the host already does to anything that leaves a
declaration — and the right end for something named "temp" (novox/hq ADR 0067).
The rename is textual and matches the QUOTED name, so the `mesh-control` inside
the image reference is not caught by it. Read back afterwards: the produced bundle
must call it the temporary name, and no other container may have been renamed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The only complete written-down copy of how a mesh is stood up was an integration
test in the lab. That is why every bootstrap gap kept being found late: an install
procedure that lives as a test fixture is exercised by whoever writes tests, never
by whoever installs. This is that procedure.
A separate binary, not a mesh-host subcommand. mesh-host says of itself that it
connects to nothing and listens on nothing and that what it applies comes from a
file, and that sentence is what makes an always-running root daemon auditable. An
installer loads images and interrogates a control plane. Same tier, different
program.
The control plane's image is carried, not built and not fetched. The forge that
holds its source runs on the mesh, so a bootstrap that had to fetch it would need
a mesh in order to raise one. Embedding breaks that cycle the way the carried
bundle breaks "copy it onto a machine and run it". The image id is read out of the
saved tar before the runtime is asked anything, which is what makes the load
idempotent: the installer can ask whether the machine already holds exactly this.
Five steps, each idempotent and each saying whether it found or changed something,
because this is run over and over by somebody getting a machine working. It stops
at a running substrate with a control plane that replies — enrolment, the module
catalogue and assignment are the next stage and are deliberately absent.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A manifest digest is assigned by a registry on push, so insisting on one meant a
registry had to exist before the thing that lets a mesh have a registry could
start — a dependency the pinning rule created by accident, not a pin. The mesh's
own control plane is built from source and lives in no public registry.
A bare sha256:... names an image the machine already holds, by the digest of its
own configuration: immutable and unforgeable in exactly the way the rule asks
for. Absent, it says so plainly rather than failing at a pull nothing serves.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A schedule: container (ADR 0053) is installed as present state and never run at apply — the
Scheduler fires it later on its cadence. But a service or run-once container only gets its
image as a side effect of docker run, so a scheduled step's image was not pulled until its
first scheduled fire: absent from the node right after a successful apply, so the first run
paid the whole pull latency and tooling that expects the image present after apply found it
missing.
applyContainer now probes the runtime and ensures the pinned image present for a scheduled
step before recording it. A new ensureImage helper inspects the image and pulls it only if
absent, then reads back (ADR 0018). Ensuring an image is not running it: no docker run fires
the container, so the no-run invariant of ADR 0053 holds. The runtime probe, previously
skipped for a schedule, now runs because a pull needs it — the schedule.go comment is updated
to match.
Tests: the install-does-not-run test is extended to allow the image-ensure while asserting no
fire and no needless pull; a new test applies a scheduled container whose image is absent and
asserts it is pulled and still not started. go build, go vet, go test ./... all pass.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The recurring twin of run-once, one modifier over: a container marked
schedule: "<cron>" is run to completion on its cadence, not started as a
service and not run once as a gate.
The gating rule is deliberately reversed. Installing a schedule records it
as present state and reports the node current at once (applySchedule) --
it never runs the container and does not gate what follows. A Scheduler,
held for the life of the daemon and re-established from each applied
declaration (the declaration is the source of truth, ADR 0018), fires the
container off an injected clock. A run that exits non-zero is logged and
never fails the apply or flips the node's state, because it happens
outside the apply and the store entirely. Runs never stack: a run still
going when the next is due is skipped, not started as a second copy.
No new host shape and no new action -- schedule is a string on the
container the host already has, and the host process runs the container
itself rather than installing a system timer (the rejected option 1). A
minimal five-field cron (declaration/cron.go) validates on arrival and
computes the next due minute; time is injected so the scheduler is tested
without the wall clock.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module can declare state but not a step that runs at first boot. This adds
`run-once: true` to the container shape: the host runs it in the foreground,
requires it to exit 0, and records that it did — as the digest of the
declaration, so a re-apply does not re-run it unless the declaration changed.
Because the declaration is applied in order and a failed run-once step gates the
apply the way a failed action does, whatever is declared after the step starts
only once it has completed. That is how "before the broker starts" is enforced,
with no dependency graph the host must resolve (ADR 0005): the step is declared
first, and the container that needs it is never reached until it is done.
No new host shape and no arbitrary host command — a run-once container is
strictly less powerful than an action. Validation refuses run-once with
restart-on (contradictory lifecycles). Six unit tests; go test ./... green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The tenth shape (novox/hq ADR 0051). Shared, pre-existing data — a media
library, a download spool several modules use — is the operator's, not
the mesh's. A `directory` resource is the host's own: it creates it,
chowns it, sets its mode and removes it when empty. An access is the
opposite on every axis.
Add the `access` type to the vocabulary. Its applier confirms the path is
present and changes nothing: it does not create, chown, reconcile or set
a mode. Absent is refused clearly — the operator must provide it — rather
than created, because a bind mount whose source is missing is made as
root by the container runtime with the wrong ownership (04-ISSUES/026).
Undeclaring an access forgets the record and never touches the path,
which is the data loss ADR 0030 prevents, on a directory the mesh never
made.
Full hosts speak it (it gates a bind mount, which needs the container
runtime); the vocabulary guard test records the decision that made it the
tenth shape. Unit tests cover present, absent-refused, and
undeclared-left-alone.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container reads a mounted file once, at start; its spec (image, env, volumes)
does not include a mounted file's content, so a settings change that re-renders the
file left the running process holding the old value while every check passed. Give
Container the restart-on field a Service already has, and recreate the container
when a named resource changed this pass. Unit-tested (recreated on change, left
alone otherwise) and proven in the mesh-lab: a running grafana runtime picked up a
token change on the next push.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The mesh decided "has this machine caught up" by comparing its send
time to the report's arrival, and lost the race it invited: an apply
started under the previous declaration finishes after the next one is
sent, its report lands newer than the send, and the machine reads as
caught up with words it has not read yet. The lab hit exactly that —
one test's closing push was still being applied when the next test's
push recorded its send, and the next test then read files that were
never going to be there yet.
Clocks cannot answer "which". The report now carries the digest of the
exact bytes it applied — the same bytes, hashed the same way, that the
mesh recorded when it sent them — and which-declaration becomes an
equality the mesh checks rather than an ordering it hopes.
A directory a module mounts into its container belongs to whoever runs
inside — grafana's 472, redis's 999, www-data's 33 — and none of those
has a row in the machine's passwd. Owner-by-name refused them all,
which looked principled and meant every module whose container drops
privileges could not own its own data.
The lab showed both coats of it in one run: the store's config file was
unreadable to the store, restarting forever on permission denied, and
the forge could not traverse into the 0700 root-owned directory that
held its files — a directory that had only become root-owned when
declaring it fixed 04-ISSUES/026, because Docker used to create it
0755. A fix that tightens ownership without a way to say whose it
should be moves the fault, not removes it.
"uid:gid" and bare "uid" are numeric and chowned as given; a name still
resolves as before, and a name with a colon is refused rather than
half-read.
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.
The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.
What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.
Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
The shape was added everywhere except the one list that decides whether
a host can actually apply it, so every declaration carrying a network
was refused whole — correctly, and with the reason stated:
resource "umami.net" is a network, and the arch host does not
implement that shape
The mesh behaved as designed throughout. A host that applied the parts
it understood would leave a machine that looks configured and is not, so
it refused the lot and said why. What was missing was anybody reading
the host's log.
The vocabulary test did not catch it because it checks what the language
has, not what a host can do — those are different lists and only one of
them was updated. There is now a test that a host claiming to do
everything implements every shape the language has. It fails with the
message above when the registration is removed.
A network needs the same runtime a container does, so it belongs to a
full host and not to the portable floor.
Two gaps found by writing the first real module's manifest rather than
by reasoning about one. Both are fields on existing shapes, so the
vocabulary is still nine.
**env-file on a container.** A declaration reaches a node over the
broker and `env` is plain text in it, so a password there is a password
the broker sees — the transitive trust refused everywhere else. A sealed
file arrives unreadable, the host writes it, the runtime reads it. It is
also simply how third-party software takes credentials: nothing shipping
in a container will read a path the mesh invented, and every one of them
reads its environment.
**secrets in a file's content.** A program wanting its token inside a
JSON document cannot be handed a file that is entirely a token, and the
mesh cannot compose the document because it discarded the value. So the
module supplies the document with `${secret:name}` in it, the mesh
delivers the value sealed, and the host is the only thing that ever
holds both.
Substitution is textual and the host learns no formats. Deliberate: a
mechanism that understood JSON would be asked to understand YAML next,
and then INI, which is how the arrangement this replaces became
something nobody could hold in their head. The module knows its own
format because it wrote the rest of the file. The sharp edge is stated
rather than left to be discovered — a value containing a quote is not
escaped for whatever surrounds it.
Refused in both directions, because both are somebody being wrong about
where a credential is: a placeholder with nothing to fill it would write
`${secret:x}` into a config file, and a secret the content never uses
means somebody believes a credential is in a file where it is not.
A file that carries one is 0600 unless the module said otherwise.
Found by asking what the conversion needs, and it is the one failure in
this system that cannot be undone.
Unassigning a module made its directory an orphan, and an orphan
directory was deleted with everything under it — os.RemoveAll — while
the report said "removed". A database's files, a mail spool, somebody's
uploads. Reproduced before fixing: assign a module, let a service write
into its directory, unassign the module, and the file is gone.
Now a directory that still holds something is kept and said so, naming
how many items are in it.
What makes that safe rather than merely cautious is the removal order,
which was already right. Everything the mesh puts in a directory is
itself a declared resource, and orphans are removed in reverse
declaration order — so what the mesh wrote is already gone by the time
the directory is reached. Anything still there was put there by
something else, which is the definition of data.
It is the host's own line applied to the one shape where getting it
wrong does not recover: it removes what it made and leaves what it
merely configured. An empty directory is what it made; a full one is
not, and an empty one is still removed so nothing accumulates.
Files are unchanged. A declared file is the mesh's own, and losing a
config file is not the failure this is about.
novox/hq ADR 0029, and work breakdown 1.3. A module of several
containers had no way to let them reach each other by name: a container
declaration could join a network and nothing could create one.
An action was the obvious alternative and is refused on removal —
"an action has no footprint the host can undo", so a network made that
way outlives every module that is ever unassigned, and the mesh cannot
tell. A resource the mesh can create and never clean up is one it should
not create.
A name and nothing else. Not a driver, a subnet or a gateway: each is
something a module would have to know about the machine it lands on, and
a module naming a subnet collides with whatever else chose the same one.
It needs no new ordering rule. Resources apply in declaration order and
orphans are removed in reverse, so a network written before the
containers that join it is created first and removed last — after they
are gone. A runtime refusing to remove one still in use is reported
rather than swallowed, because that means something undeclared is
holding it.
The vocabulary guard fired on the change, as designed, and now names the
record instead of a number: nine shapes, with the argument beside the
count.
Creation reads back rather than trusting an exit status (ADR 0018): a
runtime that reports success and made nothing leaves every container
that joins it failing to start, one step from the cause.
Half of novox/hq work breakdown 1.3, and it needed no change: the apply
loop walks d.Resources and sorts nothing, so a module that needs one
thing before another says so by writing it first.
Asserted because it is the kind of property a later change breaks
silently. Sorting the resources for any good reason at all — by type,
by identity, for a tidier report — would still pass every other test in
this package.
It is sequence, not readiness. A container started is not a container
ready, and nothing here waits: what depends on something being usable
retries, which is what both example provisioners do and is the more
robust answer anyway, because a dependency can restart long after
everything was applied.
Two mistakes worth keeping in the test's own comments. The first
version stubbed the runner to always succeed, so verify passed, every
action counted as already done, and nothing ran — the assertion was
measuring an empty list. The second declared the actions over the link,
which refuses them: only a bundle may carry an action (ADR 0005).
From auditing the decision records: of 28, only 12 were named by any
test, so "which decisions are defended" could not be answered without
reading everything. ADR 0017 says a test names the decision it defends —
that rule was itself unenforced.
Most of the gap was citation, not coverage. Drift detection was tested
in several places without naming ADR 0011; the archive refusal without
naming 0012; forged declarations without naming 0002. Named now, so the
question is answerable by grep.
The bundle was the real gap: nothing tested substrate-first-node.lock at
all. It is what a machine becomes when there is no mesh to ask — the one
declaration applied with nothing to verify it against — and it was
edited by hand and read by nothing but a running host.
Two tests now assert what it carries: exactly postgres, lavinmq and the
control plane. That defends ADR 0028, which removed the object store
from the substrate after it had been a member for months on the strength
of "it cannot grant itself a bucket" — true, and the answer to only half
the test. Nothing counted what the bundle held.
Fault-injected, and the first attempt did not bite: the injection landed
on a comment line, which stripComments discards. Injecting into the
image field fails as it should.
novox/hq 04-ISSUES/002, which was recorded against HAL and is present here: a
machine asking the mirrors for a version they have already replaced gets a 404
from every one of them. The package exists and the declaration is correct — it
is the machine's view that is old — and reported as a generic install failure
it sends somebody to check the manifest, which is the one thing that is right.
It is deliberately not fixed by syncing. `pacman -Sy <pkg>` installs a package
built against libraries the machine does not have: a partial upgrade, which
this distribution does not support and which surfaces much later as something
apparently unrelated. The remedy is a full upgrade, which is a decision about
the whole machine rather than something a host does silently while applying one
resource. So this says which of the two it is looking at, and leaves the
decision where it belongs.
Every mirror, not one: a single mirror timing out is transient and retrying is
the answer.
And the package manager's own words were being discarded entirely — the output
was read into `_`. Whatever it said is now part of the failure, which is the
rule everywhere else here and was not being followed in the one place the
reason only exists in the output.
The commit before this said "told where to resolve names" and passed --dns,
which is not what it ended up doing. This is that correction: a container is
given the names themselves, written into its own hosts file by the runtime.
The reason for the change is the decision the mesh already made about names — a
file rather than a resolver, because it works on every runtime, needs no
package and has no failure mode of its own. Passing a resolver address would
have required a resolver to exist, which at that point none did.
A resolver is coming, for the case a file genuinely cannot express: a service
named under a machine, postgres.novox.internal, where the wildcard cannot be
enumerated in advance. When it arrives it will need this field back under its
own name. It is not being kept in the meantime — a field nothing fills is a
field nobody can trust, and the vocabulary is asserted by a count for exactly
that reason.
A container does not inherit the machine's names. It gets its own /etc/hosts
holding its own hostname, and a runtime rewrites resolv.conf — so every
internal name the mesh wrote for that machine is invisible to what the machine
is running.
That was hit for real, in the lab: a database client on one node could not
resolve another node, on a mesh where both names were correct and present on
both machines. It was worked around by resolving on the host and passing an
address, which is the kind of workaround that should not be needed twice.
A field on an existing shape, not a ninth shape — the vocabulary is still the
eight the count asserts.
Per container rather than by editing the machine's resolver configuration: that
file belongs to something else on most machines, and a host that edited it
would be fighting whatever owns it on every boot — the fault this host exists
to avoid, in the place it would be hardest to see.
A container told nothing is run exactly as before. Most containers should
resolve whatever the machine resolves, and passing an empty flag would be a
change of behaviour dressed up as a default.
A suspended laptop's connection is dead the moment it wakes, and the socket
looks perfectly healthy from inside the process — no error, no close, because
nothing has tried to send anything. Heartbeats find out twenty or thirty
seconds later. For that time the node believes it is in a mesh it has left,
which is the one state this design says must never be indistinguishable from
being connected. The machine knew immediately.
So being roused ends the current attempt rather than only shortening the wait
after it: shortening the wait would do nothing at all, because the process is
not waiting — it is sitting inside a connection that will not return.
A signal, because nothing may listen on a node (novox/hq ADR 0004). A socket
for this would be a control surface on every machine, reachable by anything
that can reach the machine, in exchange for saving twenty seconds — and the
whole security argument rests on there not being one.
Two rouses in the same instant are one: a machine suspending and resuming
repeatedly must not build a backlog of reconnections to work through. And the
backoff is not reset by being roused — that says the machine changed, not that
whatever was refusing the connection has stopped, and a laptop woken on a
network with no route would otherwise retry at full speed for as long as
somebody keeps opening the lid.
The dispatcher acts on the events that change where packets go and not on
`down`: the link is already gone there, reconnecting will fail, and the backoff
exists for exactly that.
Each context owns its own database (novox/hq ADR 0008), so a third context is a
third database, created and named the same way — which is the whole of adding
one to the bootstrap, and is why the count is not something the substrate has
an opinion about.
The schema step verifies all three now. It checked two while creating three,
which would have reported success for a context whose tables were never made.
While the store initialises it runs a temporary server on the unix socket only,
then stops it and starts the real one. A socket check sees that temporary
server, the action exits happy, and the verify a moment later lands in the gap
between the two and fails — reported as "the action ran without error and its
own verify still fails", which is true and names nothing.
Intermittent, so it read as a slow machine. Both the action's own loop and its
verify now ask the same question, over the port the init phase deliberately
does not open: an action and its verify asking different questions is an action
that can succeed into a state its verify rejects.
A JSON decoder reads one value and stops, so a file holding a declaration and
then anything else parsed as the declaration and the rest was never looked at.
The machine applies something, reports success, and what it applied is not what
the file says — the same fault this host refuses everywhere else, in its
quietest form.
Not hypothetical. A test harness had been appending a line to the substrate
bundle by accident; every apply kept working and nothing said so for as long as
it was wrong. That is how the bug was found, and it is the argument for the
refusal: a file with something after it may be a truncated rewrite or two
declarations run together, and applying the first would be applying something
nobody wrote.
Trailing whitespace is not "something after it".
PKCS#8 PEM, not this host's own base64. The mesh delivers a PEM certificate
beside it and every TLS server there is reads PEM: nginx's ssl_certificate_key,
Go's LoadX509KeyPair, openssl s_server. Stored the other way the file was
intact, present, correctly permissioned, and unusable — the machine failed at
the moment something connected, which the lab found by connecting.
A key in the old encoding is refused by name rather than called corrupt: it is
replaced by enrolling again, and that is a different remedy from a damaged
file.
A fourth key, reported at enrolment like the others. The reasoning is the
one this file's neighbours already give twice: a key used for two
purposes is one rotation away from breaking the other.
The private half never leaves the machine. The mesh is told the public
half and signs a certificate binding it to this node's name inside the
mesh — so there is nothing to seal, and a copy of what the mesh holds
certifies nothing it did not already certify.
It does not make one on demand, for the same reason the sealing key does
not: a key the mesh has never certified is a key nothing will trust, so a
node that quietly generated one would serve a certificate for a key it no
longer has and fail in a way that names neither.
The previous commit continued past every failure, and the lab found the
cost immediately: the bootstrap's store-readiness gate failed, the apply
carried on and started the broker and control plane against a machine
that was not ready, and the database still initialising was shut down.
An action is the only shape whose purpose is to make something true
before the next thing needs it — which is why it is the only one with a
verify. The bootstrap is a row of them. Everything else is independent
state, and stopping there is what made one broken module hold a whole
machine hostage.
The report says which happened: "these things failed" and "these things
failed and the rest was never tried" are different machines.
Found in the lab while proving something else. A machine assigned a
module declaring a package that does not exist applied NOTHING on every
later push, for ever — the broker's queues were empty, so the declaration
had been delivered and read; the machine stopped at the first failing
resource and never reached the rest.
A machine with one bad module and nine good ones ran none of the nine,
and the mesh reported "failed" without saying the rest were never
attempted. Nothing that re-pushes to machines that are behind could
recover it either: it would retry a permanent failure for ever and make
no progress on anything else. And which nine a broken module blocks is an
accident of resolution order.
The behaviour had a test asserting it, citing ADR 0010. That record does
not decide this — it argues about pipelines against reconcilers, and says
nothing about whether one resource failing should stop the next being
attempted. The citation was doing more work than the record supports.
So: everything is attempted, every failure is reported, and the first
line says how many. The case for stopping was that a later resource may
depend on an earlier one. It still may — and it then fails its own check
and is reported, which is more information than skipping it. This host
reads back after every write precisely so that is caught rather than
assumed.
Unchanged: a declaration that cannot be PARSED is still refused whole.
That is a different thing — "this machine could not do it" against "this
was never a declaration" — and they are fixed in different places.
Recorded as novox/hq 04-ISSUES/011 with the evidence.
The bootstrap's readiness wait failed on a loaded machine and reported
`docker exited 1:` with nothing after the colon. The file's own comment
already records this failing three times before and being fixed by
running it again — "the worst kind, because it teaches people to run
things twice".
Raising the timeout a second time would treat the symptom. What makes a
retry the only available response is a timeout that reports nothing, so
the wait now prints what pg_isready says and the store's own last lines
before giving up.
not a service
A shell, a terminal, a chat client, a desktop are a package plus
configuration in somebody's home. A mesh with no notion of a user can own
/etc and nothing anybody looks at, which is most of the reason to manage
a machine at all.
Three shapes, and the vocabulary test asserts the count precisely because
widening it widens what a compromised control plane can express:
user a login, its shell and its groups
archive a set of files, fetched by digest and unpacked
(file) gains `bytes` for what is not text, and `owner`
`user` also makes "zsh is my login shell" declared state. chsh is a
command, the link may not carry one, and a shell settable only by hand is
a shell the mesh cannot manage.
Groups are additive and never pruned — usermod without --append REPLACES
them, which would silently remove every group that makes a login able to
use the machine. A machine's own groups are not the mesh's to know about.
The archive is the one place this host reaches out on its own; everywhere
else it holds one outbound connection and fetches nothing. So it carries
the discipline the bootstrap already uses for images: pinned by digest,
and the digest checked before a single file is written.
Two decisions in the unpacker worth naming:
- an entry naming a path outside the archive is REFUSED, not sanitised.
Rewriting it to land inside would put a file somewhere nobody asked for
and report success. Found by the test: the first version quietly
relocated it.
- symlinks and device nodes are refused rather than skipped, or an
archive that needed one arrives silently incomplete.
A partial host does archives and refuses users: an archive needs a
filesystem and a way to fetch; a user needs a user database it is allowed
to write.
Found raising two machines: `apply <file>` refused the bundle example in
this repository with `invalid character '/'`. The bundle strips whole-line
comments; apply handed the raw bytes to the parser. So a file this repo
ships could be built into a binary and not applied from disk.
This is the third instance of one fault. There is already a test here
named "what validates is what is applied", written when `mesh-host
bundle` said yes and `reconcile` said no about the same artefact — two
paths to one thing, disagreeing. Fixing that instance left the shape
intact, so it came back somewhere else.
So the fix is structural rather than local: `declaration.ParseFileTrusted`
is the one way to read a declaration from disk, and the bundle and apply
both use it. Comment handling and its test now live in one place, since
having them in two is how it came to be done in two.
The wire format is untouched — over the link it stays exactly JSON,
because a format with a second thing to strip is a format with a second
thing to disagree about. Asserted, and confirmed to fail if the link
starts stripping.
Found by raising a mesh end to end for the first time. Enrolment's own
help says the token "is the only thing it needs", and it also needed
--name, with no default. Without it the failure is:
cannot reach the broker at 192.0.2.10:5671 as : username or password
not allowed
An empty username, and nothing about the cause.
The node cannot work its own name out. The broker account it
authenticates as is named after it and exists before this machine has
been told anything, so the name has to arrive with the rest. It is not a
secret and the issuer already knows it.
--name stays, as an override for a token issued before the name
travelled in one, and says so when it is needed rather than failing at
the broker.
Also corrects the bundle example, which claimed to stop before the
control plane runs and has raised one for some time. A comment about what
something does not do is a comment nobody updates.
The enrolment request is a struct in each repository, and this node now
reports a third key — the one its secrets are sealed to. That wiring had
tests on each side and had never been run across the join, where a
renamed field fails silently: enrolment succeeds, the key is absent, and
the node looks joined until the first thing sealed to it cannot be
opened.
So this writes a real one — keys generated the way enrolment generates
them, not typed as literals — and the private half of the sealing key
beside it, so the other side can prove what it sealed is openable rather
than merely present.
The mirror of the declaration check that already runs the other way.
Everything else in a declaration is visible to whatever carried it. The
message is signed so it cannot be forged, and signing does not make it
unreadable — a password in `content` is a password the broker sees, which
is the transitive trust this design refuses everywhere else.
So a node generates a third key at enrolment and reports the public half,
exactly as it does for its identity and its overlay key. A file may
arrive `sealed` instead of `content`; the host opens it with that key and
writes the result. The control plane can then store a credential it
cannot use, and the broker relays a blob it cannot read.
A third key rather than reusing one of the two. The identity key signs
and is Ed25519; the overlay key is WireGuard's and is tied to being on
the private network, which a machine may not be. A key used for two
purposes is one rotation away from breaking the other.
Details that are not incidental:
- sealed and content together is refused, so "was this the secret or the
placeholder" is answerable by looking
- a sealed file defaults to 0600 rather than 0644, because the
consequence differs; an explicit mode still wins
- a node with no sealing key refuses the file rather than skipping it. A
machine that quietly omits the one resource carrying a credential looks
configured and cannot connect
- what is recorded is a digest of what was written, so drift on a
credential is still detected without the node keeping the value, and
the report that goes back over the broker carries neither
The key is made at enrolment rather than on first use. One made later is
one the mesh was never told about, so nothing could ever be sealed to it,
and the node would look fine and receive nothing.
This is why sealing was borrowed from another mesh's mistakes rather than
its design: there, credentials sit encrypted in the control plane's
database — which guards the database file and nothing else, since the
same value is also in each node's environment file in plain text and
inside every connection string composed from it. Its own tooling has to
search by value rather than by name to find the copies, and says the ones
inside composed URLs are usually the only copies in use.
"The host needs no new vocabulary" is the load-bearing claim behind every
computed and contributed resource on the other side, and it had never
been tested against this parser — only asserted.
Skipped unless MESH_EMITTED names a file, so it stays a check somebody
runs deliberately rather than a dependency between two repositories.
mesh-control plan <node> --json > /tmp/d.json
MESH_EMITTED=/tmp/d.json go test ./internal/declaration/ -v
Confirmed to fail when the declaration carries an action, which is the
thing this parser exists to refuse.
Asked how the mesh would know if somebody edited their hosts file. It would
not. The file was rewritten within five minutes and the outcome said
"updated" -- which is exactly what the mesh changing its own mind looks like.
So the change vanished, nothing anywhere said why, and the obvious thing to do
is edit it again.
The host now records a digest of what it wrote, which is enough to tell the two
apart on the next pass:
the file matches the declaration unchanged
it matches what was last written updated -- the mesh changed its mind
it matches neither corrected -- somebody changed it here
The machine is put back either way, because holding it to what it was told is
the point. What changes is that it says so.
A digest rather than the content: the store is read on every reconcile and sits
beside the state on disk, and keeping every managed file twice would make it
grow with the size of the machine rather than with the number of resources.
Two questions that had been answered by one capability. graphical-session asks
whether a session is running now; seat asks whether one could ever run here.
Assignment needs the second -- a headless server can never have a display
server, a workstation with nothing installed yet can, and until now those
looked identical. The mesh would have assigned xorg to the server and found out
at apply time.
"seat" rather than "display", and the distinction matters most on a phone. An
Android device plainly has a screen and has no seat: nothing there is going to
take a DRM device and present an X or Wayland session on it. A capability
called display would answer yes and be useless. This one answers no, which is
true and useful.
Read from what the kernel reports about its connectors, which distinguishes the
two ways of not having one: a machine with a graphics card and nothing plugged
in is a different thing from a machine with no graphics at all, and somebody
deciding where a desktop goes wants to know which they are looking at.
Verified against this workstation: it reports card1-DP-1 and card1-DP-2, which
are the two monitors actually connected, and ignores the DisplayPort and HDMI
that are not -- and the writeback connector reporting "unknown", which counting
would have handed a seat to machines that have none.
The launcher already ran `host run`, and `run` on a machine with no identity
exited with an error. So a freshly installed host, sitting exactly as intended
waiting for somebody to bring it a token, would have counted three failed
starts and rolled back its own installation.
It waits now, and says what it is waiting for. That is the *hosted* state from
the lifecycle: the host is running, it has no identity, and there is nobody to
link to. Every machine passes through it.
An identity that exists and cannot be read is still a fault rather than a wait.
Treating that as "not enrolled yet" would leave a node sitting quietly for ever
while the mesh believes it is a member.
Also: a node now says it is there once a minute. Nothing but its name, because
anything more would be a report, and reports are rare where this is constant --
reading one as the other would make a quiet node look like a stale one. Not
published mandatory, unlike a report: losing one is nothing, the next is a
minute away, and the mesh reads a gap rather than counting arrivals.
Verified in the lab: a node was stopped and the mesh said "out of touch 4m",
then it was started and the mesh said "here" again, without anything else being
touched.
Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.
Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.
A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.
And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.
The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.
With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.
Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.
That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.
Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.
Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.
Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
Curve25519, which is what WireGuard uses. The private half never leaves the
machine and is written to a file of its own, so the interface configuration the
mesh composes can point at it without ever carrying it.
Separate from the identity keypair on purpose. One signs messages to the mesh
and the other encrypts traffic between nodes -- different things verified by
different parties at different times, and a key used for two purposes is one
rotation away from breaking the other.
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.
State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.
Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.
Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
Pushed a failing test in the last commit -- my own gate reported it and I read
the count rather than the result. The failure was real and worth having.
Adding the membership requirement to Load made Generate produce an identity that
Save would write and Load would then refuse. A file that cannot be read back is
the worst shape this could take: it is read back on the next start, on a machine
nobody is watching, and by then the token that could have fixed it is spent.
Save now refuses exactly what Load refuses, and writes nothing when it does. The
round-trip test covers the membership too, since that is the half that lets a
node come back on its own.
The loop the whole thing exists for: told, apply, report.
`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.
Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.
A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.
Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.
That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.
Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.
The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.
An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.
Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.
The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.
Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.
The host side of enrolment. It parses a token the control plane issued, dials
the broker, refuses anything but the pinned certificate, and generates an
Ed25519 keypair whose private half never leaves the machine.
Verified against a real LavinMQ serving a real certificate: the pin matched and
the node proceeded. Then against a second broker with a different certificate
on another port, which was refused -- with an error that says retrying will not
help, because it does not mean the network is down, it means the mesh was
substituted.
InsecureSkipVerify is set and that is the point rather than a weakening. At
bootstrap the broker is self-signed and reached at an address, so there is no
authority to trace and no name to match. Chain and hostname checks are replaced
with something stricter: this exact certificate or nothing, checked in
VerifyPeerCertificate, which runs before the handshake completes -- so nothing
is sent to the wrong broker. There is a test that counts the bytes an impostor
receives, and it is zero.
The token format is defined separately here and in the control plane, because
this binary requires nothing present and does not import it. They are held
together by a test on each side asserting the exact field names, so a rename
breaks both immediately rather than at enrolment on a real machine.
Two distinctions the identity file has to keep. A machine that never joined has
no identity, which is an ordinary state and not a fault. A machine whose
identity cannot be read is a different thing entirely, and must not take the
same path -- re-enrolling would discard the identity the mesh still believes and
need a person with a new token. Fault injection found the second case untested:
the corrupt-file test was passing on the parse check, so the read-error path had
nothing defending it. It does now.
An already-enrolled machine refuses to enrol again rather than quietly
acquiring a second identity.
What is not built is the link. Enrolment stops after verifying the broker and
generating the identity, having saved nothing, so it can be run again unchanged.
132 tests, plus 32 launcher and 9 rollback.
Five steps now instead of three. A sealed machine goes from bare to a container
runtime, a store, the inventory database, that database's schema applied by
mesh-control, and LavinMQ running and answering.
The database is called inventory rather than mesh. ADR 0008 grants a context
only what it exclusively owns and ADR 0006 says the mesh database names a thing
that will not exist -- so one database per context, and there is one context.
The broker is in the bundle because ADR 0006 now says it must be: the control
plane reaches a node only over the link, the link is the broker, so nothing can
provision the broker. Two images, which is the cost that record accepts.
Verified by reading the system rather than the report: inventory present and
mesh absent, the node table with its indexes, the migration row, lavinmqctl
answering, 5672 listening. Sealed confirmed both ways -- the internet times out,
the lab registry returns 200.
Then rebooted, which was the part worth doing rather than assuming. Everything
returned: docker from boot: enabled, both containers because this host creates
every container --restart unless-stopped, the schema intact in its volume. Three
reconciles before and one after all report no change.
One thing that reads as a success and was not: the first sealed check said the
machine could reach example.com. It was the test that was wrong -- a helper
script pasted arguments into a shell line, so a command with quotes was re-split
and ran on the workstation. The machine had been sealed the whole time. The
helper now requotes each argument.
The first three steps of the substrate bootstrap, run on a lab machine
confirmed to have no route out. It went from bare to a container runtime
installed and enabled, PostgreSQL running from an image pinned by digest, and
the control plane's database created inside it -- from the file the host
carries, with nothing to ask.
Second run changed nothing. `owned` lists all five afterwards, and `mesh` is in
the store.
It stops before the last two steps because there is no control plane yet: its
schema cannot be loaded and its image does not exist. The bundle says so rather
than naming something that cannot be applied.
Two things fixed on the way.
`make host BUNDLE=...` still swapped a file called substrate.lock, which the
per-system split had renamed months of decisions ago -- it now takes SYSTEM and
replaces that system's bundle. And the .lock files still cited ADR 0060, since
the renumbering pass only covered .md, .go, .ts and .sh.
One thing learned by it failing first: a directory the host creates is owned by
root, and a database inside a container runs as somebody else, so it could not
write and the container crash-looped. The store's data is a named volume now,
which lets the image set up its own ownership and outlives the container --
which is what you want for the thing holding the mesh's state.
Worth noting the failure was caught by the action's verify rather than by the
container step. `docker inspect` reported the container running because it was,
briefly, between restarts. Running is not working, and the thing that knew the
difference was the step that asked the database whether it would answer.
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.
Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
Two gaps.
The bundle's contents are per system even though its mechanism is not, so there
are now three: substrate-arch.lock, substrate-alpine.lock and
substrate-android.lock. All three are embedded and a host reads only the one it
was built for. Arch and Alpine remain placeholders -- the closure for a one-node
mesh is still research 011/012's open question, and inventing it here would be
worse than an honest placeholder.
Android's is not a placeholder. It says a partial host cannot raise a mesh and
why: every step of a bootstrap is a package, a container, or an action against
one, and those are exactly the shapes it refuses. So a partial host can JOIN a
mesh and cannot BE the first node. That belongs where somebody looking for the
android bundle will find it.
Also separated two things that were being conflated: "this system has no
bundle" and "this system was never built". Loading a bundle for debian is not
ErrEmpty, and the test asserts they differ.
0062 -- a host may be episodic. There is no way to keep a process running on an
ordinary Android device: init needs root, a foreground service can be killed
for memory. The answer is not to fight that. It is that being killed IS
disconnection, which ADR 0036 already made an ordinary situation -- and
everything the design does for a laptop that closes is what an episodic host
needs, at a shorter period. An authoritative local store, reconcile on start,
last-heard-from reported without an alarm.
So the gap closes by requiring less rather than building something. No keep-
alive, no Android daemon, no fighting the platform's process management.
Two consequences recorded rather than glossed. Last-heard-from is a much weaker
signal on an episodic host, so a healthy phone reads as a dead server unless
the reader knows which kind it is looking at. And a declaration may take a long
time to land, which makes 0058's separation of outstanding from failed
load-bearing rather than tidy.
Left open deliberately: how an episodic host is actually started, and -- first --
what an Android node is for. Building the start mechanism before deciding that
would be building it for nobody.
ADR 0060, built. `make hosts` produces mesh-host-arch, mesh-host-alpine and
mesh-host-android, each pinned to its system at link time.
The claim that "almost all of it is shared" held up. All 36 existing apply
tests pass unchanged -- the only edit was naming which system they run against,
which was previously implicit. What moved into internal/system is two appliers'
worth of code and the probes that go with them.
Each system's differences are real and needed re-deriving rather than
translating:
apk reports absence by EMPTY OUTPUT and exits zero either way, where pacman
exits non-zero. Reading apk's exit code the way pacman's is read reports every
package as installed. That is the single most dangerous difference between the
two and it is invisible until it bites.
OpenRC has no LoadState, so "the service does not exist" is read from its prose
rather than a field. Same distinction, different evidence -- and this is exactly
what an interface spanning both would have had to drop, which is why 0060
rejected one.
OpenRC has no is-enabled either. Boot state comes from the runlevel listing:
"does it start at boot" becomes "does it appear in rc-update show default".
Android is a partial host and that is the point. It implements file, directory
and action -- the shapes needing only a filesystem and a way to run something --
and refuses the other three by name, before anything is applied. Its unreachable
appliers return ErrUnsupported rather than a zero value, so "unreachable" fails
loudly if it stops being true.
A host also confirms it is on the machine it was built for, once, at the start.
The alpine host on this Arch machine says "this machine is not Alpine" instead
of failing later inside a package manager that is not there. And a host built
without -X main.builtFor refuses everything, naming the hosts that exist.
Two test problems found by injecting faults. One injection did not compile, so
the check now reports that separately from a pass. The other passed with the
behaviour removed: the missing-service assertion matched "does not exist", which
the FALL-THROUGH error also contains because it echoes the raw output. It now
asserts the diagnosis, which only the correct branch produces.
Verified with the real binaries: android refuses a package naming what it does
support; alpine on Arch refuses the machine; arch applies and is idempotent; a
system-less build refuses everything.
Jochen: "I thought we did not want to run the host under a systemd/openrc/init
loop, but instead had our own host-init program?" -- and that was right. I had
moved the give-up logic out of unit files and left RESTART in them, with the
launcher exec'ing the host and disappearing. So init still decided when the
host came back, which is the arrangement 0061 exists to remove.
The launcher now stays and supervises: starts the host as a child, waits,
decides. Init is asked for one thing, run this at boot. There is an OpenRC
script beside the systemd unit now, four lines each, which is the point --
a second init is transcription rather than a port.
The cost of not exec'ing is signals. A supervisor that exits while its child
runs leaves the host to be killed rather than to stop, and an apply interrupted
that way is the half-configured machine this project is about. So SIGTERM is
trapped, passed down, and waited on.
Two bugs, both found by the tests rather than by review:
A clean exit was counted as a failure. The host exits cleanly to stand aside
for a new binary after an upgrade (0057), so a host that upgraded itself three
times rolled itself back having worked perfectly every time. The counter now
counts CONSECUTIVE FAILURES, incremented after the wait rather than before the
start.
And when rolling back I reset the counter file but not the variable, so the
next failure counted from the old value -- the rolled-back version got one
attempt instead of three.
Also: the host now clears the counter when it completes a reconcile, at the
same moment it records known-good and for the same reason. Without it the count
only climbs, and a node up for months rolls itself back on its third ordinary
restart -- a healthy machine undone by its own recovery.
One test expectation was tightened rather than fixed: "resets the counter after
rolling back" asserted exactly 0, which was true only under the old
count-before-start semantics. It now asserts the property -- below the limit --
since 1 is correct after a rollback plus one failure.
32 launcher tests, all confirmed to bite.
Two gaps found by testing podman rather than reasoning about it.
The service shape could not say "starts at boot". It ran `systemctl start`, so
`service: docker.service, running` started docker now and it would not come
back after a reboot unless something else had enabled it. A declaration that
reports success and stops being true at the next power cut.
`boot: enabled|disabled` is now a separate field, not a fourth value of
`state`, because the two are orthogonal: a unit can be enabled and stopped (it
returns at boot) or disabled and running (started by hand, gone after one).
Absent means the host asserts nothing, so a machine whose operator enabled
something is not silently disabled by a declaration that never mentioned it.
Boot state is made true BEFORE the unit is started. When an apply fails part
way, enabled-and-stopped comes back at the next boot and running-and-disabled
does not, so the more durable half goes first.
`is-enabled` has the same trap as `is-active` had. Its exit code is non-zero
for nearly everything, and `static` is neither enabled nor disabled -- the unit
has no install section and CANNOT be enabled. Reading it as "disabled" would
have the host try, fail, and blame the wrong thing, which is the same shape as
reading a missing unit as "stopped".
The container applier no longer calls `docker` literally. Verified on this
machine against podman 6.1.0:
docker info --format '{{.ServerVersion}}' -> 29.7.2
podman info --format '{{.ServerVersion}}' -> Error: can't evaluate field
ServerVersion
podman info --format '{{.Version.Version}}' -> 6.1.0
So one probe cannot find both, and a host using docker's would report a machine
running podman as having no container runtime at all. Everything else IS
compatible -- run, rm -f, and docker's own Go template syntax for reading state
and labels all work unchanged on podman, confirmed by running them. That is why
this is a two-entry lookup rather than an interface: only the probe differs.
Detected rather than declared, because adoption keeps what the machine already
has (research 012), which hardcoding one runtime contradicts.
A machine with neither now says so, naming both: "docker: command not found" on
a machine deliberately running podman sends the reader after the wrong thing.
Verified end to end against real docker (container created, running, labelled)
and against an empty PATH (refused, naming both runtimes).
Two injections per behaviour, all confirmed to bite. One injection produced a
build failure that my check read as "no bite" for the third time, so the check
now distinguishes them.
ADR 0061. Recovery was the most systemd-specific part of the host, and it is
the part that must work on a machine where nothing else does -- which made
unit-file syntax a poor place for it, because syntax cannot be tested and the
one time it runs is the one time nobody can afford it wrong.
So StartLimitBurst and OnFailure move into a launcher script that init starts
instead of the host. The unit drops to start-at-boot and restart-on-exit, which
OpenRC, runit, s6 and an Android init.rc can all express. Everything 0059
decided is kept: two watchdogs, roll back once, recovery is local, the rollback
shares no code with the host.
The counter is the whole mechanism, so it is what the tests are mostly about.
Three real problems came out of writing them:
A counter file holding "1 2" became "12" -- `tr -d [:space:]` concatenates
rather than rejecting -- which is past the limit, so a HEALTHY node rolled
itself back. Now it reads the first field and insists on a plain integer.
The corrupt-counter test used "not-a-number", which shell arithmetic happens to
evaluate to 0, so it passed with the guard removed and proved nothing. Replaced
with values that discriminate: "5x" errors under set -e and kills the launcher,
and "0x10" is read as HEX 16 -- past the limit, so again a healthy node rolls
back.
And the test harness itself was wrong. With `set -e` and a bare launcher call,
removing a guard killed the script at the first corrupt case and silently
skipped everything after -- reporting a full pass over tests that never ran.
Every launcher call now records its failure instead of aborting. Same class as
the placebo assertion found last time, and the reason to keep injecting faults
rather than trusting green.
Both scripts run in `make check`. 27 launcher tests, 9 rollback tests, all
confirmed to bite.
ADR 0059's recovery path: the pieces that run when the host will not start.
internal/upgrade -- two facts, neither of them the host judging its health.
Whether the executable this process started from has been replaced on disk, and
which version last completed a reconcile.
The first design was wrong and the tests caught it, not review. It asked
/proc/self/exe whether it was marked deleted. That is Linux procfs behaviour
rather than a fact about files, and it catches only unlink -- a binary swapped
by rename onto the same path reads as untouched, which is exactly what a
package manager does. Now the identity is captured at start and compared later:
no procfs, and neither case missed.
known-good is one bare line. The reader is a shell script on a machine where
the host is failing to start, so it must not need a parser to be present and
working. Written only after a clean apply, which is the whole claim -- not
health, because a disconnected node is ordinary and a failing resource is the
machine's problem rather than the binary's.
packaging/ -- the unit, the rollback unit, and the rollback script. The script
shares no code with the host and calls none of it: a binary that cannot start
cannot be its own recovery. POSIX sh, nothing that has to be installed. The
unit carries Restart=always with a comment saying why on-failure would break
every upgrade.
Both are tested and both sets of tests were confirmed to bite. Injecting five
faults broke exactly the intended tests -- except one, and chasing why it did
not found a placebo assertion I had written: `check "exits zero" ... "0" "0"`
compares a literal to itself and can never fail. Replaced with the real exit
code, after which the injection bites.
Also caught: an injection that produced a build failure rather than a test
failure, which my grep read as "no failure". Re-run so it compiled, and the
test did bite.
The script test runs in `make check`, so it is a gate rather than something
that was run once.
Verified against the real binary: known-good is written beside the store after
a clean apply and is NOT written after a failed one.
Jochen asked why we don't simply have dedicated structs. We should, and the
flat struct was me extending an existing pattern rather than questioning it.
Before: one Resource struct carrying path, content, mode, unit, state, package,
image, name, env, ports, volumes, args, command, verify and in. Because a file
and a container shared it, nothing stopped {"type":"file","image":"postgres"},
so a `uses` map listed which fields each kind was allowed to carry -- a second
place to keep current, and the kind nobody updates is the one that silently
accepts a field the host will never read.
Now: Directory, File, Service, Package, Container and Action are separate
structs behind a Resource interface. File has no Image field, so the mistake is
not detected -- it is unrepresentable. Adding a field to a kind is the whole of
adding it; there is nowhere else that has to agree.
Parsing is two passes: read the envelope and each resource's raw bytes, peek at
"type" to choose the struct, then decode into it. Peeking is lenient on purpose
-- reading strictly there would report an unknown field before knowing which
fields are known.
Unknown fields are found by comparing the JSON keys against the struct's own
json tags rather than by catching the decoder's error. The decoder stops at the
first unknown field, and RefusalError promises every problem at once: a caller
fixing one field at a time learns the next only by running again. Caught by
testing the refactor against a real declaration -- a container carrying both
`unit` and `mode` reported only one of them.
apply.go switches on the concrete type instead of a string, so a new kind that
has no applier is a compile error rather than a runtime default branch.
No behaviour change otherwise. All existing tests pass unmodified except two
that reached for fields the interface no longer exposes.
The three shapes the substrate bootstrap needs and the host did not have. Until
now tier 1 could not be raised at all -- step 0 is a package, step 1 a
container, steps 2 and 3 actions -- so every line of the tier 1 and 2 designs
was unbuildable.
package -- present, never upgraded, never uninstalled. Removal is "forgotten",
not "removed": the host cannot know what else needs the package, uninstalling a
container runtime because a declaration changed would stop every container on
the node, and the machine may have had it before the mesh saw it. Reporting it
removed would claim an effect the host declined to have.
container -- identified by a label carrying a digest of the declaration that
made it. Comparing every field the runtime reports cannot be done reliably: a
runtime normalises, defaults and reorders what it is given, and that is
indistinguishable from real drift. There is no in-place update; a container's
configuration is fixed at creation, so any change is a replacement, and saying
so beats a partial update that leaves the running thing half-declared. This is
the one shape the host removes, because it is the one the host created.
action -- bundle-only, per ADR 0047. Verify is mandatory and does double duty:
it is the idempotency check as well as the read-back. The host does not know
what a database is, so "is it already there" is a question only the declaration
can ask. `in` runs the action inside a named container, which steps 2 and 3
need.
Parse now refuses actions; ParseTrusted permits them. The safe path is the
default and the permissive one has to be named. The bundle and a local file
handed to a root process use ParseTrusted; the link will use Parse.
Also replaced the per-type "fields this type ignores" check with a field-set
diff stated as what each type USES. The negative form needs every type revisited
whenever a field is added, and the one nobody revisits silently accepts a field
it will never read.
Images must be pinned by digest (ADR 0046). A bundle naming a tag pins nothing.
Verified against a real machine, not only fakes: an action ran and was
idempotent on the second apply; an action that exits zero and satisfies nothing
fails the apply; a real container was created, labelled, replaced when its
declaration changed, exec'd into, and removed; a real package query round-
tripped. Each new test was also confirmed to fail on an injected fault -- five
injections, each breaking exactly its own test.
One existing test changed: a vanished unit is now reported "forgotten" rather
than "removed", which is what actually happened.
novox/hq ADR 0038: one behaviour, two sources of declaration. This is the source
that does not need a mesh — the first node's path.
The bundle is embedded in the binary rather than shipped beside it, because
"copy it onto a machine and run it is the whole installation" stops being true
the moment a second file has to arrive with it. `make host BUNDLE=...` builds a
host carrying one; `mesh-host reconcile` applies it; `mesh-host bundle` shows it.
A default build carries nothing and REFUSES to reconcile, saying why. A host
that applied nothing and reported success would look exactly like one that
raised a first node, and the difference would surface later as a mesh that never
came up with nothing to point at.
Proved on a sealed machine: no route out, no name resolution, one binary copied
on, and it configured itself from what it carried. Idempotent on the second run.
One bug found by running rather than reasoning, and it is a shape worth naming:
`mesh-host bundle` validated the carried bundle through a path that strips
comments, while `reconcile` handed the raw bytes to the parser. So the command
whose whole job is to check the bundle said yes, and the command that uses it
said no. Two paths to one artefact, disagreeing. There is one path now, and a
test asserts that what validates is what is applied.
What this does NOT prove is stated in the README rather than left implied: the
claim under stage 2 is that one host can raise the substrate alone, and the
substrate is four container services. There is no container type, because a
container needs an image and where images come from is open; what belongs in a
substrate is not known, because the closure for a one-node mesh is what research
011 and 012 exist to answer; and the machine used to test this cannot install a
container runtime through a sealed network.
The mechanism is finished. The claim is not, and shipping a host that claimed a
substrate it has never raised would be the fault this whole project is about.
65 tests.
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.
It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.
It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.
The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.
Three faults found by running inside a raised machine rather than by reasoning:
A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.
Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.
And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.
30 new tests, 55 in total.
Tier 0's first slice, per novox/hq 03-DESIGN/01-to-be/05-the-node-host.md. It
applies nothing, connects to nothing, listens on nothing. 2.9 MB, static, no
dynamic dependencies: copy it onto a machine and run it is the whole install,
which is the property ADR 0041 rests on.
A capability is detected, never assumed. Every detector runs something that only
succeeds if the thing FUNCTIONS — the daemon is asked for its version, the
package database is queried, the firewall is asked to list a ruleset, which
needs the privilege as well as the tool. 04-ISSUES/007 is the fault this
prevents: a client on disk with its daemon down looks exactly like a working
runtime, and a node assigned work on that basis fails when the work arrives.
Every verdict carries the reason and the method. A capability reported absent
with no reason is the same fault in a new place: something nobody can act on.
Two bugs found by running rather than reasoning, both silent:
systemctl is-system-running exits non-zero for every state except `running` —
including `degraded`, which means units failed and the init is emphatically
there. Reading the exit code reported NO service manager on a machine whose init
it was. That is 007 in the mirror, and both directions place work wrongly. A
verdict now reads what a tool says about itself, not only how it exited.
And `mesh-host inventory --json` printed text: the standard library stops
parsing at the first non-flag argument, so the flag sat unread and the command
exited 0 having ignored what was asked. The parser now takes the subcommand off
the front, and a stray or mistyped argument is refused rather than dropped.
Detection deliberately does NOT follow ADR 0008. That rule governs applying
state, where a failed step means the machine is not what was asked for. A failed
probe is a finding — "absent, because the probe failed" — and aborting would
replace one legible absence with total ignorance of the rest.
25 tests: structure and logic with a fake runner, and the same detectors against
this machine, because a test that fakes the system under detection asserts only
that the fake behaves as expected.