An operator ran `mesh-host reconcile` on an adopted control-node with twelve
modules assigned. It applied the bundle the host carries — the genesis
declaration, foundation only, converged: recreated the store, failed on the
broker's held port, wrote the converged base filter and started its service,
and stopped at the first failing action. The filter closed the machine for
forty-five minutes. The host reported the node adopted in every report, the
declaration said converged, and nothing compared the two; nothing was printed
before acting (hq issue 104).
The host now records the node's mode — from every declaration the mesh sends,
and at genesis from what the operator said — and refuses, at the point of
application, a declaration that says the other mode, naming both and the act
that changes it. Only a declaration the link delivers, signed, changes the
mode: that is how `converge` and `adopt` arrive, so the flip still works and
nothing else can do it. Genesis marks the bundle consumed, with the digest of
what it applied, so `reconcile` holds a node the mesh has spoken to against
what the mesh last said and never the bundle, and refuses the carried bytes
when they are not what genesis applied. A file is refused when it is not what
the mesh last said: a declaration carries no sequence and no issued-at, so the
host cannot tell older from newer, and says so. Both commands print what they
would change — a hold, a removal, an action named as one — before touching
anything, and --dry-run is that list and nothing more.
031: a window of unacknowledged declarations is drained to the newest; the
rest are set aside and reported as superseded. 035: a file resource may say
create-once — written when absent, kept untouched when present (ADR 0087).
054: the bundle installs nftables and loads a base ruleset before the store
and broker, in the table the filter module later replaces (ADR 0088).
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Found by being asked whether processes and containers handle environment the
same way. They do not, and the difference is not cosmetic.
Docker passes --env through literally. A unit file reads three things out of a
value that nothing else does, and a module's environment routinely contains all
three because a generated password is arbitrary bytes:
- % begins a specifier. %H is the hostname. A password containing one is
silently replaced, and it fails later as an authentication error nobody can
explain by reading the declaration.
- whitespace separates assignments. Unquoted, K=a b sets K to "a" and reads
"b" as another assignment.
- a newline ends the line, and what follows is read as a unit DIRECTIVE.
The first two are escaped: quoted, with quotes and backslashes escaped and
percent doubled. The third cannot be — a unit's environment has no way to carry
a line break — so it is refused in validation, near whoever wrote it. Without
that, an environment value could write ExecStart= and have the machine run
something nobody declared.
Ordinary awkward values stay accepted, because refusing those too would leave a
module unable to hold a generated password.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first cut of this added a `daemon` for the long-running case alone. That
would have meant a new vocabulary entry for each of the others — a scheduled
task, a run-once migration, a health check — when they are one thing run at
different cadences. That is a field, not four entries in a vocabulary where every
entry widens what a compromised control plane can express.
So it mirrors a container exactly, because it IS a container's twin: the same
intent, hosted by the machine's own supervisor instead of a runtime. Stays up,
runs once, or runs on a schedule.
Tools, hooks and event consumers are not further modes. They are loaded by a tool
host, which is itself a process that stays up — so the generic case already
covers them, which is the test of whether it is generic.
A scheduled process gets a timer and a unit that finishes; a long-running one
gets a unit that is restarted when it exits. Getting that wrong either way is a
second copy running continuously between fires, or a schedule that never fires.
The modes are exclusive and validation says so near the author: something that
runs once does not run on a schedule, and something not running between fires
cannot be restarted when a file changes.
A missed fire happens when the machine comes back rather than being skipped,
which is the difference between a machine that was down and a schedule that
quietly stopped.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mechanism was leaking into every module. Code of one's own meant a container
and therefore an image; a script meant a service and a unit somebody else had to
install. One intent — run this and keep it running — expressed two unrelated
ways, with the hosting chosen before anything could be declared.
A daemon names a bundle and a command. The host fetches it, refuses it unless it
hashes to what was declared, unpacks it where the mesh keeps such things, writes
the unit and puts it in the state asked for. The unit is the mesh's, generated
whole and saying so, because an edit that survives until the next declaration and
then vanishes is worse than one that is refused.
Its identity is the bytes AND how it is run: two daemons from one bundle
differing only in their command are different daemons, and tracking the digest
alone would call the second unchanged and leave the first running. The unit is
rendered deterministically for the same reason — environment from a map would be
written in Go's iteration order, so every apply would see a different unit and
restart an unchanged daemon for ever.
restart-on is honoured as a service's is: a running process does not re-read its
configuration, so replacing a file and finding the daemon already up leaves the
machine behaving as before while every check passes.
A full-host shape, not a portable one: it needs a process supervisor to install
into. It does NOT need a container runtime, which is the point.
Two guards caught this properly and both were updated deliberately rather than
silenced: the vocabulary count, which exists because every addition widens what a
compromised control plane can express, and the shape test that catches a kind the
language has and a host cannot apply — added after `network` did exactly that.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two faults that both reported success while being wrong, found while proving
the firewall module actually delivers.
A unit whose job is to apply something and exit — load a rule set, set a
sysctl — is inactive the instant it succeeds. Reading that as stopped made it
permanently unsatisfiable: the host started it, it worked, the host read back
stopped and reported the machine as not doing what it was told, on every apply,
for ever, with the rules correctly in place the whole time. That is what the
firewall has been doing on every machine it was assigned to, and why the
four-machine bed was red.
And a container took its identity from its own fields, not from the files it
reads. A file written in an earlier apply — or before the container declared it
as a dependency — left a process holding a credential the mesh had already
replaced, with everything reporting success (novox/hq 04-ISSUES/045). What a
container reads is now part of what it is, so the comparison is a standing one
rather than a tripwire that fires during one apply and never again.
A manifest digest is assigned by a registry on push, so insisting on one meant a
registry had to exist before the thing that lets a mesh have a registry could
start — a dependency the pinning rule created by accident, not a pin. The mesh's
own control plane is built from source and lives in no public registry.
A bare sha256:... names an image the machine already holds, by the digest of its
own configuration: immutable and unforgeable in exactly the way the rule asks
for. Absent, it says so plainly rather than failing at a pull nothing serves.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A schedule: container (ADR 0053) is installed as present state and never run at apply — the
Scheduler fires it later on its cadence. But a service or run-once container only gets its
image as a side effect of docker run, so a scheduled step's image was not pulled until its
first scheduled fire: absent from the node right after a successful apply, so the first run
paid the whole pull latency and tooling that expects the image present after apply found it
missing.
applyContainer now probes the runtime and ensures the pinned image present for a scheduled
step before recording it. A new ensureImage helper inspects the image and pulls it only if
absent, then reads back (ADR 0018). Ensuring an image is not running it: no docker run fires
the container, so the no-run invariant of ADR 0053 holds. The runtime probe, previously
skipped for a schedule, now runs because a pull needs it — the schedule.go comment is updated
to match.
Tests: the install-does-not-run test is extended to allow the image-ensure while asserting no
fire and no needless pull; a new test applies a scheduled container whose image is absent and
asserts it is pulled and still not started. go build, go vet, go test ./... all pass.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The recurring twin of run-once, one modifier over: a container marked
schedule: "<cron>" is run to completion on its cadence, not started as a
service and not run once as a gate.
The gating rule is deliberately reversed. Installing a schedule records it
as present state and reports the node current at once (applySchedule) --
it never runs the container and does not gate what follows. A Scheduler,
held for the life of the daemon and re-established from each applied
declaration (the declaration is the source of truth, ADR 0018), fires the
container off an injected clock. A run that exits non-zero is logged and
never fails the apply or flips the node's state, because it happens
outside the apply and the store entirely. Runs never stack: a run still
going when the next is due is skipped, not started as a second copy.
No new host shape and no new action -- schedule is a string on the
container the host already has, and the host process runs the container
itself rather than installing a system timer (the rejected option 1). A
minimal five-field cron (declaration/cron.go) validates on arrival and
computes the next due minute; time is injected so the scheduler is tested
without the wall clock.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module can declare state but not a step that runs at first boot. This adds
`run-once: true` to the container shape: the host runs it in the foreground,
requires it to exit 0, and records that it did — as the digest of the
declaration, so a re-apply does not re-run it unless the declaration changed.
Because the declaration is applied in order and a failed run-once step gates the
apply the way a failed action does, whatever is declared after the step starts
only once it has completed. That is how "before the broker starts" is enforced,
with no dependency graph the host must resolve (ADR 0005): the step is declared
first, and the container that needs it is never reached until it is done.
No new host shape and no arbitrary host command — a run-once container is
strictly less powerful than an action. Validation refuses run-once with
restart-on (contradictory lifecycles). Six unit tests; go test ./... green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The tenth shape (novox/hq ADR 0051). Shared, pre-existing data — a media
library, a download spool several modules use — is the operator's, not
the mesh's. A `directory` resource is the host's own: it creates it,
chowns it, sets its mode and removes it when empty. An access is the
opposite on every axis.
Add the `access` type to the vocabulary. Its applier confirms the path is
present and changes nothing: it does not create, chown, reconcile or set
a mode. Absent is refused clearly — the operator must provide it — rather
than created, because a bind mount whose source is missing is made as
root by the container runtime with the wrong ownership (04-ISSUES/026).
Undeclaring an access forgets the record and never touches the path,
which is the data loss ADR 0030 prevents, on a directory the mesh never
made.
Full hosts speak it (it gates a bind mount, which needs the container
runtime); the vocabulary guard test records the decision that made it the
tenth shape. Unit tests cover present, absent-refused, and
undeclared-left-alone.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container reads a mounted file once, at start; its spec (image, env, volumes)
does not include a mounted file's content, so a settings change that re-renders the
file left the running process holding the old value while every check passed. Give
Container the restart-on field a Service already has, and recreate the container
when a named resource changed this pass. Unit-tested (recreated on change, left
alone otherwise) and proven in the mesh-lab: a running grafana runtime picked up a
token change on the next push.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A directory a module mounts into its container belongs to whoever runs
inside — grafana's 472, redis's 999, www-data's 33 — and none of those
has a row in the machine's passwd. Owner-by-name refused them all,
which looked principled and meant every module whose container drops
privileges could not own its own data.
The lab showed both coats of it in one run: the store's config file was
unreadable to the store, restarting forever on permission denied, and
the forge could not traverse into the 0700 root-owned directory that
held its files — a directory that had only become root-owned when
declaring it fixed 04-ISSUES/026, because Docker used to create it
0755. A fix that tightens ownership without a way to say whose it
should be moves the fault, not removes it.
"uid:gid" and bare "uid" are numeric and chowned as given; a name still
resolves as before, and a name with a colon is refused rather than
half-read.
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.
The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.
What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.
Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
Two gaps found by writing the first real module's manifest rather than
by reasoning about one. Both are fields on existing shapes, so the
vocabulary is still nine.
**env-file on a container.** A declaration reaches a node over the
broker and `env` is plain text in it, so a password there is a password
the broker sees — the transitive trust refused everywhere else. A sealed
file arrives unreadable, the host writes it, the runtime reads it. It is
also simply how third-party software takes credentials: nothing shipping
in a container will read a path the mesh invented, and every one of them
reads its environment.
**secrets in a file's content.** A program wanting its token inside a
JSON document cannot be handed a file that is entirely a token, and the
mesh cannot compose the document because it discarded the value. So the
module supplies the document with `${secret:name}` in it, the mesh
delivers the value sealed, and the host is the only thing that ever
holds both.
Substitution is textual and the host learns no formats. Deliberate: a
mechanism that understood JSON would be asked to understand YAML next,
and then INI, which is how the arrangement this replaces became
something nobody could hold in their head. The module knows its own
format because it wrote the rest of the file. The sharp edge is stated
rather than left to be discovered — a value containing a quote is not
escaped for whatever surrounds it.
Refused in both directions, because both are somebody being wrong about
where a credential is: a placeholder with nothing to fill it would write
`${secret:x}` into a config file, and a secret the content never uses
means somebody believes a credential is in a file where it is not.
A file that carries one is 0600 unless the module said otherwise.
Found by asking what the conversion needs, and it is the one failure in
this system that cannot be undone.
Unassigning a module made its directory an orphan, and an orphan
directory was deleted with everything under it — os.RemoveAll — while
the report said "removed". A database's files, a mail spool, somebody's
uploads. Reproduced before fixing: assign a module, let a service write
into its directory, unassign the module, and the file is gone.
Now a directory that still holds something is kept and said so, naming
how many items are in it.
What makes that safe rather than merely cautious is the removal order,
which was already right. Everything the mesh puts in a directory is
itself a declared resource, and orphans are removed in reverse
declaration order — so what the mesh wrote is already gone by the time
the directory is reached. Anything still there was put there by
something else, which is the definition of data.
It is the host's own line applied to the one shape where getting it
wrong does not recover: it removes what it made and leaves what it
merely configured. An empty directory is what it made; a full one is
not, and an empty one is still removed so nothing accumulates.
Files are unchanged. A declared file is the mesh's own, and losing a
config file is not the failure this is about.
novox/hq ADR 0029, and work breakdown 1.3. A module of several
containers had no way to let them reach each other by name: a container
declaration could join a network and nothing could create one.
An action was the obvious alternative and is refused on removal —
"an action has no footprint the host can undo", so a network made that
way outlives every module that is ever unassigned, and the mesh cannot
tell. A resource the mesh can create and never clean up is one it should
not create.
A name and nothing else. Not a driver, a subnet or a gateway: each is
something a module would have to know about the machine it lands on, and
a module naming a subnet collides with whatever else chose the same one.
It needs no new ordering rule. Resources apply in declaration order and
orphans are removed in reverse, so a network written before the
containers that join it is created first and removed last — after they
are gone. A runtime refusing to remove one still in use is reported
rather than swallowed, because that means something undeclared is
holding it.
The vocabulary guard fired on the change, as designed, and now names the
record instead of a number: nine shapes, with the argument beside the
count.
Creation reads back rather than trusting an exit status (ADR 0018): a
runtime that reports success and made nothing leaves every container
that joins it failing to start, one step from the cause.
Half of novox/hq work breakdown 1.3, and it needed no change: the apply
loop walks d.Resources and sorts nothing, so a module that needs one
thing before another says so by writing it first.
Asserted because it is the kind of property a later change breaks
silently. Sorting the resources for any good reason at all — by type,
by identity, for a tidier report — would still pass every other test in
this package.
It is sequence, not readiness. A container started is not a container
ready, and nothing here waits: what depends on something being usable
retries, which is what both example provisioners do and is the more
robust answer anyway, because a dependency can restart long after
everything was applied.
Two mistakes worth keeping in the test's own comments. The first
version stubbed the runner to always succeed, so verify passed, every
action counted as already done, and nothing ran — the assertion was
measuring an empty list. The second declared the actions over the link,
which refuses them: only a bundle may carry an action (ADR 0005).
From auditing the decision records: of 28, only 12 were named by any
test, so "which decisions are defended" could not be answered without
reading everything. ADR 0017 says a test names the decision it defends —
that rule was itself unenforced.
Most of the gap was citation, not coverage. Drift detection was tested
in several places without naming ADR 0011; the archive refusal without
naming 0012; forged declarations without naming 0002. Named now, so the
question is answerable by grep.
The bundle was the real gap: nothing tested substrate-first-node.lock at
all. It is what a machine becomes when there is no mesh to ask — the one
declaration applied with nothing to verify it against — and it was
edited by hand and read by nothing but a running host.
Two tests now assert what it carries: exactly postgres, lavinmq and the
control plane. That defends ADR 0028, which removed the object store
from the substrate after it had been a member for months on the strength
of "it cannot grant itself a bucket" — true, and the answer to only half
the test. Nothing counted what the bundle held.
Fault-injected, and the first attempt did not bite: the injection landed
on a comment line, which stripComments discards. Injecting into the
image field fails as it should.
The commit before this said "told where to resolve names" and passed --dns,
which is not what it ended up doing. This is that correction: a container is
given the names themselves, written into its own hosts file by the runtime.
The reason for the change is the decision the mesh already made about names — a
file rather than a resolver, because it works on every runtime, needs no
package and has no failure mode of its own. Passing a resolver address would
have required a resolver to exist, which at that point none did.
A resolver is coming, for the case a file genuinely cannot express: a service
named under a machine, postgres.novox.internal, where the wildcard cannot be
enumerated in advance. When it arrives it will need this field back under its
own name. It is not being kept in the meantime — a field nothing fills is a
field nobody can trust, and the vocabulary is asserted by a count for exactly
that reason.
A container does not inherit the machine's names. It gets its own /etc/hosts
holding its own hostname, and a runtime rewrites resolv.conf — so every
internal name the mesh wrote for that machine is invisible to what the machine
is running.
That was hit for real, in the lab: a database client on one node could not
resolve another node, on a mesh where both names were correct and present on
both machines. It was worked around by resolving on the host and passing an
address, which is the kind of workaround that should not be needed twice.
A field on an existing shape, not a ninth shape — the vocabulary is still the
eight the count asserts.
Per container rather than by editing the machine's resolver configuration: that
file belongs to something else on most machines, and a host that edited it
would be fighting whatever owns it on every boot — the fault this host exists
to avoid, in the place it would be hardest to see.
A container told nothing is run exactly as before. Most containers should
resolve whatever the machine resolves, and passing an empty flag would be a
change of behaviour dressed up as a default.
The previous commit continued past every failure, and the lab found the
cost immediately: the bootstrap's store-readiness gate failed, the apply
carried on and started the broker and control plane against a machine
that was not ready, and the database still initialising was shut down.
An action is the only shape whose purpose is to make something true
before the next thing needs it — which is why it is the only one with a
verify. The bootstrap is a row of them. Everything else is independent
state, and stopping there is what made one broken module hold a whole
machine hostage.
The report says which happened: "these things failed" and "these things
failed and the rest was never tried" are different machines.