The controller is the way in for tool calls (novox/hq ADR 0095); the first user list must say so,
or a mesh raised from nothing refuses its own first ask. Kept in step with the controller's
composition by the test that reads this file.
The mesh runs on the seat's bus alone (novox/hq ADR 0131, design 28 task 5.5). The host's old
dialling and enrolment paths are deleted with the switch that chose between them; a membership or
a token naming another bus is refused before anything is sent, rather than dialled on a transport
that no longer exists.
The rescue path: an operator writes the membership file by hand on a machine no
bus can reach — rotated while it still held the old password — and restarts the
host. Same file, same check as the declaration path.
Two things the controller's own composition now derives and the installer's
carried list did not: the delivery subjects of the controller's consumers, and
JetStream enabled on the account — without which the first bound consumer is
refused. Found live on a mesh moved rather than raised; a fresh genesis would
have met both at first start.
The mesh names a machine's bus user node.<name>; the host dialled with the bare
name and every machine was refused the first time it reached the server:
"authentication error - User". Same string as the composed user list, or nothing
connects.
Until now a membership — bus address, fingerprint, password — was written once,
at enrolment, and nothing ever rewrote it, so a machine already enrolled could not
be moved to another bus at all (novox/hq design 28, task 5.2). The mesh now
delivers a membership for the new bus as a sealed file in the declaration, like
any secret; the host reads it after the declaration has applied, saves it as its
identity, and exits cleanly so the service manager restarts it dialling the bus it
names. The same path a machine takes after a reboot — so nothing new has to be
right for it to work. A membership carries its transport; empty means the bus the
mesh ran on before, so nothing written earlier reads as unset.
The same twelve steps, with the difference that matters: the mesh composes its own user
list and at genesis there is none, so this carries the first one — the controller's
account at a bootstrap password, rotated with the store's and replaced by the
controller's own composition from its first start.
The server's settings and the user list are separate files in one directory. Separate
because the settings belong to whoever raises the server and the users belong to the
mesh; in one directory of necessity, because an include path resolves relative to the
including file's own directory, so an absolute one sends the server looking underneath
that directory and it refuses to start.
No `verify` in the TLS block. That makes the server demand a client certificate and
nothing in the mesh presents one — a host pins this server's exact certificate and
authenticates with a password.
The controller's permissions here are checked against what the controller derives, by a
test in its own repository reading this file. They are two statements of one fact, and a
template that granted less than the controller needs would produce a mesh that comes up
and is refused on its first act.
The last of the host's link that still named a transport. `Asking` is one
enrolment conversation — a connection made with the token, a question asked, and
an answer waited for — and it is its own seam rather than part of `Link` because
almost nothing about it is the same: the credential is a one-time secret, there
is no declaration to hear, and a node that fails here is not in the mesh at all,
where a node that fails in `Link` has merely lost touch with one it belongs to.
`Enrol`'s thirteen arguments became an `Approach` — where, which certificate,
which bus — and the request it already had. The token says nothing about which
bus, and does not need to: every token names the one the mesh runs on today until
the rollout.
**The reply address is the whole of what changes on the new bus**, and it is
forced rather than preferred. Verified against a running server, both halves: the
answer reaches the node at the address its request carried in the payload, and
the transport's own reply field held something else entirely by the time the
consumer saw it — the consumer's ack address, exactly as design 25 §2 says. The
test asserts the field is *not* the node's inbox, so a future server that stopped
claiming it would fail this rather than let the reason quietly become folklore.
The inbox is under `_INBOX.enrol.<node>.`, which is exactly what the enrolling
user may subscribe and no wider, with a random tail per attempt: a reply left
over from an attempt that timed out is not the answer to this question, which is
what the correlation id does on the other transport. Subscribed before anything
is published, because a node that published first could miss an answer to a
question nobody was listening for.
The outbound half went behind `Bus` and a node's two statements stopped naming a
transport. This is the other half, and where the transport reached furthest: the
run loop selected on a channel of the client library's own delivery type, so
every part of holding a node in its mesh knew which bus it was on.
`Link` is dialling, hearing and saying in one interface, because dialling is
where the transport is chosen and choosing it twice is how one half of a node
ends up on a different bus from the other. `Declaration` has one way of being
done rather than two: a declaration set aside for a newer one is settled exactly
as an applied one is, on both buses, and the difference is a fact the report
carries.
Four things this settled.
**The host declares nothing on the new bus.** On the bus the mesh has it declares
its own queue, because a queue that is not there means a node that hears
nothing. Here it binds to a consumer the mesh made when the node enrolled, and a
missing one is said as the mesh's to answer rather than quietly created with
whatever this client happens to default to.
**The pin is easier here than in the tool runtime, not harder.** The Go client
takes a *tls.Config, so the same PinnedConfig with the same VerifyPeerCertificate
does the work — the subject-alternative-name constraint recorded against the
runtime's client is that client's, because it takes PEM strings with no verify
hook. A host checks the fingerprint and nothing else.
**Binding needs the subject as well as the consumer.** An empty subject is
refused rather than taken to mean "whatever that consumer delivers", which the
server said plainly and only when asked.
**Reconnection stays the caller's.** Hold already decides when to try again and
how long to wait; a client reconnecting underneath it would make that reasoning
a duplicate of the library's.
The drain keeps its live half and loses its catch-up half, as it said it would:
verified that three declarations pushed to an absent node leave one on the
stream, and it is the newest.
One test-harness lesson worth the comment it got: delete-then-add is not a reset.
A test that did that inherited the previous test's messages, and the symptom was
a declaration counted as delivered twice — which reads as a redelivery bug in the
code under test rather than as a dirty stream.
Put back by hand, it was found afresh and its copy became the hold's original, so a machine that
kept restoring it kept growing copies and lost which one was first. The first original now stays
the record's, content that differs is kept once beside it, and the note says a rollback means
unassigning the private network. Nothing is removed unless it is <wireguard dir>/<iface>.conf,
not a path the mesh writes, and not a link, which would leave the key-bearing target behind.
Kept on disk it was the take's fallback; once the mesh's interface is up in its place and a peer
has handshaken with it, it is an unmaintained way back onto the network, held for ever. It is now
removed from where its unit reads it, its kept original verified first and left as it is, and the
hold ends. Until proven — no handshake, or wg not answering — it is kept and the report says why.
The retirement is recorded apart from holds, so later applies, an undeclare, and a reassignment
find it retired rather than missing, and nothing writes it back.
Records written before Found existed left the adoption guard and the converge filter loaded on
undeclare, then deleted their unit files from under them; a unit whose own file the mesh created
is now the mesh's, whatever its record says. Found is kept apart the moment it is read, so a
first apply that enabled and then failed is not read back as the machine's; boot is found the
first time the mesh sets it; a service once stateless, or moved to another unit, is found afresh
(the old unit given back). The unit is read after the reload that loads a file written in the
same apply, and removal reports what it actually did.
removeProcess deletes filepath.Join(daemonRoot, name) whole; a process named ".." made that
/var/lib/mesh. The declaration and the removal now hold the name to one rule.
A service undeclared used to be stopped: unassigning the private network stopped the container
runtime, unassigning sshd would stop ssh, an uplink module would take the machine offline. The
host now records the unit's state when it first applies it and restores that on undeclare —
found running stays running; started by the mesh (the converge filter) is stopped again; nothing
is started on the way out; a pre-existing record leaves the unit alone.
An undeclared process had no removal at all and failed every apply on its node; its unit, timer
and bundle are now removed.
Step 3.5's first half, mirroring the controller's. A host says exactly two
things unprompted, and the difference between them is the whole interface:
a report must arrive, and a heartbeat must not be insisted on. So a report
goes through JetStream — it is the message the store-window guarantee is
about — and a heartbeat stays on core, because a heartbeat in a stream is
the mesh's least valuable message competing for retention with its most
valuable.
The host still imports nothing of the mesh's own (ADR 0005): this is its
own interface over its own libraries. It agrees with the controller because
a fixture holds both to one envelope, which is the only agreement that
survives two repositories.
Also recorded, where the next person reads it rather than in a plan: the
"newest wins" window narrows at the rollout and does not disappear. Last-
per-subject makes the catch-up half the stream's, and sequence orders them
definitively — but three pushes to a connected node are still three
deliveries. Saying which half goes is worth more than "can probably be
removed", which is how a load-bearing window gets deleted in a hurry.
The empty-resources guard refused every empty body as a likely mistake,
with no way to say emptiness was meant — so the control plane could
never tell a node to drop its last resource. The envelope gains
owns_nothing: with it, an empty declaration is applied (the node drops
what the mesh owned); without it, empty is still refused, so a
truncated or mis-composed body cannot silently strip a machine. One
test, both directions.
The host parses MTU from the found [Interface] and reports it, so the
mesh's interface can come up with the same MTU when it takes the tunnel
over. A path tuned to 1380 regresses to the 1420 default otherwise —
invisible to ping, fatal to TLS handshakes and transfers over that path
(novox/hq: the mesh had no MTU concept). Zero when the config named
none, and the mesh writes no MTU line then.
dns and ip were declared, validated, handed to the runtime — and part of
no comparison, so their first deployment compared every container equal
and changed nothing, silently. The same shape as 04-ISSUES/045: a field
that is not in the spec is a field that can never reach a container that
already runs.
Mailu's 2024.06 admin refuses to serve behind a resolver that does not
validate DNSSEC, and the runtime's own forwarder (127.0.0.11) validates
nothing — so a module shipping its own validating resolver had a
resolver nothing could be pointed at. Found live, blocking a cutover:
the admin sat unhealthy, submission answered 454, and the declaration
language had no words for the fix.
Two fields on a container, both handed to the runtime verbatim: dns —
the resolvers it asks — and ip, its static address on its user-defined
network, which exists for exactly one shape: a container others must
reach before name resolution works, the resolver itself being the case
that forced it. Both take only addresses and are refused on arrival
otherwise — a name here would reach the runtime verbatim and be refused
at create, after the old container was already gone.
docker inspect <name> resolves across every object kind, not just
containers. A module regularly names a network the same as the
container that joins it (keycloak does this today, ordinarily) — so
when the container does not exist yet but the same-named network
already does, the bare form answers with the network's JSON instead
of reporting the container absent, and the template these callers use
(.State.Running) fails to execute against it entirely.
Live on novox tonight: minio's LB container, named the same as its
network ("minio"), could never be created — every apply crashed on
"the container runtime could not say whether minio is here", stuck
since first push, because the check itself never got a clean answer.
Fixed at every call site asking a container's state by name
(containerState, inspectFound, NamesFree, raiseGiteaServer,
containerRunning) by scoping to `docker container inspect`, matching
the type-scoped form this codebase already uses correctly for
networks, volumes and images elsewhere. Also scoped the one image
inspect that was still bare (publish.go), for the same reason.
mesh-host runs as a host-level service (nox-mesh-host.service), not a
Docker module — merging this does not redeploy it. The live novox
failure persists until the service itself is rebuilt and updated.
Review of the ADR 0105 build (hq ADR 0105). The takeover stopped the found
unit and then found out whether the mesh's interface would do; a start that
failed left the machine with no tunnel at all.
Now nothing is stopped until the declared interface listens on the found port
at the found address and the key file it names holds the found key — the
refusal names the remedy — and a mesh interface that fails to start after the
takeover has the found unit started again, with the account saying so. The
account has three states (not taken, taken, down) and is given on every
takeover, failure included. An interface raised by hand is looked at again
for a moment and then refused naming `wg-quick down`. A found unit started
again by hand beside the mesh's is said, not stopped: on the hub it cannot
hold the port, and on a spoke two interfaces with one key would fight.
`mesh-host overlay take --tunnel <iface>` is the path for a node that
enrolled before the mesh knew to take a tunnel over: the found key becomes its
overlay key — identity, sealing and serving keys untouched, so nothing sealed
to the node is remade — and the mesh is told with a rekey signed by the
identity key, over the key left, the key taken and the tunnel. Told first,
written second, so a run again puts right whichever half did not happen.
The plan says what an apply would change from the declaration and the
record, before the machine is touched. The apply now recreates a container
when the content of a file it reads at creation changed, and the plan said
"check" for every recorded container — true, but a preview that hides the
one step somebody asked about.
So a recorded container whose record of what it read differs from what this
apply will hand it — a plain file declared here, by its declared content;
otherwise what this host last wrote at that path — is planned as an update
naming the file, the same comparison applyContainer makes. What the record
cannot settle stays a check: a file neither declared nor recorded is read
from the machine by the apply, not by the plan; and a container with no
record of what it read was labelled before the host kept that record and is
accepted as it is.
novox/hq 04-ISSUES/103, 104
Review of the first cut found four things.
A directory mounted into a container is no longer looked inside, not even
for the files this host wrote there. The controller records every
provider's received and contributions file as a plain file under a mounted
directory, so folding those in would have recreated the route proxy — which
re-reads its routes live, by design — on every route change, and killed
every provisioner sidecar, which polls what it receives, mid-reconcile on
every grant. Whether a service reads a file under its directory once or
watches it is the service's; restart-on is how a module says "once", and it
stays the opt-in. Env-files and files mounted directly remain by content.
Genesis wrote the superuser secret as `value\n`; `secret accept` strips the
line ending by design, so the postgres module declared `value` — and with
a mounted file's content in the spec, phase three would have recreated the
store it meant to adopt in place, with the temporary control plane
connected to it. Genesis now writes the value alone. readCredentialFile
tolerated both endings already. Pinned with the bytes the genesis code
path writes, then the module's declaration of the same container: it must
reconcile.
A container carrying a label from before the host folded in what it reads
is accepted rather than recreated, when that label matches the spec as it
used to be computed: what it reads is recorded then, a change is caught
from that record from the next apply on, and the label is renewed at the
next genuine recreate. Recreating them all would have been a restart storm
across the mesh in declaration order, the store first. The trade-off is
stated in the code: a container already stale at upgrade time is not
caught, and could not have been either way.
The record of what a container read is looked up by its name when its
declared id has none — the bundle's `store` becomes `postgres.server` for
the same container — so a change on the day it is adopted still names the
file. The by-target lookup takes the most recently applied record, since
the bundle's record for the same target is never removed by the mesh's.
novox/hq 04-ISSUES/103
The host decided whether a container was still the one declared by a digest
of its declaration, and the declaration names an env-file's path and a
mount's path — never what is in them. So when the store was given a new
port, the host rewrote the forge's and the analytics service's environment
files, correctly, and left both containers running with the old port in
their environment: a container reads its env-file when it is CREATED, and
`docker restart` hands it the same environment again. Both looked healthy
until they answered 502.
What a running container takes in at creation is now part of its spec, by
content: every env-file, a file bind-mounted into it, and every file this
host wrote at or under a directory bind-mounted into it — the secrets,
bindings and configs under a module's state directories. The digest is the
one the store already records for a file the host wrote (`wrote`), read
from the state as it stands when the container is reached, so a file
rewritten earlier in the same apply is already the new one; a file the host
has no record of — an env-file a predecessor left, the superuser secret
genesis writes before any declaration names it — is read from disk, which
is what keeps adopting a running store in place a reconcile and not a
recreate.
Deliberately not part of it: what else is in a bind-mounted directory,
which is the service's own data and changes while it runs; a named volume;
a seed created once, which digests as the seed the host wrote and not as
what has grown in it; and a step — a run-once or scheduled container reads
its files when it runs and runs fresh each time. On an adopted node a held
container is held before any of this is looked at.
The host records what each container was created reading, per file, so
the recreate can say which file changed — "recreated: <file> changed" in
the report and, now with its detail, in the log. A container made before
this record existed is recreated once and says so.
novox/hq 04-ISSUES/103
Review of the fix for hq issue 104 found three faults in it. A file applied
on an enrolled node — the mesh's own last declaration included — is applied
as the bundle is, so its resources are recorded as the machine's own and
what the mesh declared reads as undeclared: the plan removed the foundation.
`apply FILE` is for a machine the mesh has not spoken to, and is now refused
saying so whenever declared.json exists. The plan looked at what is held
before what the declaration says is taken, so the one cutover ADR 0100 says
must be previewed read as a hold; it now decides in holdOnAdopted's order,
models a step run inside a held container, and a test holds the plan's
sequence to the apply's outcomes. Genesis wrote the mode on every run, so a
re-run after `converge` left the state saying adopted while the kept,
signed declaration said converged, and the reconcile loop refused every five
minutes with no delivery coming to end it: genesis now writes the mode only
when none is recorded, and where the state and the verified kept declaration
disagree, the kept declaration wins and the repair is said.
Also: a file lock beside the state, taken by the link service, the host's
own commands and the installer alike, so a `reconcile` run by hand no
longer races the loop's save — chosen over refusing while a named service is
active, which would miss a `mesh-host run` started by hand; `--json
--dry-run` emits {plan} like an apply emits {plan, report}; the README's
duplicate flag line; and the bundle refusal is about the digest, not a claim
the carried bytes can never match what genesis applied.