Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.
That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.
Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.
Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.
Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.
State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.
Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.
Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
The loop the whole thing exists for: told, apply, report.
`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.
Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.
A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.
Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.
That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.
Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.
The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.
An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.
Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.
The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.
Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.
Five steps now instead of three. A sealed machine goes from bare to a container
runtime, a store, the inventory database, that database's schema applied by
mesh-control, and LavinMQ running and answering.
The database is called inventory rather than mesh. ADR 0008 grants a context
only what it exclusively owns and ADR 0006 says the mesh database names a thing
that will not exist -- so one database per context, and there is one context.
The broker is in the bundle because ADR 0006 now says it must be: the control
plane reaches a node only over the link, the link is the broker, so nothing can
provision the broker. Two images, which is the cost that record accepts.
Verified by reading the system rather than the report: inventory present and
mesh absent, the node table with its indexes, the migration row, lavinmqctl
answering, 5672 listening. Sealed confirmed both ways -- the internet times out,
the lab registry returns 200.
Then rebooted, which was the part worth doing rather than assuming. Everything
returned: docker from boot: enabled, both containers because this host creates
every container --restart unless-stopped, the schema intact in its volume. Three
reconciles before and one after all report no change.
One thing that reads as a success and was not: the first sealed check said the
machine could reach example.com. It was the test that was wrong -- a helper
script pasted arguments into a shell line, so a command with quotes was re-split
and ran on the workstation. The machine had been sealed the whole time. The
helper now requotes each argument.
The first three steps of the substrate bootstrap, run on a lab machine
confirmed to have no route out. It went from bare to a container runtime
installed and enabled, PostgreSQL running from an image pinned by digest, and
the control plane's database created inside it -- from the file the host
carries, with nothing to ask.
Second run changed nothing. `owned` lists all five afterwards, and `mesh` is in
the store.
It stops before the last two steps because there is no control plane yet: its
schema cannot be loaded and its image does not exist. The bundle says so rather
than naming something that cannot be applied.
Two things fixed on the way.
`make host BUNDLE=...` still swapped a file called substrate.lock, which the
per-system split had renamed months of decisions ago -- it now takes SYSTEM and
replaces that system's bundle. And the .lock files still cited ADR 0060, since
the renumbering pass only covered .md, .go, .ts and .sh.
One thing learned by it failing first: a directory the host creates is owned by
root, and a database inside a container runs as somebody else, so it could not
write and the container crash-looped. The store's data is a named volume now,
which lets the image set up its own ownership and outlives the container --
which is what you want for the thing holding the mesh's state.
Worth noting the failure was caught by the action's verify rather than by the
container step. `docker inspect` reported the container running because it was,
briefly, between restarts. Running is not working, and the thing that knew the
difference was the step that asked the database whether it would answer.