A fourth key, reported at enrolment like the others. The reasoning is the
one this file's neighbours already give twice: a key used for two
purposes is one rotation away from breaking the other.
The private half never leaves the machine. The mesh is told the public
half and signs a certificate binding it to this node's name inside the
mesh — so there is nothing to seal, and a copy of what the mesh holds
certifies nothing it did not already certify.
It does not make one on demand, for the same reason the sealing key does
not: a key the mesh has never certified is a key nothing will trust, so a
node that quietly generated one would serve a certificate for a key it no
longer has and fail in a way that names neither.
Everything else in a declaration is visible to whatever carried it. The
message is signed so it cannot be forged, and signing does not make it
unreadable — a password in `content` is a password the broker sees, which
is the transitive trust this design refuses everywhere else.
So a node generates a third key at enrolment and reports the public half,
exactly as it does for its identity and its overlay key. A file may
arrive `sealed` instead of `content`; the host opens it with that key and
writes the result. The control plane can then store a credential it
cannot use, and the broker relays a blob it cannot read.
A third key rather than reusing one of the two. The identity key signs
and is Ed25519; the overlay key is WireGuard's and is tied to being on
the private network, which a machine may not be. A key used for two
purposes is one rotation away from breaking the other.
Details that are not incidental:
- sealed and content together is refused, so "was this the secret or the
placeholder" is answerable by looking
- a sealed file defaults to 0600 rather than 0644, because the
consequence differs; an explicit mode still wins
- a node with no sealing key refuses the file rather than skipping it. A
machine that quietly omits the one resource carrying a credential looks
configured and cannot connect
- what is recorded is a digest of what was written, so drift on a
credential is still detected without the node keeping the value, and
the report that goes back over the broker carries neither
The key is made at enrolment rather than on first use. One made later is
one the mesh was never told about, so nothing could ever be sealed to it,
and the node would look fine and receive nothing.
This is why sealing was borrowed from another mesh's mistakes rather than
its design: there, credentials sit encrypted in the control plane's
database — which guards the database file and nothing else, since the
same value is also in each node's environment file in plain text and
inside every connection string composed from it. Its own tooling has to
search by value rather than by name to find the copies, and says the ones
inside composed URLs are usually the only copies in use.
Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.
That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.
Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.
Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.
Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
The loop the whole thing exists for: told, apply, report.
`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.
Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.
A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.
Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.
That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.
Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.
The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.
An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.
Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.
The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.
Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.