PKCS#8 PEM, not this host's own base64. The mesh delivers a PEM certificate
beside it and every TLS server there is reads PEM: nginx's ssl_certificate_key,
Go's LoadX509KeyPair, openssl s_server. Stored the other way the file was
intact, present, correctly permissioned, and unusable — the machine failed at
the moment something connected, which the lab found by connecting.
A key in the old encoding is refused by name rather than called corrupt: it is
replaced by enrolling again, and that is a different remedy from a damaged
file.
A fourth key, reported at enrolment like the others. The reasoning is the
one this file's neighbours already give twice: a key used for two
purposes is one rotation away from breaking the other.
The private half never leaves the machine. The mesh is told the public
half and signs a certificate binding it to this node's name inside the
mesh — so there is nothing to seal, and a copy of what the mesh holds
certifies nothing it did not already certify.
It does not make one on demand, for the same reason the sealing key does
not: a key the mesh has never certified is a key nothing will trust, so a
node that quietly generated one would serve a certificate for a key it no
longer has and fail in a way that names neither.
Found by raising a mesh end to end for the first time. Enrolment's own
help says the token "is the only thing it needs", and it also needed
--name, with no default. Without it the failure is:
cannot reach the broker at 192.0.2.10:5671 as : username or password
not allowed
An empty username, and nothing about the cause.
The node cannot work its own name out. The broker account it
authenticates as is named after it and exists before this machine has
been told anything, so the name has to arrive with the rest. It is not a
secret and the issuer already knows it.
--name stays, as an override for a token issued before the name
travelled in one, and says so when it is needed rather than failing at
the broker.
Also corrects the bundle example, which claimed to stop before the
control plane runs and has raised one for some time. A comment about what
something does not do is a comment nobody updates.
Everything else in a declaration is visible to whatever carried it. The
message is signed so it cannot be forged, and signing does not make it
unreadable — a password in `content` is a password the broker sees, which
is the transitive trust this design refuses everywhere else.
So a node generates a third key at enrolment and reports the public half,
exactly as it does for its identity and its overlay key. A file may
arrive `sealed` instead of `content`; the host opens it with that key and
writes the result. The control plane can then store a credential it
cannot use, and the broker relays a blob it cannot read.
A third key rather than reusing one of the two. The identity key signs
and is Ed25519; the overlay key is WireGuard's and is tied to being on
the private network, which a machine may not be. A key used for two
purposes is one rotation away from breaking the other.
Details that are not incidental:
- sealed and content together is refused, so "was this the secret or the
placeholder" is answerable by looking
- a sealed file defaults to 0600 rather than 0644, because the
consequence differs; an explicit mode still wins
- a node with no sealing key refuses the file rather than skipping it. A
machine that quietly omits the one resource carrying a credential looks
configured and cannot connect
- what is recorded is a digest of what was written, so drift on a
credential is still detected without the node keeping the value, and
the report that goes back over the broker carries neither
The key is made at enrolment rather than on first use. One made later is
one the mesh was never told about, so nothing could ever be sealed to it,
and the node would look fine and receive nothing.
This is why sealing was borrowed from another mesh's mistakes rather than
its design: there, credentials sit encrypted in the control plane's
database — which guards the database file and nothing else, since the
same value is also in each node's environment file in plain text and
inside every connection string composed from it. Its own tooling has to
search by value rather than by name to find the copies, and says the ones
inside composed URLs are usually the only copies in use.
Curve25519, which is what WireGuard uses. The private half never leaves the
machine and is written to a file of its own, so the interface configuration the
mesh composes can point at it without ever carrying it.
Separate from the identity keypair on purpose. One signs messages to the mesh
and the other encrypts traffic between nodes -- different things verified by
different parties at different times, and a key used for two purposes is one
rotation away from breaking the other.
Pushed a failing test in the last commit -- my own gate reported it and I read
the count rather than the result. The failure was real and worth having.
Adding the membership requirement to Load made Generate produce an identity that
Save would write and Load would then refuse. A file that cannot be read back is
the worst shape this could take: it is read back on the next start, on a machine
nobody is watching, and by then the token that could have fixed it is spent.
Save now refuses exactly what Load refuses, and writes nothing when it does. The
round-trip test covers the membership too, since that is the half that lets a
node come back on its own.
The loop the whole thing exists for: told, apply, report.
`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.
Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.
A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.
Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.
That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.
The host side of enrolment. It parses a token the control plane issued, dials
the broker, refuses anything but the pinned certificate, and generates an
Ed25519 keypair whose private half never leaves the machine.
Verified against a real LavinMQ serving a real certificate: the pin matched and
the node proceeded. Then against a second broker with a different certificate
on another port, which was refused -- with an error that says retrying will not
help, because it does not mean the network is down, it means the mesh was
substituted.
InsecureSkipVerify is set and that is the point rather than a weakening. At
bootstrap the broker is self-signed and reached at an address, so there is no
authority to trace and no name to match. Chain and hostname checks are replaced
with something stricter: this exact certificate or nothing, checked in
VerifyPeerCertificate, which runs before the handshake completes -- so nothing
is sent to the wrong broker. There is a test that counts the bytes an impostor
receives, and it is zero.
The token format is defined separately here and in the control plane, because
this binary requires nothing present and does not import it. They are held
together by a test on each side asserting the exact field names, so a rename
breaks both immediately rather than at enrolment on a real machine.
Two distinctions the identity file has to keep. A machine that never joined has
no identity, which is an ordinary state and not a fault. A machine whose
identity cannot be read is a different thing entirely, and must not take the
same path -- re-enrolling would discard the identity the mesh still believes and
need a person with a new token. Fault injection found the second case untested:
the corrupt-file test was passing on the parse check, so the read-error path had
nothing defending it. It does now.
An already-enrolled machine refuses to enrol again rather than quietly
acquiring a second identity.
What is not built is the link. Enrolment stops after verifying the broker and
generating the identity, having saved nothing, so it can be run again unchanged.
132 tests, plus 32 launcher and 9 rollback.