Commit Graph
20 Commits
Author SHA1 Message Date
jschoubben 770f589401 Report what an adopted node holds, its firewall and what is reachable, and speak unasked when that changes (hq ADR 0100) 2026-09-22 17:22:31 +02:00
jschoubben 406a5559b0 An enrolling node signs its request with the identity it just generated, so the mesh can tell it from anyone who knows its public key (novox/hq issue 083) 2026-09-22 14:33:31 +02:00
jschoubben eaebae7b36 Review of 083: the give-up message says to wait out the mesh's hold before asking again; the installer no longer says the token is spent when it may not be 2026-09-22 14:23:20 +02:00
jschoubben c8bbdb2da5 An enrolling node asks again, with the same request, while the mesh says it cannot answer — for as long as the mesh holds the token for it (novox/hq issue 083) 2026-09-22 14:07:25 +02:00
jschoubben b72b71a989 The host applies the newest declaration, a file may be created once, the foundation filters first
031: a window of unacknowledged declarations is drained to the newest; the
rest are set aside and reported as superseded. 035: a file resource may say
create-once — written when absent, kept untouched when present (ADR 0087).
054: the bundle installs nftables and loads a base ruleset before the store
and broker, in the table the filter module later replaces (ADR 0088).
2026-09-21 12:11:52 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 8211d8b6fb A report says which declaration it is about
The mesh decided "has this machine caught up" by comparing its send
time to the report's arrival, and lost the race it invited: an apply
started under the previous declaration finishes after the next one is
sent, its report lands newer than the send, and the machine reads as
caught up with words it has not read yet. The lab hit exactly that —
one test's closing push was still being applied when the next test's
push recorded its send, and the next test then read files that were
never going to be there yet.

Clocks cannot answer "which". The report now carries the digest of the
exact bytes it applied — the same bytes, hashed the same way, that the
mesh recorded when it sent them — and which-declaration becomes an
equality the mesh checks rather than an ordering it hopes.
2026-09-02 00:01:12 +02:00
jschoubben b91342a6bd A machine says which ports it already holds
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.

The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.

What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.

Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
2026-09-01 18:29:39 +02:00
jschoubben 8e12b3c9e4 Name the decisions these tests defend, and check the bundle at all
From auditing the decision records: of 28, only 12 were named by any
test, so "which decisions are defended" could not be answered without
reading everything. ADR 0017 says a test names the decision it defends —
that rule was itself unenforced.

Most of the gap was citation, not coverage. Drift detection was tested
in several places without naming ADR 0011; the archive refusal without
naming 0012; forged declarations without naming 0002. Named now, so the
question is answerable by grep.

The bundle was the real gap: nothing tested substrate-first-node.lock at
all. It is what a machine becomes when there is no mesh to ask — the one
declaration applied with nothing to verify it against — and it was
edited by hand and read by nothing but a running host.

Two tests now assert what it carries: exactly postgres, lavinmq and the
control plane. That defends ADR 0028, which removed the object store
from the substrate after it had been a member for months on the strength
of "it cannot grant itself a bucket" — true, and the answer to only half
the test. Nothing counted what the bundle held.

Fault-injected, and the first attempt did not bite: the injection landed
on a comment line, which stripComments discards. Injecting into the
image field fails as it should.
2026-08-31 17:33:34 +02:00
jschoubben 8fcfa88fe0 A machine that wakes or moves says so, instead of waiting to be told
A suspended laptop's connection is dead the moment it wakes, and the socket
looks perfectly healthy from inside the process — no error, no close, because
nothing has tried to send anything. Heartbeats find out twenty or thirty
seconds later. For that time the node believes it is in a mesh it has left,
which is the one state this design says must never be indistinguishable from
being connected. The machine knew immediately.

So being roused ends the current attempt rather than only shortening the wait
after it: shortening the wait would do nothing at all, because the process is
not waiting — it is sitting inside a connection that will not return.

A signal, because nothing may listen on a node (novox/hq ADR 0004). A socket
for this would be a control surface on every machine, reachable by anything
that can reach the machine, in exchange for saving twenty seconds — and the
whole security argument rests on there not being one.

Two rouses in the same instant are one: a machine suspending and resuming
repeatedly must not build a backlog of reconnections to work through. And the
backoff is not reset by being roused — that says the machine changed, not that
whatever was refusing the connection has stopped, and a laptop woken on a
network with no route would otherwise retry at full speed for as long as
somebody keeps opening the lid.

The dispatcher acts on the events that change where packets go and not on
`down`: the link is already gone there, reconnecting will fail, and the backoff
exists for exactly that.
2026-08-31 10:21:26 +02:00
jschoubben 5237944473 A node generates the key it serves TLS with
A fourth key, reported at enrolment like the others. The reasoning is the
one this file's neighbours already give twice: a key used for two
purposes is one rotation away from breaking the other.

The private half never leaves the machine. The mesh is told the public
half and signs a certificate binding it to this node's name inside the
mesh — so there is nothing to seal, and a copy of what the mesh holds
certifies nothing it did not already certify.

It does not make one on demand, for the same reason the sealing key does
not: a key the mesh has never certified is a key nothing will trust, so a
node that quietly generated one would serve a certificate for a key it no
longer has and fail in a way that names neither.
2026-08-31 00:09:15 +02:00
jschoubben ef0d96a1a8 Write out what this node says when it joins, for the mesh to read
The enrolment request is a struct in each repository, and this node now
reports a third key — the one its secrets are sealed to. That wiring had
tests on each side and had never been run across the join, where a
renamed field fails silently: enrolment succeeds, the key is absent, and
the node looks joined until the first thing sealed to it cannot be
opened.

So this writes a real one — keys generated the way enrolment generates
them, not typed as literals — and the private half of the sealing key
beside it, so the other side can prove what it sealed is openable rather
than merely present.

The mirror of the declaration check that already runs the other way.
2026-08-30 02:20:43 +02:00
jschoubben a752fc514b A file the mesh can deliver and cannot read
Everything else in a declaration is visible to whatever carried it. The
message is signed so it cannot be forged, and signing does not make it
unreadable — a password in `content` is a password the broker sees, which
is the transitive trust this design refuses everywhere else.

So a node generates a third key at enrolment and reports the public half,
exactly as it does for its identity and its overlay key. A file may
arrive `sealed` instead of `content`; the host opens it with that key and
writes the result. The control plane can then store a credential it
cannot use, and the broker relays a blob it cannot read.

A third key rather than reusing one of the two. The identity key signs
and is Ed25519; the overlay key is WireGuard's and is tied to being on
the private network, which a machine may not be. A key used for two
purposes is one rotation away from breaking the other.

Details that are not incidental:

- sealed and content together is refused, so "was this the secret or the
  placeholder" is answerable by looking
- a sealed file defaults to 0600 rather than 0644, because the
  consequence differs; an explicit mode still wins
- a node with no sealing key refuses the file rather than skipping it. A
  machine that quietly omits the one resource carrying a credential looks
  configured and cannot connect
- what is recorded is a digest of what was written, so drift on a
  credential is still detected without the node keeping the value, and
  the report that goes back over the broker carries neither

The key is made at enrolment rather than on first use. One made later is
one the mesh was never told about, so nothing could ever be sealed to it,
and the node would look fine and receive nothing.

This is why sealing was borrowed from another mesh's mistakes rather than
its design: there, credentials sit encrypted in the control plane's
database — which guards the database file and nothing else, since the
same value is also in each node's environment file in plain text and
inside every connection string composed from it. Its own tooling has to
search by value rather than by name to find the copies, and says the ones
inside composed URLs are usually the only copies in use.
2026-08-30 00:12:22 +02:00
jschoubben 4bff67ec69 A machine waiting to be enrolled is not a broken one
The launcher already ran `host run`, and `run` on a machine with no identity
exited with an error. So a freshly installed host, sitting exactly as intended
waiting for somebody to bring it a token, would have counted three failed
starts and rolled back its own installation.

It waits now, and says what it is waiting for. That is the *hosted* state from
the lifecycle: the host is running, it has no identity, and there is nobody to
link to. Every machine passes through it.

An identity that exists and cannot be read is still a fault rather than a wait.
Treating that as "not enrolled yet" would leave a node sitting quietly for ever
while the mesh believes it is a member.

Also: a node now says it is there once a minute. Nothing but its name, because
anything more would be a report, and reports are rare where this is constant --
reading one as the other would make a quiet node look like a stale one. Not
published mandatory, unlike a report: losing one is nothing, the next is a
minute away, and the mesh reads a gap rather than counting arrivals.

Verified in the lab: a node was stopped and the mesh said "out of touch 4m",
then it was started and the mesh said "here" again, without anything else being
touched.
2026-08-29 20:32:18 +02:00
jschoubben 980e12a850 A node that loses its mesh comes back on its own
Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.

Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.

A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.

And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.

The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.

With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.

Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
2026-08-29 20:17:49 +02:00
jschoubben 1bc97ed50d A service can be declared to reflect a file
Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.

That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.

Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.

Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.

Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
2026-08-29 18:04:16 +02:00
jschoubben fa48b5825e The bundle and the mesh stop removing each other
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.

State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.

Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.

Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
2026-08-29 16:43:46 +02:00
jschoubben a488c76b5e A node holds its link open, and applies what the mesh signs
The loop the whole thing exists for: told, apply, report.

`run` holds one outbound connection open and consumes the node's own queue.
Every declaration is verified against the control plane's signing key before a
byte of it is read as an instruction -- not once at connect, every time. The
transport being pinned is a different question from the instruction being
genuine, and pinning only the first would make the second transitive: a
compromised broker could forge declarations, and this host applies whatever the
link delivers.

Malformed and forged are reported differently, because ADR 0004 requires a host
to tell "this is not from the mesh I joined" from "this is broken". One means
somebody is trying and the other means something needs fixing.

A node now keeps what it needs to come back on its own: the broker's address
and fingerprint, the signing key it believes, and its own broker password --
which the mesh issues at enrolment to replace the token's secret, so the
one-time thing stays one-time and the credential it holds for years is not the
one that was pasted into a terminal.

Verified in the lab end to end. The node enrolled, held its link, received a
signed declaration and applied it -- the file is on the machine with the right
contents, and the host's own record lists both resources.

That run also found issue 010, which is recorded in novox/hq: the declaration
removed every container on the machine, including the control plane that sent
it. Correct reconciliation, shared store, and the first thing that happens.
2026-08-29 16:23:27 +02:00
jschoubben a4445f5c0a A machine joins the mesh it raised
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.

Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.

The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.

An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.

Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.

The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.

Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.
2026-08-29 16:03:15 +02:00
jschoubben 65d896d96e A node makes its own identity, and checks the broker before speaking
The host side of enrolment. It parses a token the control plane issued, dials
the broker, refuses anything but the pinned certificate, and generates an
Ed25519 keypair whose private half never leaves the machine.

Verified against a real LavinMQ serving a real certificate: the pin matched and
the node proceeded. Then against a second broker with a different certificate
on another port, which was refused -- with an error that says retrying will not
help, because it does not mean the network is down, it means the mesh was
substituted.

InsecureSkipVerify is set and that is the point rather than a weakening. At
bootstrap the broker is self-signed and reached at an address, so there is no
authority to trace and no name to match. Chain and hostname checks are replaced
with something stricter: this exact certificate or nothing, checked in
VerifyPeerCertificate, which runs before the handshake completes -- so nothing
is sent to the wrong broker. There is a test that counts the bytes an impostor
receives, and it is zero.

The token format is defined separately here and in the control plane, because
this binary requires nothing present and does not import it. They are held
together by a test on each side asserting the exact field names, so a rename
breaks both immediately rather than at enrolment on a real machine.

Two distinctions the identity file has to keep. A machine that never joined has
no identity, which is an ordinary state and not a fault. A machine whose
identity cannot be read is a different thing entirely, and must not take the
same path -- re-enrolling would discard the identity the mesh still believes and
need a person with a new token. Fault injection found the second case untested:
the corrupt-file test was passing on the parse check, so the read-error path had
nothing defending it. It does now.

An already-enrolled machine refuses to enrol again rather than quietly
acquiring a second identity.

What is not built is the link. Enrolment stops after verifying the broker and
generating the identity, having saved nothing, so it can be run again unchanged.

132 tests, plus 32 launcher and 9 rollback.
2026-08-29 15:38:19 +02:00