Commit Graph
5 Commits
Author SHA1 Message Date
jschoubben 8b974deb42 A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all
nine paths open. The mesh computes the graph, delivers it as a declaration, and
the nodes bring it up.

Every fault below looked like success from inside the mesh: the graph was
right, the files were right, the services were up, every node reported it had
applied. None was reachable by reasoning.

A running interface does not re-read its configuration. A node joins, every
existing node's peer list changes, the file is replaced -- and the service is
already running, so nothing reloads it. Fixed as declared state rather than a
command: the service must reflect the file. A command to restart would be an
action, and the link may not carry one. The host refused exactly that, which is
how this shape was arrived at.

A hub sharing a site with a spoke appeared twice in that spoke's peer list --
once as a direct peer, once as the route of last resort. WireGuard takes one
entry per key and refuses the file. The ordinary shape of a small mesh, and in
none of the tests written before it ran.

Two nodes at one site that neither can be dialled were peered directly. Nobody
opens the path, and the direct route is more specific than the hub's, so it
wins and blackholes -- this design's own warning arriving in its
implementation. They now route through the hub unless one end can be dialled.

And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled
carried nothing between its spokes. The substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so. The hub inserts
its own rule above those chains and removes it on the way down.

Two weak tests found by injection along the way: one asserted the keepalive
rule only against the hub, whose peer entries happen not to set that field at
all, so it tested an absence; the other checked the firewall rules by looking
for FORWARD anywhere, which the PostDown line satisfies on its own.
2026-08-29 18:04:15 +02:00
jschoubben f563ababa1 The mesh keeps a copy of what each node owns
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host
reports what it owns and the mesh keeps the last report. A backup, never a
source -- nothing decides anything from it, and a node that disagrees with it
wins, because the node is the one that can see the machine.

Its point is the orphans. A node that loses its state file currently strands
whatever it applied: nothing on the machine knows those resources were the
mesh's doing, so nothing removes them. With this, a rebuilt node receives both
the declaration and the record of what it previously owned.

Never reported and reported nothing are kept apart, and that is the whole care
in it. A node that applied nothing holds nothing; a node that has never spoken
is unknown -- and handing back an empty list for the second would tell a
rebuilding node it owns nothing and have it remove whatever it found.

The age comes back with the answer rather than being left for the caller to go
and find. An answer about a machine is worth much less without one, and this
repository has already been bitten by a cache with no age on it.

A refusal or a partial failure moves last_seen and nothing else: neither is an
account of what the machine holds, and recording one as though it were would
tell a rebuilding node to remove what it still has.
2026-08-29 16:51:54 +02:00
jschoubben 0e116d2d65 Bind every key a node may publish
The control queue was bound to enrol and not to report, so every report a node
sent was accepted by the broker, matched no binding, and dropped. The publisher
saw success and the consumer saw nothing, for an afternoon.

The refactor that was meant to bind both never applied -- it left behind a
helper nothing called, which compiled and passed vet. The loop is now where the
bind is, so there is one place to forget rather than two.
2026-08-29 16:44:05 +02:00
jschoubben bbbc860188 The control plane declares, and hears back
`declare` sends a node a signed declaration; `serve` now also consumes reports.

Signed over the exact bytes published, which is what the node verifies. Anything
re-encoding in between would sign one thing and check another, and a difference
in key order alone would have a node refuse a declaration that was genuinely
the mesh's.

Sent to the node's queue directly rather than through the exchange: a
declaration is for one node, and routing by name through a shared exchange
means a binding per node that nothing removes when a node is retired.

Enrolment now issues the node its own broker password, replacing the token's
secret, and tells it the broker address, the fingerprint and the signing key --
so a node can reconnect after a restart without a person and a new token, which
is what makes disconnection ordinary rather than a crisis.

A report is a statement, not a write. What a node says it applied is its own
account of its own machine, kept as a copy for recovery rather than as a source.
2026-08-29 16:23:37 +02:00
jschoubben 46e760fc94 The control plane serves, and a node can join
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates
the node's account when it issues the token, and the one-time secret is that
account's password. A joining node's first connection is already authenticated;
enrolment is what it says once it is in. I had been treating this as a decision
that needed taking, and it did not.

The account is per node and scoped: it may read its own queue, write to the one
exchange, and configure nothing else. The patterns are anchored and the node
name is constrained to characters that cannot widen them, because a name
carrying a dot or a star would silently let that node read everybody's queues.

`serve` is the control plane running: one connection, one queue, one consumer.
One deliberately -- two consumers on a queue get round-robined and each receives
half of what it expects, which has happened on this project before, between a
module's daemon and its capability server.

Enrolment spends the token first, in the single statement that both finds and
marks it, and only then records the key. That order is the order things become
irreversible: recording a key for a node whose token turned out to be spent
would leave the mesh believing a machine that never had the right to join.

Refusals are one message for every reason. The log says which, where an
operator can see it; the node is told only that the token cannot be used.

Verified in the lab, on a sealed machine, through the whole first-node path.
2026-08-29 16:03:14 +02:00