Commit Graph
14 Commits
Author SHA1 Message Date
jschoubben 047830a6b5 A report the store cannot take yet is handed back to the broker, not acknowledged and lost (novox/hq issue 082) 2026-09-22 13:25:13 +02:00
jschoubben 1ab0364705 The link knows a superseded report (issue 031) 2026-09-21 12:11:52 +02:00
jschoubben c3b88b9148 Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben cae9a3e54f A bare alive moves last_seen and nothing else
A node says it is there every minute and describes what it applied rarely,
and both went through Heard, which wrote every one down as a report. So a
bare alive replaced the node's last real apply with an empty one -- clearing
the declaration digest `current` is measured against, the carried ports a
push assigns around, and the clean-or-failed outcome. A node that had just
caught up read as behind within the minute, and never converged.

Whether it converged in time was a race the node's own apply set: the link's
one loop applies a declaration to completion before it can send the pending
heartbeat, so a fast apply (catalogue-small) leaves the digest standing the
~60s until the next beat -- long enough for the lab to see `current` -- while
a heavy wave whose apply outran the first beat (mongodb + unifi + marrytts)
had the alive fire milliseconds after the report and never showed `current`
at all, timing out settle even at 1200s.

Heard now returns after moving last_seen for a report that carries no account
of what the machine did -- nothing applied, nothing refused, nothing failed,
which is exactly a bare alive. A real report always carries one. This is what
the commit that began hearing alives said it did and did not: "a bare word
that a node is there moves last_seen and touches nothing else."

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:37:53 +02:00
jschoubben 1b63e21c0f Caught up is an equality, not an ordering
The report carries the digest of the declaration it applied (mesh-host
8211d8b), and the mesh stores it beside the outcome. `reported` rows in
the status JSON now say `current`: whether the machine's last word
names the declaration last sent.

Not derivable from the timestamps beside it, which is why they were
not enough: an apply begun under the previous declaration reports
after the next send — newer, and still about the old words. The lab
lost exactly that race between one test's closing push and the next
test's opening one.

Empty digests — every host from before reports carried one — read as
not current, which errs toward waiting rather than toward asserting on
files that are not there yet.
2026-09-02 00:02:44 +02:00
jschoubben 41f7c51032 Assign around what a machine already holds
The other half of ADR 0038, and what 04-ISSUES/028 was actually about.
A module can now avoid colliding with another module; until this it
could not avoid colliding with the mesh itself.

The substrate is not a module. A node raises it from the bundle it
carries before any mesh exists, so the control plane had never heard of
the store, the broker, or its own container — and handed a database
module 5432, which the store already had.

So the machine says. The host records what each resource binds,
distinguishing what it carried from what the mesh sent — a distinction
that already existed so the two never remove each other — and reports
the carried ones. The node states and this context writes, which is the
shape of every message between them.

What the declaration binds, not what is open. A machine's open ports are
a moving target, and assigning around them would mean a port that was
free when it was asked for and taken when it was used.

Replaced whole each time rather than merged: a machine that gave a port
back must be believed about that too, and a set that only grows keeps a
port reserved for something no longer there.

Tested against a real database, and the tests bite — removing the check
hands the module 20000, which the machine had said it holds.
2026-09-01 18:32:38 +02:00
jschoubben 646609c1b2 The mesh certifies names inside it
08-connectivity keeps two authorities apart on purpose: a public one for
names the outside world reaches, and the mesh's own for names only the
mesh knows. Nothing implemented the second, so anything between machines
was plaintext or trust-on-first-use — which the design refuses everywhere
else.

A node now generates a fourth key at enrolment and reports the public
half. A fourth, because a key used for two purposes is one rotation away
from breaking the other: the identity key signs messages to the mesh and
would do for TLS, and reusing it would mean rotating a node's identity
every time its certificate is replaced.

**Nothing secret travels and nothing is sealed.** A certificate authority
says "this name belongs to the holder of this key", so the mesh signs a
public half it cannot use, and the certificate it issues is public. A
module asks for one and is given the certificate and, if it wants,
the mesh's own — the private key is a path to a file the machine already
has, the same arrangement the private network's key uses.

Asserted by verifying rather than inspecting, because a certificate that
parses and does not chain fails at the moment something connects:

- what the mesh issues verifies against the mesh, for the name asked for
- the name is in the subject alternative names, since a certificate
  carrying it only in the common name is refused by every modern client
- it certifies the key the node generated and no other
- another mesh's certificate does not verify, which is the whole point of
  two authorities being separate
- the authority cannot sign another authority — one that could is one
  that can be delegated without anybody deciding to
- two control planes starting together agree on one authority, or a mesh
  has certificates half its machines refuse

Certificates last ten years, which is a choice: a short life needs
something to renew it, and a renewal that fails silently is a mesh that
stops trusting itself on a date nobody wrote down. What makes one
replaceable is that the mesh reissues on demand, not that it expires.
2026-08-31 00:09:13 +02:00
jschoubben 9681b288aa Keep what each machine did, so status can say what is wrong
A node reports back after applying a declaration: it worked, some of it
failed, or the whole thing was refused. A refusal or a failure moved
last_seen and the reason went to a log line — so "which machine is not
doing what it was told" had no answer the next morning, which is the
question a mesh exists to answer.

Refused and failed are kept as different things, because they are
different situations with different remedies: refused means the machine
is exactly as it was and what is wrong is in what was sent; failed means
it is in a state nobody declared and what is wrong is on the machine. One
word for both would make the record say less than the node did.

One row per node, replaced. The question is the machine's current state —
"this failed an hour ago and then succeeded" is not a machine anybody
needs to look at, and a table of every report would bury the ones that
matter under the ones that do not.

`status` now answers three questions in the order somebody asks them: is
anything broken, is anything not answering, is anything out of date. The
first has consequences now, the third is a plan for later, and a status
leading with the third would bury the first. A machine that has never
spoken is reported as quiet rather than as broken — new, switched off and
unreachable are not the same as tried and could not.

The mapping from a report to an outcome had no test at all, which the
injection caught: it is the code deciding which of those situations a
machine is in. It has four now, including that a partial report never
becomes the account of what the machine holds — the fault that destroyed
a substrate once.
2026-08-30 18:08:59 +02:00
jschoubben 20f78cd5f1 Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.

Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.

So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.

Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.

It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.

Two tests found passing for the wrong reason, both caught because their
injection came back clean:

- the provider's copy was asserted non-empty, which reads the same
  whichever column is selected. It now opens the blob with the
  provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
  the next read makes one anyway. Removed, and a second path to the same
  act is how two ends come to disagree.

And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
2026-08-30 00:21:18 +02:00
jschoubben f0cff88172 The mesh knows who is out of touch
09-the-node-lifecycle asks for this in as many words -- *how long it has been
disconnected is a fact the mesh must hold, and nothing holds it today. Without
it, a node running last month's assignments looks exactly like one that is
current.* Now it holds it.

`node list` says "here", "out of touch 4m", or "never spoken", and the third is
kept distinct from the second on purpose: a node that has never spoken did not
finish joining, and a node last heard from a month ago is running a month-old
picture of the mesh. Those need different responses from a person.

A bare word that a node is there moves last_seen and touches nothing else. It
is not an account of what the machine holds, and recording it as one would
replace the recovery copy with an empty list every minute -- so a rebuilding
node would then be told it owns nothing and remove whatever it found. There is
a test for exactly that.

Heard is silent in the log. A node saying it is there every minute would fill
the log with the ordinary case, and a log where the ordinary case is loud is a
log nobody reads.

Verified in the lab across the threshold, both directions.
2026-08-29 20:32:27 +02:00
jschoubben 8b974deb42 A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all
nine paths open. The mesh computes the graph, delivers it as a declaration, and
the nodes bring it up.

Every fault below looked like success from inside the mesh: the graph was
right, the files were right, the services were up, every node reported it had
applied. None was reachable by reasoning.

A running interface does not re-read its configuration. A node joins, every
existing node's peer list changes, the file is replaced -- and the service is
already running, so nothing reloads it. Fixed as declared state rather than a
command: the service must reflect the file. A command to restart would be an
action, and the link may not carry one. The host refused exactly that, which is
how this shape was arrived at.

A hub sharing a site with a spoke appeared twice in that spoke's peer list --
once as a direct peer, once as the route of last resort. WireGuard takes one
entry per key and refuses the file. The ordinary shape of a small mesh, and in
none of the tests written before it ran.

Two nodes at one site that neither can be dialled were peered directly. Nobody
opens the path, and the direct route is more specific than the hub's, so it
wins and blackholes -- this design's own warning arriving in its
implementation. They now route through the hub unless one end can be dialled.

And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled
carried nothing between its spokes. The substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so. The hub inserts
its own rule above those chains and removes it on the way down.

Two weak tests found by injection along the way: one asserted the keepalive
rule only against the hub, whose peer entries happen not to set that field at
all, so it tested an absence; the other checked the firewall rules by looking
for FORWARD anywhere, which the PostDown line satisfies on its own.
2026-08-29 18:04:15 +02:00
jschoubben f563ababa1 The mesh keeps a copy of what each node owns
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host
reports what it owns and the mesh keeps the last report. A backup, never a
source -- nothing decides anything from it, and a node that disagrees with it
wins, because the node is the one that can see the machine.

Its point is the orphans. A node that loses its state file currently strands
whatever it applied: nothing on the machine knows those resources were the
mesh's doing, so nothing removes them. With this, a rebuilt node receives both
the declaration and the record of what it previously owned.

Never reported and reported nothing are kept apart, and that is the whole care
in it. A node that applied nothing holds nothing; a node that has never spoken
is unknown -- and handing back an empty list for the second would tell a
rebuilding node it owns nothing and have it remove whatever it found.

The age comes back with the answer rather than being left for the caller to go
and find. An answer about a machine is worth much less without one, and this
repository has already been bitten by a cache with no age on it.

A refusal or a partial failure moves last_seen and nothing else: neither is an
account of what the machine holds, and recording one as though it were would
tell a rebuilding node to remove what it still has.
2026-08-29 16:51:54 +02:00
jschoubben bbbc860188 The control plane declares, and hears back
`declare` sends a node a signed declaration; `serve` now also consumes reports.

Signed over the exact bytes published, which is what the node verifies. Anything
re-encoding in between would sign one thing and check another, and a difference
in key order alone would have a node refuse a declaration that was genuinely
the mesh's.

Sent to the node's queue directly rather than through the exchange: a
declaration is for one node, and routing by name through a shared exchange
means a binding per node that nothing removes when a node is retired.

Enrolment now issues the node its own broker password, replacing the token's
secret, and tells it the broker address, the fingerprint and the signing key --
so a node can reconnect after a restart without a person and a new token, which
is what makes disconnection ordinary rather than a crisis.

A report is a statement, not a write. What a node says it applied is its own
account of its own machine, kept as a copy for recovery rather than as a source.
2026-08-29 16:23:37 +02:00
jschoubben 46e760fc94 The control plane serves, and a node can join
There was no chicken-and-egg to solve. The mesh runs the broker, so it creates
the node's account when it issues the token, and the one-time secret is that
account's password. A joining node's first connection is already authenticated;
enrolment is what it says once it is in. I had been treating this as a decision
that needed taking, and it did not.

The account is per node and scoped: it may read its own queue, write to the one
exchange, and configure nothing else. The patterns are anchored and the node
name is constrained to characters that cannot widen them, because a name
carrying a dot or a star would silently let that node read everybody's queues.

`serve` is the control plane running: one connection, one queue, one consumer.
One deliberately -- two consumers on a queue get round-robined and each receives
half of what it expects, which has happened on this project before, between a
module's daemon and its capability server.

Enrolment spends the token first, in the single statement that both finds and
marks it, and only then records the key. That order is the order things become
irreversible: recording a key for a node whose token turned out to be spent
would leave the mesh believing a machine that never had the right to join.

Refusals are one message for every reason. The log says which, where an
operator can see it; the node is told only that the token cannot be used.

Verified in the lab, on a sealed machine, through the whole first-node path.
2026-08-29 16:03:14 +02:00