Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben 65ade756f2 Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one
has to live where the generator can see it. It does now: the module ships
defaults, settings go over the top by key, and the file is produced from both.
Upstream can rewrite its half freely and the keys somebody chose survive.

Two layers, both from the start. The mesh's settings for a module, then one
machine's over those. A node that differs is expressed by differing, rather
than by restating everything the rest already say -- which would pin all of it
against future changes for no reason.

An override beats a default and there is nothing to resolve. A setting is a
statement about that key made deliberately; the default was only ever what to
do in the absence of one. So when upstream changes a key somebody has set,
there is no conflict, no merge markers, and nothing to ask.

Nested blocks merge and lists are replaced whole. Setting one field of a block
must not delete its siblings, or every setting would restate the whole block
and pin all of it. A list that merged element-wise could neither be shortened
nor reordered, and there is no correct guess about which element is "the same
one".

A module can keep specific keys for itself -- a socket path its own code
depends on -- and setting one is REFUSED rather than ignored. A setting quietly
dropped is somebody believing they changed something.

Settings that reach nothing are named at the moment they would be used, not
discovered later by the machine not behaving differently.

`plan --files` prints what a machine would be given before it is sent, because
"1 resource" does not tell you whether the merge landed.

One test kept with a note that it does not defend this code: output stability
comes from Go's encoder sorting map keys, so it passes with the merging
removed. Worth having as the thing that would catch a change of encoder, but it
is not evidence about anything written here, and it was checked.
2026-08-29 23:01:09 +02:00
jschoubben 653e232f1c The mesh knows where a module came from, and whether it is behind
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.

`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.

  zsh    holds 4f2a9c1e, source has 9e3b7d2a
         running on laptop

The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.

Three things this had to get right.

A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.

A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.

And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.

Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
2026-08-29 22:32:16 +02:00
jschoubben 409cd16a09 The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.

Everything from the module conversation, built and run on real machines:

  assign laptop i3      -> accepted, brings xorg, because nothing else provides
                           it and there was no choice to make
  assign laptop sway    -> refused: xorg and wayland both claim the-seat
  assign laptop editor  -> refused: three modules provide a shell -- bash,
                           fish, zsh -- choose one
  assign laptop zsh     -> accepted, and the editor's requirement is answered
  bash, fish beside it  -> fine, nothing is claimed

Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.

Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.

Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.

Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.

One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
2026-08-29 22:00:06 +02:00
jschoubben f44e73d286 The mesh computes a private network it cannot impersonate
The first thing the control plane decides rather than relays. Every node's peer
list is derived from every node at once, which is what makes this control-plane
work by definition: no node has that view.

A hub, with direct peering between nodes at the same site. Not a full mesh, and
the reason is a property of WireGuard rather than a preference -- there is no
failover, so a more specific route to a dead endpoint blackholes instead of
falling back. A node gets exactly one path to any peer, because two would mean
one of them silently swallowing traffic. A roaming node is hub-only for the
same reason.

Reachability and the hub are declared, never inferred from an address. The
address is evidence and is not the fact: carrier-grade NAT looks public and is
not, a routable address behind a closed firewall looks public and is not, and
the regular expression that used to decide it got the lab wrong too. Hub
election by address prefix failed silently when nobody knew the convention.

No private key travels, and that is the whole design. The node generated its
own keypair and kept the private half; the configuration points at a file the
node wrote, using WireGuard's own PostUp. So the control plane composes a
complete configuration for a node it cannot pretend to be -- it knows every
public key and holds none of the private ones.

Delivered as an ordinary declaration: a package, a file and a service. The host
does not know what a private network is and does not learn one. There is a test
holding that line, because the moment connectivity needs a new shape in tier 0
is the moment the host stops being small enough to trust.

The generated file is written to be read: each peer says why it is there, a
peer with no endpoint says why it has none, and the header says not to edit it
-- an edit survives until the graph next changes and then vanishes, which is
worse than never being applied, because the machine works and then stops and
nothing changed that anybody remembers.

Fault injection found one weak test. The keepalive rule was asserted only
against the hub, whose peer entries happen not to set the field at all, so it
was testing an absence rather than the rule. It now checks two direct peers
where one is reachable and one is not.
2026-08-29 16:58:56 +02:00
jschoubben f563ababa1 The mesh keeps a copy of what each node owns
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host
reports what it owns and the mesh keeps the last report. A backup, never a
source -- nothing decides anything from it, and a node that disagrees with it
wins, because the node is the one that can see the machine.

Its point is the orphans. A node that loses its state file currently strands
whatever it applied: nothing on the machine knows those resources were the
mesh's doing, so nothing removes them. With this, a rebuilt node receives both
the declaration and the record of what it previously owned.

Never reported and reported nothing are kept apart, and that is the whole care
in it. A node that applied nothing holds nothing; a node that has never spoken
is unknown -- and handing back an empty list for the second would tell a
rebuilding node it owns nothing and have it remove whatever it found.

The age comes back with the answer rather than being left for the caller to go
and find. An answer about a machine is worth much less without one, and this
repository has already been bitten by a cache with no age on it.

A refusal or a partial failure moves last_seen and nothing else: neither is an
account of what the machine holds, and recording one as though it were would
tell a rebuilding node to remove what it still has.
2026-08-29 16:51:54 +02:00
jschoubben 66768208d2 Node records, and the right to join once
The next step after the schema: inventory now holds node records and enrolment
tokens, and mesh-control has the commands to work with them.

A token is issued for a node record, which is where re-enrolment gets decided
-- what an identity binds to is settled when the token is made, not when it is
presented, so the machine presenting one does not need to know whether it is
joining or returning.

What the token guarantees, each with a test confirmed to fail when the
behaviour is removed: the secret is 256 random bits, shown once and stored only
as a hash; it works exactly once; it stops working when it expires; issuing
again for a node invalidates the outstanding one, because two live tokens are
two machines able to join as the same node. Redemption is a single statement
that finds and spends together, so eight concurrent attempts on one secret
produce exactly one winner rather than a race between a check and a write.

Refusals are deliberately identical for unknown, spent and expired. Somebody
guessing must not learn which guess was a real token that had merely aged out.

SHA-256 rather than a password hash, and that is a choice not a shortcut: the
secret is high-entropy random, so there is nothing to guess and a slow hash
would buy nothing while making every redemption expensive.

It stops before what a node receives in exchange. What a machine presents
afterwards to prove it is that node is not decided anywhere, and a migration is
the most expensive place here to guess.

So a token carries one of the four things ADR 0004 requires. The command prints
the secret and then says exactly that -- the broker's address, its certificate
fingerprint and the control plane's signing identity do not exist yet. Better
than emitting something that looks complete and silently cannot be used.
2026-08-29 14:46:03 +02:00
jschoubben 306c4ca13b The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one
thing with it: brings its schema up to date. That is step 3 of the substrate
bootstrap -- the step the first node cannot get past.

Verified against a real PostgreSQL, with the built binary: applied 0001-nodes,
reported 'already up to date' on the second run, and the node table is there
with the index and the unique constraint the migration asks for.

Written in Go, and the image is FROM scratch holding one file. Confirmed by
unpacking it. That is the whole argument of ADR 0024: the bundle pins this
image by digest and runs it where nothing can check it, so everything in it is
something a person has to audit before trusting a first node.

Exclusive store ownership is built as a rule about credentials rather than
about intentions. There is no mesh-wide connection setting and no way to ask
for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so
reaching another context's store needs a new variable, which is visible in the
declaration that runs it.

The migration runner is mostly refusals: an edited migration that already ran,
a migration numbered below one that has run, duplicate numbers, misnamed files,
empty files. All stop rather than warn, because at the moment any of them is
true nobody knows what the database holds.

It stops before identity, deliberately. What a node presents to prove who it is
has not been decided anywhere, and a migration is the most expensive place in
this system to guess.

Two tests did not defend what they claimed, and both are fixed rather than
removed. One asked only whether Open returned an error, which it did either way
-- a bad context name and a missing credential both fail, so deleting the name
check changed nothing. The other claimed to prove the migration runs in a
transaction, but PostgreSQL already wraps a multi-statement query in one of its
own, so it passed with the transaction taken out. What the transaction actually
buys is that the schema change and the row recording it commit together, and
there is now a test for that which fails when they are split.
2026-08-29 02:44:09 +02:00