Commit Graph
160 Commits
Author SHA1 Message Date
jschoubben d1c256c2b1 The board says what is waiting too, or it disagrees with the command
Adding "not running what the mesh would send it" to `status` and not to the
board would have left two answers to one question with a person in front of
each — which is the single thing this page's design forbids, introduced by the
change that was supposed to make the question answerable.

The published JSON carries it as well, so the page, the command and anything
built against either say the same thing from the same read. Additive, because
that shape is hard to change once anything is built against it.

Never told stays separate from out of date on the page as it is everywhere
else: same remedy, and nobody has ever asked that machine to be anything.
2026-08-31 06:00:04 +02:00
jschoubben 5a28434ba8 "Behind" means not running what the mesh would send
It meant "failed or refused". So a machine that applied cleanly and whose
declaration has since changed was not behind — and novox/hq ADR 0010's
question, did my change go out?, was answerable exactly for the machines that
broke. For every machine that worked, the answer was silence whether the change
had gone out or not, which is the thing replacing a pipeline was supposed not
to cost.

The mesh now records a digest of what it last sent each machine. A digest
rather than the declaration: it can compute what a machine should be at any
moment, and keeping a copy would be a second account of it able to disagree
with the first. What cannot be recomputed is what was actually sent.

Recorded after the send, not before — a digest kept for something that failed
to send would make the machine look current for a declaration it never
received.

Never told stays separate from out of date. The remedy is the same push and the
situations are not alike: nobody has ever asked that machine to be anything.
And a machine the mesh could not work out is not reported as waiting, because
saying so would invent a comparison — that is `plan`'s answer to give.

`status` says it and `push --behind` sends it, or the flag would know something
the person reading the status does not.
2026-08-31 05:17:21 +02:00
jschoubben dfca21fa55 Keep what a machine said about itself, not just the yes
novox/hq ADR 0009: a capability's presence gates an assignment and its detail
carries a value — seat: card1-DP-1, an architecture, an amount of memory. So
'can this run here' and 'what should it be configured as' are one fact read two
ways, and the mesh was keeping the first read and discarding the second.

The reason an absent capability is absent went the same way, which is the case
a person most needs: 'this machine has no container runtime' is the answer and
'docker is not installed' is why, and only the machine knows why.

`node show` says it back. Never reported and reported nothing stay different
things there — one machine has not run the host, the other ran it and can do
nothing, and those send a person to different places.
2026-08-31 04:53:45 +02:00
jschoubben 92133c340b A board that reads through the same interfaces and holds nothing
novox/hq 03-DESIGN/01-to-be/11-a-board.md, built. The board being replaced is
one service reading every context's database directly — ADR 0008 violated by
the one component with a reason to violate it. The cost is not hypothetical: a
boundary nothing may cross can move, and one thing crossing it is enough to
freeze it. A board that reads the provisioning tables breaks when provisioning
changes them, and the change then gets weighed against the board.

So the three questions are read once, by one function, for all three ways of
saying them — a person's status, its JSON, and this page. Three
implementations of "which machine is not doing what it was told" would be three
chances to disagree.

Refused and failed stay distinct all the way to the page: refused means the
machine is exactly as it was and what is wrong is in what was sent; failed
means it is in a state nobody declared. Different places to fix, so one word
for both would send half the readers to the wrong one.

It stores nothing, changes nothing, and every action it might offer already
exists as a command. A board that cannot reach the mesh says so rather than
rendering an empty page — an empty page says "nothing is wrong" in the one
situation where nobody can know that.

One test earns its place twice: a machine's own words are the whole reason the
page is useful and the one thing on it nobody in this repository wrote, so they
are shown and are not markup.
2026-08-31 04:49:21 +02:00
jschoubben 29b336bb8d Say the remedy beside the problem in status
A status that names what is wrong and not what to do about it makes somebody go
and find the command — and the command is the whole point of having noticed.
The hint existed on one of the two paths that print this.
2026-08-31 03:22:24 +02:00
jschoubben faf5ecd70f One machine's unanswerable requirement does not remove it from the mesh
The pass that answers *what does this node offer* takes a failed resolution to
mean it learned nothing about that node. So refusing an unanswerable
requirement there made the machine disappear — and every other machine was then
told, wrongly, that the two of them shared no private network.

A wrong answer about a machine nobody asked about, caused by a fault on a
third. The lab found it: one module needing a licence that had not been added
yet made two unrelated machines look disconnected.

The second pass still refuses it, where the question is actually being asked.
2026-08-31 03:06:33 +02:00
jschoubben 005b928d36 Test the licence store, and name an unknown licence rather than a constraint
Its own store, its own test database, the same shape every other context has.
Five properties: a key with nobody to seal it to is refused rather than kept
readably; a key is sealed once per holder and the blobs differ because they are
sealed to different machines; a holder recorded afterwards has none and the
existing ones keep theirs; releasing a consumer takes its key; and a licence
nobody recorded is refused by name.

The last was the only one whose message mattered and whose message was not
checked — the database's own foreign-key error is true and mentions a
constraint, which sends somebody to read a schema instead of typing the name
they meant.

Partial sealing now says how far it got. The person holding the key is the only
one who can finish, and running it again knowing what it will do is different
from running it hoping.
2026-08-31 02:59:39 +02:00
jschoubben fbae2f1f9f Build everything behind its source, in one command
novox/hq ADR 0010 replaced a pipeline with a comparison, and named the risk:
losing the question "did my change go out?". The mesh could already answer
which modules are behind their source — and then a person read that list and
retyped each repository, which is a person being the loop, and the loop is the
thing the pipeline was doing before it was taken away.

The mirror of `push --behind`, with the same argument and the same refusal to
combine the two forms: naming a repository and asking which need building are
different requests.

One failing does not stop the others, for the same reason one broken module no
longer blocks a machine's whole declaration: a mesh where one bad repository
holds back nine good ones is a mesh where nobody dares add the tenth.

Each is built from its own recorded ref rather than the commit the mesh
happened to notice — pinning to that would quietly turn a tracked branch into
a pin.
2026-08-31 02:55:10 +02:00
jschoubben 87c6a56b81 Model access is a provision answered by a record, not a machine
novox/hq ADR 0024, gaps 1 and 2. The user's stated requirement, and the first
thing here that no machine can answer: a hosted model is on nobody's node and
is reached over the public internet, so the rule that refuses two ends sharing
no private network must not apply to it.

A licence is a named thing and the name is the operator's — *the personal
account*, *the organisation's* — because the whole point is saying which one a
given consumer uses, and an anonymous credential hanging off a provider cannot
be said. Many to many, so deliberately not a claim: two machines sharing an
account is ordinary rather than a collision.

Gap 2 is the missing verb, *accept*: take a value somebody supplied, seal it to
each holder, discard the plaintext. With the consequence stated rather than
hidden — a holder recorded after the key was supplied has no key and the mesh
cannot make one, so it is refused by name with the remedy, not silently handed
an empty file.

Refusal is felt, as the record warns: a mesh holding three ways to reach a
model refuses every consumer that has not chosen. So the refusal names the
candidates and the exact command. Being right is not the same as being usable.

Gaps 3 and 4 — a consumer that is not a machine, and switching as a reaction
rather than a declaration — remain gaps. Half-building them would put a
conditional in the declaration language, which is what ADR 0024 says plainly to
avoid.

Its own context, with its own store and its own credential: a licence is a
different aggregate from anything inventory owns, and it refers to nodes by
name because that is what crossing a context boundary may carry.
2026-08-31 02:50:38 +02:00
jschoubben d0c0ee8dab A route is a grant, and a provider is told where its consumer is
novox/hq 08-connectivity §3, built. The mirror of a database grant: there the
consumer supplies a name and receives credentials; here it supplies a target
and receives a name. Nothing new in the vocabulary — a route is a provision
like any other.

One field was missing and it is the one that matters for anything reaching
back: a contribution now carries where the mesh says that machine is. A reverse
proxy is told to send traffic to a consumer and has to open a connection, so
without it every provider implementing a provision would have to know how the
mesh names machines — a convention leaking into every module.

The proxy itself is an example, not part of the control plane: the contract is
the file, not this program. It replaces its table whole rather than merging,
because the file is the whole truth about who has a route and merging would
keep serving a name whose module was unassigned — the stale-route fault
08-connectivity lists as open, reintroduced one level down. A name it does not
serve is refused by saying which it does: a route withdrawn and a name that
never existed are different things.
2026-08-31 02:43:19 +02:00
jschoubben ebcfd37b92 Rotate a credential and move both ends together
The invariant novox/hq ADR 0001 records as unowned, and it was measurably
false in HAL: a provision documented as never rotating minted a new password on
every adoption and updated only the provider's row. Consumers on three nodes
held dead credentials for two days while the mesh reported success. Nothing
enumerated who held the old one.

Three things make that impossible here. The holders are a set the mesh can name
— each pair has its own credential, so rotating one consumer touches one role
and the affected list is a query rather than an assumption. Both ends are
pushed by this command rather than a later one, because leaving the sending to
whoever remembered is the fault exactly. And it is all-or-nothing: if any
affected machine cannot be resolved, nothing is sent and the old credential
keeps working, which is a mesh that has not rotated rather than one that has
half-rotated.

The window is stated rather than hidden: a role's password changes on the
provider and the file changes on the consumer, and they cannot be simultaneous.

The provisioner now takes its superuser password from the file the mesh wrote,
which is how the mesh delivers one. Passing it through the environment needed a
person in the middle of the one path that exists so there is not one — and put
a superuser password where `docker inspect` prints it.
2026-08-31 02:38:52 +02:00
jschoubben e2916ee3db The pin is compared in the spelling the mesh writes it
The builder hashed the certificate and compared a bare digest against a
fingerprint written as `sha256:` followed by 64 hex characters. It could never
match — and it failed as "this is not the broker this builder was told about",
which is the one thing this check exists to report truthfully. A check that
cries wolf on every correct broker is worse than no check, because the first
thing anybody does is remove it.

The error now prints what was expected beside what arrived, the way the host's
has always done: without both, the message describes a mismatch nobody can
confirm.

And the pin check has its own test, driven against a real TLS handshake — it
accepts the certificate whose fingerprint the mesh wrote and refuses another.
A pin only ever exercised through a live broker is a pin nothing tests.
2026-08-31 01:45:09 +02:00
jschoubben 8b2d1bce9d Something answered on this machine is still bound
A binding was skipped when the provider turned out to be on the same node,
reasoning that a file saying "it is on this node" is a fact nobody needs. That
is right about the location and wrong about everything beside it: a binding
also carries what the provider said a consumer must know, which is the port,
and a consumer cannot invent that.

A build machine sharing a node with the registry it pushes to sat in a loop
saying it could not read its own binding. Nothing was wrong with the machine,
the module, the credential or the provision — the file was never written, and
the absence looked exactly like a mistake in the module.

The original intent is kept where it was right: a provision whose provider said
nothing a consumer must know is still not written. A shell is answered here and
there is nothing to say about it. A registry is answered here and the port is
still unguessable.

The address is this machine's name on the private network, or loopback when it
has none — a machine off the network still reaches itself, and a name nothing
resolves is worse than an address that always works.
2026-08-31 01:35:23 +02:00
jschoubben a5b11fd6b2 A build machine is told what to check the broker against
The credential was a URL and nothing else, so the builder verified the broker
the ordinary way — against public roots. A mesh's broker presents a certificate
of the mesh's own, which is in no trust store anywhere, so the connection could
only ever succeed against a broker somebody else vouches for. It failed at TLS
with an error about an unknown authority rather than about a missing pin, and
the container sat there running: up, credential on disk, connected to nothing.

So the sealed credential now carries the URL and the broker's fingerprint —
the same two facts a node's token carries, for the same reason, delivered out
of band relative to the thing being trusted. The builder pins it: the standard
chain check is replaced rather than removed, and what replaces it is stricter,
accepting one certificate instead of every certificate a public authority
would sign.

A file holding only a URL still works, for a builder somebody runs by hand
against a broker with an ordinary certificate.
2026-08-31 00:55:33 +02:00
jschoubben 1eb1b69cae What the mesh computes is applied before what the module declared
The host does not sort — order is stated (novox/hq ADR 0005) — so the order the
mesh writes down is the order a machine applies. Certificates, credentials,
bound files and the rule set were appended after a module's own resources, so a
service or container that depends on one was applied before it existed.

It failed and the next reconcile fixed it, which is why nothing caught it. A
fault that repairs itself on the second attempt is worse than one that does
not: what gets remembered is that it works.

Nothing the mesh computes depends on a module's resources, so putting all of it
first is unconditionally right. Merged after the computed-resources branch,
which replaces a module's resources wholesale and would otherwise discard them.
2026-08-31 00:37:01 +02:00
jschoubben d9bee18d44 A mesh on both address families renders both
nftables matches ip and ip6 separately and one set holding both is a syntax
error, so the file would not load: the service reports a configuration fault
and the machine filters nothing. Also a make target for the builder image,
which the lab now stocks.
2026-08-31 00:34:20 +02:00
jschoubben 58c8ab7747 A secret the mesh was given is not one the mesh can reinvent
Two kinds live in module_secret and they behaved identically, which is right
for one of them. A made secret is the mesh's: when a node regenerates its
sealing key the mesh makes another and nothing is lost, because nothing else
ever knew the old one.

An accepted secret is not. A broker account's password exists because the
broker was told about it. Regenerating one puts 32 random bytes where a working
credential was — and the machine applies it, reports success, and the program
reading it fails to authenticate somewhere else entirely, with the mesh
insisting the secret was delivered, which it was.

The row now records where the value came from, and a rejoined machine asking
for an accepted one is refused with the remedy named: issue it again. No amount
of pushing produces a password the broker has never heard of.

Found while making the builder a module, which is the first thing to hold one.
2026-08-31 00:32:11 +02:00
jschoubben 48735171eb A machine's filtering is computed from what it was assigned
A rule nobody derives is a rule somebody keeps in step by hand, and five HAL
manifests carry a `scope:` key that reads as a restriction and restricts
nothing. Both halves are closed here.

Manifests are parsed strictly. An unknown key is refused, which is the
discipline the host's declaration parser has always had; `scope:` survived
because nothing rejected it.

A module says what it listens on and who may reach it, and saying from where is
required — a rule with no source is open, and must say so rather than appear to
restrict something. The mesh gathers every assigned module's ports, widens
where two overlap, names every module that wanted each one, and renders one
nftables file per node. What no module declared is closed.

Three things it deliberately does not do: it writes no forward policy, because
what a machine routes is the container runtime's business and dropping there
stops every container on the node; it never flushes the whole ruleset, only
its own table; and it carries no command to load itself, because the link may
not carry an action. A service declares `restart-on` the file instead, which is
the shape that rule leaves.

Also fixes a fault the lab found: certificateFor asked where every node is
without the catalogue, so nothing resolved, every machine looked like it was on
no private network, and every certificate the mesh was asked for was refused
with a reason that was not true. Asking that question without the catalogue is
now refused rather than answered wrongly.
2026-08-31 00:25:12 +02:00
jschoubben 646609c1b2 The mesh certifies names inside it
08-connectivity keeps two authorities apart on purpose: a public one for
names the outside world reaches, and the mesh's own for names only the
mesh knows. Nothing implemented the second, so anything between machines
was plaintext or trust-on-first-use — which the design refuses everywhere
else.

A node now generates a fourth key at enrolment and reports the public
half. A fourth, because a key used for two purposes is one rotation away
from breaking the other: the identity key signs messages to the mesh and
would do for TLS, and reusing it would mean rotating a node's identity
every time its certificate is replaced.

**Nothing secret travels and nothing is sealed.** A certificate authority
says "this name belongs to the holder of this key", so the mesh signs a
public half it cannot use, and the certificate it issues is public. A
module asks for one and is given the certificate and, if it wants,
the mesh's own — the private key is a path to a file the machine already
has, the same arrangement the private network's key uses.

Asserted by verifying rather than inspecting, because a certificate that
parses and does not chain fails at the moment something connects:

- what the mesh issues verifies against the mesh, for the name asked for
- the name is in the subject alternative names, since a certificate
  carrying it only in the common name is refused by every modern client
- it certifies the key the node generated and no other
- another mesh's certificate does not verify, which is the whole point of
  two authorities being separate
- the authority cannot sign another authority — one that could is one
  that can be delegated without anybody deciding to
- two control planes starting together agree on one authority, or a mesh
  has certificates half its machines refuse

Certificates last ten years, which is a choice: a short life needs
something to renew it, and a renewal that fails silently is a mesh that
stops trusting itself on a date nobody wrote down. What makes one
replaceable is that the mesh reissues on demand, not that it expires.
2026-08-31 00:09:13 +02:00
jschoubben 0262873254 status --json, so a board has something to read
A board reads through interfaces and holds nothing. Everything it needs
is already answered — as text, for people, which is not something a page
can read.

`--json` rather than a serving API, because nothing needs one yet:
whatever serves a board runs the command, and the constraint holds either
way — the board never touches a context's store. An API is the larger
thing and should wait until something asks for it.

Both forms are gathered from the same reads before either says anything,
so they answer the same questions rather than being two implementations
that can drift. That was not true of the first version: the JSON printed
after the text, because the branch was too late.

Four properties, each asserted and each confirmed to fail when removed:

- refused and failed stay distinct all the way out. They are fixed in
  different places, so one word for both sends half a page's readers to
  the wrong one — and how much DID apply is carried, since "three of
  eight" and "none of eight" are different machines
- a machine that never spoke carries no time at all, rather than a zero
  one that any page would format as a date in 1970
- nothing is null. A page distinguishing "no machines are wrong" from
  "this field is missing" has to handle both, and null is the one that
  gets forgotten
- no field is named like a secret. Everything here comes from records
  that hold no readable one, but a shape a page is built against is
  exactly where one would eventually be added for convenience
2026-08-30 20:22:04 +02:00
jschoubben 79d6ade4c8 push --behind, and a builder told where to publish
`status` says which machines are not doing what they were told, and
nothing acted on it: a machine that refused or failed stayed wrong until
somebody ran push again naming it.

`push --behind` sends only to machines whose last report was not a clean
apply. A command rather than a timer, deliberately: a scheduler is then a
scheduler over this, where building the scheduler first would have meant
two paths to one act with nothing to compare them against.

Naming a machine and asking which machines need one are different
requests, so `push <node> --behind` is refused rather than guessed. With
nothing behind it says so, because "nothing needed one" and "this did not
run" must never look the same. A machine failing the same way for six
hours is pushed to anyway and said about — refusing would leave no way to
retry after fixing the cause, and this is a command somebody ran.

Proven in the lab: a machine is broken with a package that does not
exist, `push --behind` names it and not the machine that is fine, the
module is corrected, and the machine recovers without anybody naming it.

And the builder can be told where to publish rather than configured. A
builder that is a module requires an artifact store, and the mesh writes
it the same binding any consumer of any provision gets. A binding with no
address is refused rather than falling back to anything — that would
publish to a store on the wrong machine and be found out much later. The
variable remains for a builder run by a person, which is how it is still
run while being developed.
2026-08-30 19:47:02 +02:00
jschoubben 45c3853f4e A consumer that stopped asking is withdrawn
Found by testing removal, which is the half nobody tests.

A grant was emitted for every secret the mesh held, whether or not the
machine still asked for it. So a consumer that was unassigned kept
appearing in its provider's manifest — and the provisioner's rule about
removing what nobody asks for can only fire if the mesh stops asking. The
login would have stayed live for ever, and nothing would have said so.

Skipped where the declaration is built rather than where grants are
gathered, so the rule holds whoever gathers them. No credential file is
written for a withdrawn consumer either, or the provisioner would find a
file its manifest does not mention and have to guess what that means.

The secret itself is deliberately kept. It is sealed and unusable to the
mesh, and a machine that comes back gets what it had — what withdraws the
login is the manifest, which is the thing that reconciles.
2026-08-30 19:17:13 +02:00
jschoubben 15fd70e3ce A module may mirror an image it did not write
A module usually runs software somebody else built: a database module
ships configuration and a provisioner and does not build a database. It
could name the upstream reference directly, and then every machine needs
a route to a public registry and the reference is a tag somebody else can
move — which is what pinning exists to prevent.

So an artifact may be `upstream`: pulled by the reference the module
names, pushed into the mesh's own registry, and pinned by the digest that
registry assigns. This is what the bootstrap already does by hand; it is
now something a module can say.

Refused: an upstream reference with no tag or digest, because what gets
mirrored would be whatever `latest` means today and a module pinned to
that is not pinned. And the rule that a build reads only its own
repository does not apply to it — applying it anyway refused every
reference with a registry host in it, which the test caught.

Written by trying to write a real postgres module and finding it could
not be said. It can now: two directories, two containers pinned by
digest, a superuser password sealed to the machine, and the grants
manifest — six resources from one assignment, all accepted by the host's
own parser.

That exercise also found my manifest wrong rather than the host: a
container declared `restart-on`, which is a service field, and the host
refused it by name. It is right to. A container whose own definition
changes is recreated, and a file it mounts is read by the process inside,
which is that image's business.
2026-08-30 18:28:05 +02:00
jschoubben c37d368f65 A module may need a secret of its own, and the provisioner watches
Two things, both found by trying to write a real postgres module and
discovering it could not be said.

A database has a superuser password, a broker an administrator, a
registry an account. None of them is *for* anybody — they are not the
credential a consumer is given, and the mechanism that hands those out
has a consumer in the middle of it. So a module may declare what it needs
and where to put it, and the mesh generates one per node, seals it, and
reads it no more than it reads any other.

Per node, deliberately: a module running on three machines has three
passwords. One in the manifest instead would put the same secret on every
machine that ever runs it, in a file anybody can read, for ever. Made
once and kept, or a running database would be handed a password it was
not started with; remade when the machine's sealing key changes, like
everything else sealed here.

A need declared and not made is refused rather than skipped, because a
module whose own credential is silently absent starts, fails to
authenticate, and the reason is three layers from the machine reporting
it.

And the provisioner can watch. That is what lets it be a module rather
than a binary somebody places: run once, it needs invoking after every
declaration by a timer or a unit wired to a file; watching, it is an
ordinary long-running service the host already supervises. It polls
rather than watching the filesystem, because the host writes atomically —
the file is replaced, so a watch on the path stops seeing anything after
the first replacement, and a watcher that silently stops working is worse
than a poll. Credentials are compared by digest and never held: this runs
for as long as the machine is up.
2026-08-30 18:22:05 +02:00
jschoubben 9681b288aa Keep what each machine did, so status can say what is wrong
A node reports back after applying a declaration: it worked, some of it
failed, or the whole thing was refused. A refusal or a failure moved
last_seen and the reason went to a log line — so "which machine is not
doing what it was told" had no answer the next morning, which is the
question a mesh exists to answer.

Refused and failed are kept as different things, because they are
different situations with different remedies: refused means the machine
is exactly as it was and what is wrong is in what was sent; failed means
it is in a state nobody declared and what is wrong is on the machine. One
word for both would make the record say less than the node did.

One row per node, replaced. The question is the machine's current state —
"this failed an hour ago and then succeeded" is not a machine anybody
needs to look at, and a table of every report would bury the ones that
matter under the ones that do not.

`status` now answers three questions in the order somebody asks them: is
anything broken, is anything not answering, is anything out of date. The
first has consequences now, the third is a plan for later, and a status
leading with the third would bury the first. A machine that has never
spoken is reported as quiet rather than as broken — new, switched off and
unreachable are not the same as tried and could not.

The mapping from a report to an outcome had no test at all, which the
injection caught: it is the code deciding which of those situations a
machine is in. It has four now, including that a partial report never
becomes the account of what the machine holds — the fault that destroyed
a substrate once.
2026-08-30 18:08:59 +02:00
jschoubben 3195634441 A build machine gets its own credential, scoped to build work
The builder was documented as holding its own broker credential and
nothing else, and nothing issued one — so in practice it used whatever it
was handed, which was the broker's administrative account. A program
documented as holding its own credential and given somebody else's is
worse than one with no story at all.

`builder issue <name>` creates an account that may read the build queue
and write to the mesh exchange. Not a node account: a build machine is
not a node, and a node's queue carries its declarations.

Two faults found by running it, both about the answer path:

- the reply queue was left for the broker to name, and the account was
  scoped to `amq.gen-*` — one broker's convention. The builder built,
  could not answer, and the connection closed. Reply queues are named
  here now, deterministically.
- the answer then went via the DEFAULT exchange, where permission is
  granted per exchange rather than per queue. A builder allowed to use it
  could publish into any node's queue, which is the privilege a build
  machine most obviously should not have. Answers go through the mesh
  exchange, which it already may use, and an asker binds its reply queue
  to the same key and filters by correlation.

Verified against a real broker: a builder cannot consume a node's queue
and cannot publish to the default exchange. That check nearly reported
the opposite — an unconfirmed publish is asynchronous, so the refusal
arrives as a channel close afterwards and a naive test sees success. With
publisher confirms it is immediate. A negative security assertion made
against an asynchronous call is not an assertion.

Redelivery was observed working while fixing this: builders that died
before answering left their work on the queue, and the next builder did
all of it.

Also: the queue and exchange names exist in both `broker` and `link`,
because `link` imports `broker`. A test in an external package keeps them
agreeing — a builder scoped to a queue nothing publishes to takes no work
and says nothing about why.
2026-08-30 10:28:41 +02:00
jschoubben 0bbb5c6838 Builds have a history, and failures are rows like any other
A build result was answered to whoever asked and kept nowhere. So "when
did this last build", "why did it fail" and "which machine built what is
running" had no answer, and a build nobody was waiting for was reported
into the void — which is the same as not reporting it.

Failures are recorded too, and that is the point rather than a detail: a
failed build that leaves no trace is indistinguishable from one nobody
asked for, and the difference is the whole of whether somebody should be
looking at something. A build that never learned what it was building
keeps the repository, because that is what a person goes and looks at.

Recording is idempotent on the correlation id, because a result can
arrive twice — as the answer to whoever asked, and on the exchange when
nobody was. Two rows would show one build as two, and which is real is
not answerable afterwards.

The serving control plane now binds `built` as well, so results from
builds it did not ask for are kept. It refuses them loudly when it has
nowhere to put them rather than dropping them, so the broker's own
counters show something arriving that nothing handles.

`builds [<module>]` reads it: what happened lately across the mesh, or
what has happened to one module — the first asked after something goes
wrong, the second when deciding whether to trust something.

What was published is kept with the build, so a digest traces back to
what made it without holding the manifest twice in a place that can
disagree with the first.
2026-08-30 10:18:23 +02:00
jschoubben 78c5b653cf A secret the mesh is given, not one it made
The last of the four gaps ADR 0024 names. Everything the mesh handles
today it generated itself, sealed to both ends, and discarded. An API key
for a hosted service comes from a person, and carrying it needs a verb
the mesh did not have.

Accept seals it on the way in and keeps no plaintext — the same storage
and the same property as a generated one, only a different origin. That
is the whole difference from the arrangement being replaced, where an
operator-supplied key sits in a column the control plane can read, which
makes a copy of the database a copy of every account the mesh touches.

The consequence is deliberate: the mesh cannot show it back. Somebody who
loses the key gets a new one from wherever it came from. There is no
reveal and there cannot be one, because a mesh that can reveal a secret
is a mesh that holds it — asserted as a test, because it is a property
somebody will eventually ask to break.

An empty value is refused. A credential that exists, authenticates
nowhere and looks exactly like a working one is the failure this whole
mechanism is arranged to prevent.
2026-08-30 03:47:04 +02:00
jschoubben 421fe73dce The mesh builds: a machine takes the work, and the catalogue shows it
A build is work, not state. Everything else the control plane sends a
node is a declaration — this is what you should be — reconciled forever.
A build happens once and is finished. Putting it in a declaration would
mean rebuilding on every reconcile, or a declaration carrying "and I
already did this", which is state about an event rather than about a
machine.

So it travels on its own queue and the answer comes back correlated. One
queue, so several build machines share the work and each request is done
exactly once — which a per-machine routing key would not give.

mesh-builder is the program a build machine runs. Not the control plane,
which must not run commands on a machine; not the host, which would then
need a container runtime and git everywhere to do something almost no
machine will ever do. It holds its own broker credential and nothing
else.

Three properties that are decisions:

- a request is acknowledged only once the answer is away, so a builder
  that dies mid-build leaves the work for another machine rather than
  losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five
  slower than it would have finished the first, and the queue is what
  shares work between machines
- a failure is a RESULT. A build that fails silently is
  indistinguishable from a builder that is not running, and those want
  different responses

And `module list` is a catalogue: what exists, at which version, built
from which commit or handed over by hand or shipped with the control
plane, whether it is behind its source, and which machines run it. All of
that was recorded from the first build and none of it was shown, so "is
this current?" could only be answered by reading the database.

Proven against a real broker, registry and store: the mesh asked, a
builder consumed, built, published, answered; the manifest was recorded
with its commit; the source moved and the catalogue said "behind";
rebuilding caught it up with a new digest because the content changed.
2026-08-30 03:46:02 +02:00
jschoubben 7d033ad9f6 Publish to the registry, and a command that builds a repository
One store, and it is the registry the bootstrap already pulls from. An
OCI registry is a content-addressed blob store that also understands
images: PUT a blob and it is retrievable at /v2/<name>/blobs/sha256:… for
ever, by digest. An archive is a content-addressed blob.

A second store beside it was considered and is the right answer for
objects that are mutable, need per-reader access, or are not build output
— somebody's uploads, a backup, a thing with a lifecycle. None of that
describes a digest-pinned archive, and running a second service to hold
one kind of immutable blob is two things to run, two to back up, and two
ways for an artifact to be missing. Overturnable by reading: the manifest
carries a URL and a digest, and neither says what served it.

`build <repository>` clones, reads module.json, builds what it declares,
publishes, and records the manifest with the commit it came from. It is a
command rather than something the control plane does on its own, because
building runs things on a machine and what the control plane may send a
machine is bounded by the declaration language. This is the shape the
builder module takes when it is given work over the broker.

Proven end to end on a real repository and a real registry: a shell
module with a package, a user and a dotfile archive built, published,
fetched back at the digest it declared, rebuilt to the same digest, and
its manifest accepted by the host's own parser — including `user` and
`archive`, which did not exist this morning.

A tag is never accepted as a pin, and a blob already stored is not sent
again — it is named by its content, so re-uploading asks the registry to
store what it already has under the name it already has.
2026-08-30 03:36:04 +02:00
jschoubben 604b04886b A builder: a repository becomes artifacts the mesh can pin
It runs on a node, not in the control plane. Building needs a container
runtime and a working tree, and the control plane deliberately cannot run
commands on a machine — what it may send is bounded by the declaration
language, and "run this build" is not in it. So the builder is something
a node runs as a module, given work over the broker like anything else.
The alternative, the control plane holding a docker socket, would make it
the one component that can do anything anywhere, which is the property
the whole design is arranged to avoid.

A module repository has one file at its root, module.json, saying what it
is and what it builds. A convention somebody can look for beats a setting
somebody has to find.

Properties that are decisions rather than details:

- a fresh clone every time. A build reusing a working tree can succeed
  because of something a previous build left behind, and that is a build
  nobody can reproduce.
- archives are packed deterministically — sorted, and carrying no
  timestamps, uid, gid or original names. Two builds of one commit must
  produce one digest, or nothing downstream can tell "this changed" from
  "this was built again", and every rebuild looks like a change to every
  machine holding it.
- nothing is published until everything is built. Half a module in the
  store under a digest the mesh never records is reachable,
  unreferenced, and indistinguishable from something in use.

The reproducibility test was passing for the wrong reason: both builds
landed in the same second, so a packer carrying timestamps would still
have agreed. It now stamps the two trees a year apart, and a timestamp in
the header breaks it.

One line is honest about not being independently tested: the sort before
packing is belt and braces over filepath.Walk's documented lexical order,
and no injection can distinguish it.
2026-08-30 03:32:46 +02:00
jschoubben 44ba100595 A module says what it builds, and the built manifest is a different
document

The manifest in a repository names artifacts; the manifest the mesh holds
names digests. Keeping them the same file would mean a repository
carrying a digest — wrong the moment anybody edits anything, and pinning
a value nobody could have checked.

So a resource says `"artifact": "server"`, and resolving a build rewrites
it to the image reference or the archive's source and digest, removing
the build-time word entirely. The host has never heard of an artifact and
its strict decoder would refuse one, at the worst moment.

A module that builds nothing is ordinary and needs no build section —
most of what a person installs is configuration, and a field that exists
to be left blank is a field nobody fills in correctly.

Refusals worth having:

- an artifact declared and not produced blames THE BUILD, not the
  resource. Both are failures and the remedies are in different places;
  telling somebody to fix the wrong one costs an afternoon. Found by
  injection: the first version's message could not be told apart from
  the resource-level one, so the check was not actually tested.
- a build reads its own repository and nothing else. An input path
  leaving it makes what gets built depend on whatever happens to be on
  the machine building it.
- two artifacts with one name, because a resource naming it could mean
  either.
2026-08-30 03:29:17 +02:00
jschoubben d978712d7f Split resolving from rendering a declaration
resolve.go had grown to 796 lines doing four jobs: working out what a
machine should run, applying settings, collecting contributions, and
placing credentials. They answer different questions — the first is
"what", the rest are "what does that look like as resources" — and one
file doing both is how a thing starts becoming the kernel everything
imports.

Prompted by looking at why HAL's shared library became unmaintainable.
Measured while here, and the shape is the inverse of that one: the large
packages import nothing internal, and only inventory and link compose. A
change to module resolution cannot reach connectivity, because
connectivity does not import it.
2026-08-30 02:54:51 +02:00
jschoubben 02d1020bce The token carries the node's name
Found by raising a mesh end to end. The broker account a joining node
authenticates as is named after the node, and exists before that machine
has been told anything — so the node has to know its name before the mesh
can tell it. Without it, enrolment fails at the broker with an empty
username, which says nothing about why.

Not a secret, and the issuer already knows it. The wire-format test now
covers it, so a rename on either side fails in both repositories rather
than at enrolment on a real machine.
2026-08-30 02:36:57 +02:00
jschoubben b9aac2b700 Check that what a node says when it joins is what this mesh reads
The enrolment request is a struct in each repository. A node now reports
a third key — the one its secrets are sealed to — and that wiring had
unit tests on each side and had never been run across the join. A field
renamed on one side fails silently: enrolment succeeds, the key is
absent, and the node looks joined until the first thing sealed to it
cannot be opened, by which point nobody is looking at enrolment.

So the host's suite writes a real request and this one reads it, the same
way the declaration check already runs in the other direction. Both skip
with a reason when the neighbour is not checked out.

It does more than compare shapes: it seals something to the key that
arrived and opens it with the private half the host kept. Confirmed to
fail three ways — a renamed field, a value that is not a key, and a key
that is present, correctly named and simply somebody else's. Only the
last needs the sealing step, and it is the one a shape check would pass.

Also `inventory.ForTest`, because the check lives beside the link and a
second copy of the throwaway-database helper would be a second thing to
keep true.
2026-08-30 02:20:28 +02:00
jschoubben c3046dcf56 A provider is told who its consumers are, and a reference provisioner
Contributions were node-local, so a mesh-scoped provider — the one case
that most needs them — never heard from its consumers. A database was
given a password and no idea what to create it for.

Cross-node consumers now reach the provider's `receives` file, merged in
with the ones on its own machine: from the provider's side they are the
same thing, and a provider that had to read two lists would read one of
them. Each names the file its credential is in rather than carrying it,
because the mesh discarded the value and could not put it there. The
readable half therefore stays readable.

And examples/postgres-provisioner, which is the last step: it reads what
the host wrote and makes PostgreSQL accept it. Explicitly not part of the
control plane — the control plane decides and never touches a machine.
This runs on the machine and touches it, and a real one ships with the
module that ships PostgreSQL. It lives here because this is where the
contract is defined, written as something that runs so it can be read.

It reconciles rather than applying a change, because it is never told
what changed. Three things that follow, and each is a fault somebody has
shipped:

- the password is set every time, not only on creation, or a rotation
  reports success and changes nothing
- what it made and nobody asks for any more is revoked, or a departed
  consumer keeps a working login for ever
- what it did not make is left alone, or it cannot be run on a database
  that predates it

Proven in the lab against a real PostgreSQL, each assertion confirmed to
fail with the behaviour removed. The suite is in mesh-lab, which also
records the two ways the test itself was wrong first.
2026-08-30 01:31:25 +02:00
jschoubben 20f78cd5f1 Credentials the mesh delivers and cannot read
HAL keeps env vars in the registry, encrypted at rest. Its own tooling
records what that bought and what it did not. `secret_locate` matches by
value rather than by name — because the same password sits in
mesh_provisions, in module_env, in each node's .env in plain text, and
inside every connection string composed from it, and its documentation
says those URL copies "are often the only copies actually in use". And a
query against the encrypted column returns zero rows and proves nothing,
so auditing moved to the decrypted copies on the nodes.

Two faults there, and encryption at rest addresses neither: the control
plane can read what it stores, so a copy of the database is a copy of
every credential; and one secret has many homes with nothing tracking
them.

So here the mesh generates a password, seals it to each end with keys
those nodes generated, stores both blobs, and discards the plaintext. It
cannot read what it holds. Neither can the broker relaying it. And
nothing is composed centrally — a connection string is assembled on the
machine that needs one — so no copy is ever minted in a shape nothing
tracks. `Compromise of a node is compromise of that node` (ADR 0004) is
now true of secrets, not only of identity.

Two files rather than one, because the mesh cannot compose a document
containing a value it discarded: `binds` carries the readable facts,
`secrets` carries the credential alone. The readable half stays readable
in the declaration; the secret half changes only when the secret does,
which makes restart-on precise. The provider gets a directory, one file
per consumer, for the same reason.

It is made once and kept — regenerating per declaration would restart
both ends on every push, and the password a provider was told to create
would never be the one its consumer was given. It is remade when either
end's sealing key changes, and both ends learn the new one in the same
push, so there is no window where half the mesh holds a dead credential.

Two tests found passing for the wrong reason, both caught because their
injection came back clean:

- the provider's copy was asserted non-empty, which reads the same
  whichever column is selected. It now opens the blob with the
  provider's own key.
- RotateSecret deleted and re-created; the re-create was dead, because
  the next read makes one anyway. Removed, and a second path to the same
  act is how two ends come to disagree.

And one real fault: three places built a declaration, and the one behind
`--json` predated credentials, so it silently produced a declaration
missing them — a difference between what `plan` showed and what anything
reading `--json` got. There is one path now.
2026-08-30 00:21:18 +02:00
jschoubben c4782ae2fd An app is told where its database is
Knowing that a machine needs the anchor's database is useless to the
program that needs it unless the program is told. It knew; nothing was
written anywhere it could read.

Two fields, mirroring contributes/receives in the other direction:

  serves: {database: {port: 5432, driver: postgres}}   on the provider
  binds:  {database: /etc/app/database.json}           on the consumer

The provider says what a consumer needs to know; the mesh adds the half
only it has — which machine, and what that machine is called on the
private network. The file says, in itself, that it carries no credential
and why. A missing field looks like a bug; a stated absence looks like a
boundary.

Binding something answered on this machine writes nothing. A file saying
"it is on this node" is a fact nobody needs and one more thing to keep
true.

And two machines that share no private network are refused rather than
wired together. An app here and a database there with no path between
them is a mesh that reports itself configured and does not work — the
failure surfaces as a connection timing out, which is the slowest place
to find it. This is checkable now only because the network became
something a machine is given rather than something it has by having an
address.

One fault, found by running it: working out who is on the private network
resolved the mesh, and resolving the mesh asks who is on the private
network. It hung for two minutes. The comment above the function said not
to do that and the function did it anyway; it now resolves each node
locally, which is the right answer to the question regardless — whether a
machine is on the network depends on what it was assigned, not on what it
takes from others.
2026-08-30 00:02:18 +02:00
jschoubben d4064122d6 Where the answer to a requirement is allowed to live
Two different things were both written `requires`. A shell, a display
server and a private network have to be on the machine that needs them.
A database does not — it runs somewhere and is reached over the network.
Both were answered the same way, so requiring a database installed
PostgreSQL on every machine that ran a web application.

What a module provides now carries a scope, the same idea claims already
use, written short in the ordinary case:

  "provides": ["shell"]
  "provides": [{"name": "database", "scope": "mesh"}]

A mesh-scoped requirement is answered by finding the node already running
it — never by installing it here. Choosing a machine to put a database on
is a decision with consequences, and nothing resolving a web application
should make it silently. With nothing anywhere it refuses and says which
module to assign; with two it refuses and says how to choose.

Choosing is `pin <node> <provision> <from>`, kept per node because that
is the granularity the choice has. A pin at a machine that does not
provide it refuses rather than falling back — a fallback would quietly
move somebody's data. One provider does not overrule a pin either.

Resolving a node now needs to know what the others offer, and working
that out needs them resolved, so it is two passes: the first answers only
what each node offers, the second answers everything. Nothing is ever
declared from the first.

A node's plan says what it takes from elsewhere. It is the only part of a
set that stops working when a different machine goes away, and nothing
else in that output would have said so. It is also where a credential
will hang once there is a mechanism for handing one back.

One test found passing for the wrong reason: it read pins through a join
on the provider, which hides a dangling row whether or not it was cleaned
up. It counts rows now, and bites when the cascade is removed.
2026-08-29 23:51:50 +02:00
jschoubben 5a3a87e8c3 A module can tell its provider what it needs
`requires` said a thing must be there. It never said what to do with it,
so a web application requiring a reverse proxy had nowhere to put "this
name, this port". The two modules that needed it most went round the
outside and opened a connection to the control plane's database, which is
why every node holds a credential to it permanently.

Two fields close it:

  contributes: {reverse-proxy: {host: board, port: 8080}}
  receives:    {reverse-proxy: /etc/traefik/dynamic/mesh.json}

The control plane collects every contribution on a node and writes them
to the path the provider named, ordered by module so the file does not
churn. Contributing to something is requiring it — asking to be published
means a publisher must exist, and a module that had to say both would
eventually say one.

The control plane does not know what a reverse proxy is and does not
write one's configuration. It delivers facts; the module turns them into
whatever it runs. That is why swapping the proxy touches nothing that
publishes through it, and why the host needs no new vocabulary — a
received file is a file.

Settings reach a contribution the same way they reach a file, because a
hostname is exactly what differs between one mesh and the next.

Two things found by running it:

- the file had a `//` header, so it said "do not edit" to a person and
  failed to parse for the program meant to read it. The note is inside
  the document now.
- a provider with no consumers gets an empty file rather than none. It
  cannot otherwise tell "nothing asked for me" from "the mesh never
  wrote it", and those want different responses.

Also `plan <node> --json`, which is how the declaration gets handed to
the host's own parser.
2026-08-29 23:35:43 +02:00
jschoubben 44d134ba25 Networking is a module, and a domain module is how you avoid choosing
Connectivity was code beside the module system doing the module system's
job: every machine with an address was on the private network and there
was no way to keep one off.

A manifest can now say its resources are computed by the control plane,
which is what a peer list needs — it is derived from every machine at
once, so nothing could be written in advance. The network is a module
from there on: assigned, resolved, settled, and absent from a machine
nobody gave it to.

Three modules rather than one, because WireGuard is one VPN of several:

  mesh-wireguard   provides private-network, mesh-addressing
                   claims the-private-network, one per node
  mesh-names       provides name-resolution, requires mesh-addressing
  networking       requires both, and ships no files of its own

The last is the point. Most people want the network up and do not want
to choose a VPN, so `assign networking` takes the only answer to each
requirement silently. The day the catalogue holds a second one there are
two answers, the resolver refuses and names them, and choosing is
assigning the one you want. No flavor field, nothing to configure.

Names left the WireGuard declaration for their own module. They would be
identical over a different private network, and bundling them made one
module out of two things.

Three faults the walk found:

- choosing tailscale still installed WireGuard, dragged back in by the
  names needing the mesh's own addresses. Caught now by a claim: running
  two VPNs is fine, being *the* mesh network is singular.
- a requirement wanted by two modules was reported twice, identically.
- "this mesh has no hub" was reported when the real cause was that a
  node could not be resolved at all. It now names the node and the why.

And a test that asserts the manifests actually shipped, after the claim
went missing from the real one while every test stayed green.
2026-08-29 23:19:32 +02:00
jschoubben 65ade756f2 Settings: changing a module's config without editing its file
Managed files are generated and never edited, so somebody's intention about one
has to live where the generator can see it. It does now: the module ships
defaults, settings go over the top by key, and the file is produced from both.
Upstream can rewrite its half freely and the keys somebody chose survive.

Two layers, both from the start. The mesh's settings for a module, then one
machine's over those. A node that differs is expressed by differing, rather
than by restating everything the rest already say -- which would pin all of it
against future changes for no reason.

An override beats a default and there is nothing to resolve. A setting is a
statement about that key made deliberately; the default was only ever what to
do in the absence of one. So when upstream changes a key somebody has set,
there is no conflict, no merge markers, and nothing to ask.

Nested blocks merge and lists are replaced whole. Setting one field of a block
must not delete its siblings, or every setting would restate the whole block
and pin all of it. A list that merged element-wise could neither be shortened
nor reordered, and there is no correct guess about which element is "the same
one".

A module can keep specific keys for itself -- a socket path its own code
depends on -- and setting one is REFUSED rather than ignored. A setting quietly
dropped is somebody believing they changed something.

Settings that reach nothing are named at the moment they would be used, not
discovered later by the machine not behaving differently.

`plan --files` prints what a machine would be given before it is sent, because
"1 resource" does not tell you whether the merge landed.

One test kept with a note that it does not defend this code: output stability
comes from Go's encoder sorting map keys, so it passes with the merging
removed. Worth having as the thing that would catch a change of encoder, but it
is not evidence about anything written here, and it was checked.
2026-08-29 23:01:09 +02:00
jschoubben 653e232f1c The mesh knows where a module came from, and whether it is behind
Delivery is a comparison, not a pipeline: the control plane holds what source
exists and what has been built from it, and the difference is the work. Both
halves are written down now, so "is this current" is a question about two
columns rather than something you find out by building.

`status` answers "did my change go out?", which ADR 0010 names as the real risk
of replacing a pipeline with a comparison -- it is answerable today by opening
a pipeline, and something had to replace that.

  zsh    holds 4f2a9c1e, source has 9e3b7d2a
         running on laptop

The machines are the point. A module being out of date is a fact about the
catalogue; which machines are running last week's version is the thing with
consequences.

Three things this had to get right.

A module with no source is never behind -- it was handed over directly, which
is how a one-off arrives, and saying "out of date" about it would be inventing
a comparison against nothing.

A source nobody has checked is not behind either. Reporting it as behind would
put every module on the list the moment provenance was recorded, which makes
the list say nothing. Fault injection found this: my first test passed with the
guard removed, because both halves were empty strings and compared equal. The
case that actually needed it -- a known commit and an unknown head -- was
untested.

And handing over a manifest by hand does not erase where the module normally
comes from. Fixing something in a hurry is legitimate; silently forgetting its
origin is not, because that record is the only thing that would say afterwards
that a machine is running something nobody can rebuild.

Also fixed the flag parsing, which stopped at the first positional argument and
silently ignored every flag after it -- so `module add thing.json --source x`
recorded no source at all and said it had succeeded. The host's own parser
documents this exact footgun and I wrote it again anyway.
2026-08-29 22:32:16 +02:00
jschoubben 931a3a19a5 Taking a module off a node takes it off the machine
The half of the module system that was built and never proved. Unassigning i3
removed i3's file AND xorg's, because xorg was only there to satisfy i3 -- the
node's own record agrees, and the resolution the mesh sends no longer mentions
either.

That works because a declaration removes what the mesh previously declared and
nothing else, which is 04-ISSUES/010's fix carrying its weight here: the
substrate the machine raised for itself is untouched by any of it.

Tests for the storage layer, which had none. The ones worth naming:

A module a machine is running cannot be forgotten -- not a fault, it means the
mesh would lose the ability to describe what is on that machine. Removing a
node DOES take its assignments, and the asymmetry is deliberate: a node that is
gone cannot be running anything.

A node that has never reported has NO capabilities rather than all of them.
That refuses anything needing one, which is wrong but visible -- where assuming
it can do everything would assign work it cannot do and find out on the
machine. And a capability the node reported as ABSENT is not counted: reading
the list without the verdict would let a module onto a machine that said no.

`overlay push` is gone, replaced by `push`, which sends a node its network and
its modules as one declaration. Two commands that overlap is how a mesh ends up
half-configured by whichever was run. The old name answers with where to go,
and answers before opening a database -- needing one would turn a redirect into
a connection error.
2026-08-29 22:16:11 +02:00
jschoubben 409cd16a09 The mesh decides what a node runs
The gap that has been named at the end of every report for a week. Until now a
declaration came from a person handing over a file; now it comes from what was
assigned, resolved against the catalogue, and the control plane is deciding
rather than relaying.

Everything from the module conversation, built and run on real machines:

  assign laptop i3      -> accepted, brings xorg, because nothing else provides
                           it and there was no choice to make
  assign laptop sway    -> refused: xorg and wayland both claim the-seat
  assign laptop editor  -> refused: three modules provide a shell -- bash,
                           fish, zsh -- choose one
  assign laptop zsh     -> accepted, and the editor's requirement is answered
  bash, fish beside it  -> fine, nothing is claimed

Claims rather than pairwise exclusion, so a third display server would say what
it claims and need no edit to xorg or wayland. Scoped to node, site or mesh:
two DHCP servers at one site collide and at two sites do not, and the mesh-wide
one is the hub said as a claim instead of hard-coded.

Some conflicts cost no manifest field at all. The refusal above names the seat
AND the two files, because the mesh already holds every resource of every
module -- neither i3 nor sway knows the other exists.

Resource identities carry their module, so two modules may both call something
"config" without the second silently replacing the first. What a service
reflects is qualified the same way, or it would name a resource that no longer
exists and stop being restarted when its own configuration changes.

Nothing is sent until every node resolves. A push that configured three and
refused on the fourth would leave the mesh in a state nobody asked for, and the
fourth is exactly where a claim collision appears.

One real flaw found by using it rather than by testing it: assigning zsh did
not satisfy a requirement for a shell. Requirements were counted against the
catalogue without first asking what the set already offers, so "choose one and
assign it" named three modules and then ignored the one you chose. The remedy
was useless and every test passed.
2026-08-29 22:00:06 +02:00
jschoubben f0cff88172 The mesh knows who is out of touch
09-the-node-lifecycle asks for this in as many words -- *how long it has been
disconnected is a fact the mesh must hold, and nothing holds it today. Without
it, a node running last month's assignments looks exactly like one that is
current.* Now it holds it.

`node list` says "here", "out of touch 4m", or "never spoken", and the third is
kept distinct from the second on purpose: a node that has never spoken did not
finish joining, and a node last heard from a month ago is running a month-old
picture of the mesh. Those need different responses from a person.

A bare word that a node is there moves last_seen and touches nothing else. It
is not an account of what the machine holds, and recording it as one would
replace the recovery copy with an empty list every minute -- so a rebuilding
node would then be told it owns nothing and remove whatever it found. There is
a test for exactly that.

Heard is silent in the log. A node saying it is there every minute would fill
the log with the ordinary case, and a log where the ordinary case is loud is a
log nobody reads.

Verified in the lab across the threshold, both directions.
2026-08-29 20:32:27 +02:00
jschoubben fc1417be72 Names, from the same graph as the network
Step 5 of the connectivity order. Every node's internal name resolves to its
overlay address, on every node, computed centrally because it needs every node
at once.

Under `.internal`, which IANA reserved for exactly this in 2024 -- a name there
can never collide with a public one, so an internal name that leaks into a
public resolver fails rather than reaching a stranger's machine. The suffix is
settable for a mesh that wants its own.

Delivered in the same declaration as the peer list rather than a second one. A
node holding the peers and not the names, or the reverse, is half on the
network for as long as that lasts.

This is not the /etc/hosts floor the design removes. That floor existed because
a node had to reach the mesh's database before its own DNS worked -- a fallback
for a circularity that is now gone. This is the mechanism: the complete set of
names, generated whole and owned by the mesh, rather than a patch written
underneath something else. A resolver daemon becomes necessary when names are
wanted that are not one-per-node, and that is not yet true.

A node resolves its own name to its overlay address rather than a loopback,
because a service binding to the name it was given would otherwise listen
somewhere nothing else can reach -- and the failure would appear on every other
machine rather than that one.

A node with no address gets no name. A name resolving to nothing is worse than
no name: connecting to an address that does not answer hangs, where a name that
does not resolve fails at once and says which name it was.

Found while writing it: a test asserting every file in the declaration is mode
0600 would have forced /etc/hosts to 0600 and broken every lookup on the
machine, to protect a file that is not secret.

Verified in the lab: three machines, nine name lookups, each resolving to the
right overlay address and reaching it.
2026-08-29 19:58:01 +02:00
jschoubben 8b974deb42 A working private network, and four reasons it did not work
Three machines across two sites, two of them behind no reachable address, all
nine paths open. The mesh computes the graph, delivers it as a declaration, and
the nodes bring it up.

Every fault below looked like success from inside the mesh: the graph was
right, the files were right, the services were up, every node reported it had
applied. None was reachable by reasoning.

A running interface does not re-read its configuration. A node joins, every
existing node's peer list changes, the file is replaced -- and the service is
already running, so nothing reloads it. Fixed as declared state rather than a
command: the service must reflect the file. A command to restart would be an
action, and the link may not carry one. The host refused exactly that, which is
how this shape was arrived at.

A hub sharing a site with a spoke appeared twice in that spoke's peer list --
once as a direct peer, once as the route of last resort. WireGuard takes one
entry per key and refuses the file. The ordinary shape of a small mesh, and in
none of the tests written before it ran.

Two nodes at one site that neither can be dialled were peered directly. Nobody
opens the path, and the direct route is more specific than the hub's, so it
wins and blackholes -- this design's own warning arriving in its
implementation. They now route through the hub unless one end can be dialled.

And Docker sets the FORWARD policy to DROP, so a hub with ip_forward enabled
carried nothing between its spokes. The substrate at tier 1 silently breaks the
network at tier 2, and nothing in either tier's state says so. The hub inserts
its own rule above those chains and removes it on the way down.

Two weak tests found by injection along the way: one asserted the keepalive
rule only against the hub, whose peer entries happen not to set that field at
all, so it tested an absence; the other checked the firewall rules by looking
for FORWARD anywhere, which the PostDown line satisfies on its own.
2026-08-29 18:04:15 +02:00
jschoubben f44e73d286 The mesh computes a private network it cannot impersonate
The first thing the control plane decides rather than relays. Every node's peer
list is derived from every node at once, which is what makes this control-plane
work by definition: no node has that view.

A hub, with direct peering between nodes at the same site. Not a full mesh, and
the reason is a property of WireGuard rather than a preference -- there is no
failover, so a more specific route to a dead endpoint blackholes instead of
falling back. A node gets exactly one path to any peer, because two would mean
one of them silently swallowing traffic. A roaming node is hub-only for the
same reason.

Reachability and the hub are declared, never inferred from an address. The
address is evidence and is not the fact: carrier-grade NAT looks public and is
not, a routable address behind a closed firewall looks public and is not, and
the regular expression that used to decide it got the lab wrong too. Hub
election by address prefix failed silently when nobody knew the convention.

No private key travels, and that is the whole design. The node generated its
own keypair and kept the private half; the configuration points at a file the
node wrote, using WireGuard's own PostUp. So the control plane composes a
complete configuration for a node it cannot pretend to be -- it knows every
public key and holds none of the private ones.

Delivered as an ordinary declaration: a package, a file and a service. The host
does not know what a private network is and does not learn one. There is a test
holding that line, because the moment connectivity needs a new shape in tier 0
is the moment the host stops being small enough to trust.

The generated file is written to be read: each peer says why it is there, a
peer with no endpoint says why it has none, and the header says not to edit it
-- an edit survives until the graph next changes and then vanishes, which is
worse than never being applied, because the machine works and then stops and
nothing changed that anybody remembers.

Fault injection found one weak test. The keepalive rule was asserted only
against the hub, whose peer entries happen not to set the field at all, so it
was testing an absence rather than the rule. It now checks two direct peers
where one is reachable and one is not.
2026-08-29 16:58:56 +02:00
jschoubben f563ababa1 The mesh keeps a copy of what each node owns
novox/hq 09-the-node-lifecycle asks for this and it was missing: the host
reports what it owns and the mesh keeps the last report. A backup, never a
source -- nothing decides anything from it, and a node that disagrees with it
wins, because the node is the one that can see the machine.

Its point is the orphans. A node that loses its state file currently strands
whatever it applied: nothing on the machine knows those resources were the
mesh's doing, so nothing removes them. With this, a rebuilt node receives both
the declaration and the record of what it previously owned.

Never reported and reported nothing are kept apart, and that is the whole care
in it. A node that applied nothing holds nothing; a node that has never spoken
is unknown -- and handing back an empty list for the second would tell a
rebuilding node it owns nothing and have it remove whatever it found.

The age comes back with the answer rather than being left for the caller to go
and find. An answer about a machine is worth much less without one, and this
repository has already been bitten by a cache with no age on it.

A refusal or a partial failure moves last_seen and nothing else: neither is an
account of what the machine holds, and recording one as though it were would
tell a rebuilding node to remove what it still has.
2026-08-29 16:51:54 +02:00