What this actually is, and three things that were assumed

Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
This commit is contained in:
2026-08-29 13:01:36 +02:00
parent 5218b06c02
commit 918dc04916
5 changed files with 170 additions and 0 deletions
@@ -70,6 +70,37 @@ invariants were found violated simultaneously (see Consequences).
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### What this is, plainly — and what "mesh" does not mean
*Written 2026-08-29, from working through connectivity and asking whether the word still fits.*
Four layers. Naming them honestly is worth more than the word on the tin:
| | |
|---|---|
| **machines are linked by a private network** | and every machine reaches every other over it |
| **one node holds knowledge of all of them** | the control plane, and only it |
| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) |
| **agents are hired onto nodes and do the work** | the layer the other three exist to carry |
**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what
machines can reach, not how they are governed:
| | a mesh? |
|---|---|
| what a machine can reach | **yes** — genuinely any to any |
| how the traffic travels | no — anything crossing sites transits the hub |
| who decides | no. One node, declared |
**And *master* overstates it in the other direction.** A master implies the others need it in
order to function. They do not: every node holds what it was last told and runs from that copy
*always* — not as a fallback, as the only mode it has. So the control plane being gone is every
node in the ordinary disconnected situation at once, and **what is lost is change, not
operation.**
The accurate phrase is **one authority, no failover**, and both halves are deliberate
([ADR 0006](0006-the-substrate-and-the-control-plane.md)).
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
@@ -36,6 +36,21 @@ than by being ours.
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
not need to know a node exists. *Being ours does not make something infrastructure.*
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
stated nowhere.* SSH appears three times across this design and every time as something that
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
thing that knows which humans and agents exist and which nodes they may reach, which is
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
It is three questions wearing one name, and only two of them are the mesh's:
| | |
|---|---|
| **humans** | their key, on the nodes they are allowed on |
| **agents** | the same, with a lifetime — and revocation that has to actually work |
| **node to node** | **not a mesh function.** The host has no inbound control surface by decision ([ADR 0004](0004-a-node-and-how-it-joins.md)); nodes SSHing to each other would be a second control path arriving through the back door |
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
removed. Listing it as an eighth context would settle by naming what has not been settled by
@@ -46,6 +61,36 @@ arguing.
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
built, so none of it can be subtly wrong.
#### The option that would make it a real mesh, and why not
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
to reject anything.*
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
- every node holds the **whole inventory**, so there is a replication process between them;
- replication needs a writer, so one node is elected **master**, and something promotes a new one
when it drops — Redis Sentinel and its whole family of problems;
- and it still would not deliver what the name promises, because **application databases are not
replicated.** A workload's store lives where the workload lives.
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
to become **a replicated database system for everything running on it** — not for our own
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
it would be supporting.
**So there are three central roles, not one**, and it is worth seeing them separately because
only the third costs operation:
| | its loss costs |
|---|---|
| **the control plane** | nothing can be *changed*. Nothing stops running |
| **the broker** | nothing can be told anything, or report anything |
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
**Whether these are one node is not decided here.** All three must be dialable by every node, which
pushes toward one; nothing says they must be.
**That is sound rather than merely cheap**, because the design already tolerates its absence by
construction: a node reconciles from its own store and never needed to ask anybody to hold the
state it was last given. **The control plane being down is not a new failure mode — it is every
+67
View File
@@ -84,6 +84,45 @@ with a declared one, the disagreement is a **reportable condition**, not a silen
home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a
pattern match.
### Some nodes must be reachable, and this had not been said
*Written 2026-08-29, on being asked and finding no answer.*
Everything above treats reachability as a **fact to record** — which node can be dialled, so that
exposure and certificate issuance can be placed. It never said the converse, and the converse is a
hard requirement:
| role | dialled by | so it needs |
|---|---|---|
| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** |
| **the hub** | every node not co-located with its peer | the same |
**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh
confined to one network it does not — the requirement is about the nodes that exist, not about the
public internet.
**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there
is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose
address moves invalidates every token issued for it, and a node that was disconnected across the
change cannot get back.
**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it
belongs with the others rather than being discovered.
### The link stays on the underlay, and that is a repair channel
The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH,
services, node to node — and the exception is each node's own outbound link to the broker.
**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one.
**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A
repair channel carried over the thing being repaired is not a repair channel: a node whose only
path home was the overlay is gone the moment an overlay declaration is wrong.
**What it does not cost is confidentiality.** The link is already authenticated and encrypted
against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the
overlay would not protect traffic that is unprotected today.
## A filter rule names its source
`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no
@@ -122,3 +161,31 @@ forbidden is a listening thing that accepts instructions and changes the machine
mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located
peers and every assigned workload keep running.
## Open — the link over the overlay, with a fallback
*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks
and it is a decision rather than a derivation.*
The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay
is not working — so ordinary operation is private and the underlay stays as the way back.
**Two things it would have to get right:**
- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once
configured it is up whether or not the far end exists. So there is no flag to read, and failing
over means *try, fail, time out, retry elsewhere*.
- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a
node whose overlay is broken with nothing to say so, and it will stay broken because everything
still works. That is this repository's recurring fault — a failure that reads as success — and a
fallback is the easiest place in the design to reintroduce it.
**What stops it being an obvious win:** if the underlay path must stay available for the fallback,
the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted
— the gain is which network carries bytes, not what an attacker can reach.
**The version that would buy something is overlay-only**, with the broker firewalled to the overlay
in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real
trade: it exchanges the automatic way back in for a closed port.
Not decided either way here.
+16
View File
@@ -29,6 +29,22 @@ That is the same shape the host uses on a machine, one layer up:
**This is not the current coordinator repaired.** That is a state machine over stages; the value
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
### The module system is the CI/CD
*Written 2026-08-29, because this was the intention throughout and was never stated in one line.*
**There is no pipeline product beside the mesh, and there is not going to be one.** A module
declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its
source is ahead of its artifacts and builds it; the graph says what else that invalidates; the
node that should run it is told. Build, test, publish and deploy are the same reconciliation seen
at four points, not four stages wired together.
**Which is why the module system is the core of the setup rather than one component of it.** Every
other layer is carried by it: the substrate is modules the bundle raises before there is a mesh,
the control plane is a module, and an application is a module with a different manifest. A thing
that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it
is the one keeping a second delivery mechanism from growing beside this one.
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
from, which is how a verify stage was built and never scheduled, and no run to lose.
+11
View File
@@ -81,6 +81,17 @@ outbound-only and carries its own identity, so it needs nothing the overlay prov
patched with an `/etc/hosts` floor written underneath the resolver; under this design there is
nothing to patch.
**Step 1 has a precondition this document treated as a fact to record rather than a requirement:
the broker's node must be dialable by every node, at a stable address, and so must the hub**
([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Across the internet that means publicly
reachable; on one network it does not. A mesh whose nodes are all behind NAT cannot be raised, and
a broker node whose address moves invalidates every token issued for it.
**Whether the link should later move onto the overlay, with the underlay as fallback, is
[open](../../02-DECISIONS/0007-connectivity.md).** It is a decision rather than a derivation: the
gain is which network carries bytes, not what an attacker can reach, since the link is already
encrypted against a pinned fingerprint.
## 1 — The overlay
**What is decided:** the peer graph. For every node: its overlay address, which peers it holds,