Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
5 changed files with 170 additions and 0 deletions
Showing only changes of commit 918dc04916 - Show all commits
@@ -70,6 +70,37 @@ invariants were found violated simultaneously (see Consequences).
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### What this is, plainly — and what "mesh" does not mean
*Written 2026-08-29, from working through connectivity and asking whether the word still fits.*
Four layers. Naming them honestly is worth more than the word on the tin:
| | |
|---|---|
| **machines are linked by a private network** | and every machine reaches every other over it |
| **one node holds knowledge of all of them** | the control plane, and only it |
| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) |
| **agents are hired onto nodes and do the work** | the layer the other three exist to carry |
**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what
machines can reach, not how they are governed:
| | a mesh? |
|---|---|
| what a machine can reach | **yes** — genuinely any to any |
| how the traffic travels | no — anything crossing sites transits the hub |
| who decides | no. One node, declared |
**And *master* overstates it in the other direction.** A master implies the others need it in
order to function. They do not: every node holds what it was last told and runs from that copy
*always* — not as a fallback, as the only mode it has. So the control plane being gone is every
node in the ordinary disconnected situation at once, and **what is lost is change, not
operation.**
The accurate phrase is **one authority, no failover**, and both halves are deliberate
([ADR 0006](0006-the-substrate-and-the-control-plane.md)).
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
@@ -36,6 +36,21 @@ than by being ours.
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
not need to know a node exists. *Being ours does not make something infrastructure.*
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
stated nowhere.* SSH appears three times across this design and every time as something that
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
thing that knows which humans and agents exist and which nodes they may reach, which is
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
It is three questions wearing one name, and only two of them are the mesh's:
| | |
|---|---|
| **humans** | their key, on the nodes they are allowed on |
| **agents** | the same, with a lifetime — and revocation that has to actually work |
| **node to node** | **not a mesh function.** The host has no inbound control surface by decision ([ADR 0004](0004-a-node-and-how-it-joins.md)); nodes SSHing to each other would be a second control path arriving through the back door |
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
removed. Listing it as an eighth context would settle by naming what has not been settled by
@@ -46,6 +61,36 @@ arguing.
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
built, so none of it can be subtly wrong.
#### The option that would make it a real mesh, and why not
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
to reject anything.*
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
- every node holds the **whole inventory**, so there is a replication process between them;
- replication needs a writer, so one node is elected **master**, and something promotes a new one
when it drops — Redis Sentinel and its whole family of problems;
- and it still would not deliver what the name promises, because **application databases are not
replicated.** A workload's store lives where the workload lives.
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
to become **a replicated database system for everything running on it** — not for our own
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
it would be supporting.
**So there are three central roles, not one**, and it is worth seeing them separately because
only the third costs operation:
| | its loss costs |
|---|---|
| **the control plane** | nothing can be *changed*. Nothing stops running |
| **the broker** | nothing can be told anything, or report anything |
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
**Whether these are one node is not decided here.** All three must be dialable by every node, which
pushes toward one; nothing says they must be.
**That is sound rather than merely cheap**, because the design already tolerates its absence by
construction: a node reconciles from its own store and never needed to ask anybody to hold the
state it was last given. **The control plane being down is not a new failure mode — it is every
+67
View File
@@ -84,6 +84,45 @@ with a declared one, the disagreement is a **reportable condition**, not a silen
home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a
pattern match.
### Some nodes must be reachable, and this had not been said
*Written 2026-08-29, on being asked and finding no answer.*
Everything above treats reachability as a **fact to record** — which node can be dialled, so that
exposure and certificate issuance can be placed. It never said the converse, and the converse is a
hard requirement:
| role | dialled by | so it needs |
|---|---|---|
| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** |
| **the hub** | every node not co-located with its peer | the same |
**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh
confined to one network it does not — the requirement is about the nodes that exist, not about the
public internet.
**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there
is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose
address moves invalidates every token issued for it, and a node that was disconnected across the
change cannot get back.
**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it
belongs with the others rather than being discovered.
### The link stays on the underlay, and that is a repair channel
The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH,
services, node to node — and the exception is each node's own outbound link to the broker.
**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one.
**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A
repair channel carried over the thing being repaired is not a repair channel: a node whose only
path home was the overlay is gone the moment an overlay declaration is wrong.
**What it does not cost is confidentiality.** The link is already authenticated and encrypted
against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the
overlay would not protect traffic that is unprotected today.
## A filter rule names its source
`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no
@@ -122,3 +161,31 @@ forbidden is a listening thing that accepts instructions and changes the machine
mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located
peers and every assigned workload keep running.
## Open — the link over the overlay, with a fallback
*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks
and it is a decision rather than a derivation.*
The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay
is not working — so ordinary operation is private and the underlay stays as the way back.
**Two things it would have to get right:**
- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once
configured it is up whether or not the far end exists. So there is no flag to read, and failing
over means *try, fail, time out, retry elsewhere*.
- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a
node whose overlay is broken with nothing to say so, and it will stay broken because everything
still works. That is this repository's recurring fault — a failure that reads as success — and a
fallback is the easiest place in the design to reintroduce it.
**What stops it being an obvious win:** if the underlay path must stay available for the fallback,
the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted
— the gain is which network carries bytes, not what an attacker can reach.
**The version that would buy something is overlay-only**, with the broker firewalled to the overlay
in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real
trade: it exchanges the automatic way back in for a closed port.
Not decided either way here.
+16
View File
@@ -29,6 +29,22 @@ That is the same shape the host uses on a machine, one layer up:
**This is not the current coordinator repaired.** That is a state machine over stages; the value
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
### The module system is the CI/CD
*Written 2026-08-29, because this was the intention throughout and was never stated in one line.*
**There is no pipeline product beside the mesh, and there is not going to be one.** A module
declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its
source is ahead of its artifacts and builds it; the graph says what else that invalidates; the
node that should run it is told. Build, test, publish and deploy are the same reconciliation seen
at four points, not four stages wired together.
**Which is why the module system is the core of the setup rather than one component of it.** Every
other layer is carried by it: the substrate is modules the bundle raises before there is a mesh,
the control plane is a module, and an application is a module with a different manifest. A thing
that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it
is the one keeping a second delivery mechanism from growing beside this one.
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
from, which is how a verify stage was built and never scheduled, and no run to lose.
+11
View File
@@ -81,6 +81,17 @@ outbound-only and carries its own identity, so it needs nothing the overlay prov
patched with an `/etc/hosts` floor written underneath the resolver; under this design there is
nothing to patch.
**Step 1 has a precondition this document treated as a fact to record rather than a requirement:
the broker's node must be dialable by every node, at a stable address, and so must the hub**
([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Across the internet that means publicly
reachable; on one network it does not. A mesh whose nodes are all behind NAT cannot be raised, and
a broker node whose address moves invalidates every token issued for it.
**Whether the link should later move onto the overlay, with the underlay as fallback, is
[open](../../02-DECISIONS/0007-connectivity.md).** It is a decision rather than a derivation: the
gain is which network carries bytes, not what an attacker can reach, since the link is already
encrypted against a pinned fingerprint.
## 1 — The overlay
**What is decided:** the peer graph. For every node: its overlay address, which peers it holds,