Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -70,6 +70,37 @@ invariants were found violated simultaneously (see Consequences).
|
||||
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
|
||||
that think and act.** Everything else supports one of those three.
|
||||
|
||||
### What this is, plainly — and what "mesh" does not mean
|
||||
|
||||
*Written 2026-08-29, from working through connectivity and asking whether the word still fits.*
|
||||
|
||||
Four layers. Naming them honestly is worth more than the word on the tin:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **machines are linked by a private network** | and every machine reaches every other over it |
|
||||
| **one node holds knowledge of all of them** | the control plane, and only it |
|
||||
| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) |
|
||||
| **agents are hired onto nodes and do the work** | the layer the other three exist to carry |
|
||||
|
||||
**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what
|
||||
machines can reach, not how they are governed:
|
||||
|
||||
| | a mesh? |
|
||||
|---|---|
|
||||
| what a machine can reach | **yes** — genuinely any to any |
|
||||
| how the traffic travels | no — anything crossing sites transits the hub |
|
||||
| who decides | no. One node, declared |
|
||||
|
||||
**And *master* overstates it in the other direction.** A master implies the others need it in
|
||||
order to function. They do not: every node holds what it was last told and runs from that copy
|
||||
*always* — not as a fallback, as the only mode it has. So the control plane being gone is every
|
||||
node in the ordinary disconnected situation at once, and **what is lost is change, not
|
||||
operation.**
|
||||
|
||||
The accurate phrase is **one authority, no failover**, and both halves are deliberate
|
||||
([ADR 0006](0006-the-substrate-and-the-control-plane.md)).
|
||||
|
||||
### Nodes and agents are decoupled
|
||||
|
||||
A node is a place where an agent can run — that is the entire relationship. There is no
|
||||
|
||||
@@ -36,6 +36,21 @@ than by being ours.
|
||||
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
|
||||
not need to know a node exists. *Being ours does not make something infrastructure.*
|
||||
|
||||
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
|
||||
stated nowhere.* SSH appears three times across this design and every time as something that
|
||||
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
|
||||
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
|
||||
thing that knows which humans and agents exist and which nodes they may reach, which is
|
||||
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
|
||||
|
||||
It is three questions wearing one name, and only two of them are the mesh's:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **humans** | their key, on the nodes they are allowed on |
|
||||
| **agents** | the same, with a lifetime — and revocation that has to actually work |
|
||||
| **node to node** | **not a mesh function.** The host has no inbound control surface by decision ([ADR 0004](0004-a-node-and-how-it-joins.md)); nodes SSHing to each other would be a second control path arriving through the back door |
|
||||
|
||||
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
|
||||
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
|
||||
removed. Listing it as an eighth context would settle by naming what has not been settled by
|
||||
@@ -46,6 +61,36 @@ arguing.
|
||||
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
|
||||
built, so none of it can be subtly wrong.
|
||||
|
||||
#### The option that would make it a real mesh, and why not
|
||||
|
||||
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
|
||||
to reject anything.*
|
||||
|
||||
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
|
||||
|
||||
- every node holds the **whole inventory**, so there is a replication process between them;
|
||||
- replication needs a writer, so one node is elected **master**, and something promotes a new one
|
||||
when it drops — Redis Sentinel and its whole family of problems;
|
||||
- and it still would not deliver what the name promises, because **application databases are not
|
||||
replicated.** A workload's store lives where the workload lives.
|
||||
|
||||
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
|
||||
to become **a replicated database system for everything running on it** — not for our own
|
||||
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
|
||||
it would be supporting.
|
||||
|
||||
**So there are three central roles, not one**, and it is worth seeing them separately because
|
||||
only the third costs operation:
|
||||
|
||||
| | its loss costs |
|
||||
|---|---|
|
||||
| **the control plane** | nothing can be *changed*. Nothing stops running |
|
||||
| **the broker** | nothing can be told anything, or report anything |
|
||||
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
|
||||
|
||||
**Whether these are one node is not decided here.** All three must be dialable by every node, which
|
||||
pushes toward one; nothing says they must be.
|
||||
|
||||
**That is sound rather than merely cheap**, because the design already tolerates its absence by
|
||||
construction: a node reconciles from its own store and never needed to ask anybody to hold the
|
||||
state it was last given. **The control plane being down is not a new failure mode — it is every
|
||||
|
||||
@@ -84,6 +84,45 @@ with a declared one, the disagreement is a **reportable condition**, not a silen
|
||||
home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a
|
||||
pattern match.
|
||||
|
||||
### Some nodes must be reachable, and this had not been said
|
||||
|
||||
*Written 2026-08-29, on being asked and finding no answer.*
|
||||
|
||||
Everything above treats reachability as a **fact to record** — which node can be dialled, so that
|
||||
exposure and certificate issuance can be placed. It never said the converse, and the converse is a
|
||||
hard requirement:
|
||||
|
||||
| role | dialled by | so it needs |
|
||||
|---|---|---|
|
||||
| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** |
|
||||
| **the hub** | every node not co-located with its peer | the same |
|
||||
|
||||
**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh
|
||||
confined to one network it does not — the requirement is about the nodes that exist, not about the
|
||||
public internet.
|
||||
|
||||
**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there
|
||||
is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose
|
||||
address moves invalidates every token issued for it, and a node that was disconnected across the
|
||||
change cannot get back.
|
||||
|
||||
**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it
|
||||
belongs with the others rather than being discovered.
|
||||
|
||||
### The link stays on the underlay, and that is a repair channel
|
||||
|
||||
The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH,
|
||||
services, node to node — and the exception is each node's own outbound link to the broker.
|
||||
|
||||
**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one.
|
||||
**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A
|
||||
repair channel carried over the thing being repaired is not a repair channel: a node whose only
|
||||
path home was the overlay is gone the moment an overlay declaration is wrong.
|
||||
|
||||
**What it does not cost is confidentiality.** The link is already authenticated and encrypted
|
||||
against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the
|
||||
overlay would not protect traffic that is unprotected today.
|
||||
|
||||
## A filter rule names its source
|
||||
|
||||
`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no
|
||||
@@ -122,3 +161,31 @@ forbidden is a listening thing that accepts instructions and changes the machine
|
||||
mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
|
||||
- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located
|
||||
peers and every assigned workload keep running.
|
||||
|
||||
## Open — the link over the overlay, with a fallback
|
||||
|
||||
*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks
|
||||
and it is a decision rather than a derivation.*
|
||||
|
||||
The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay
|
||||
is not working — so ordinary operation is private and the underlay stays as the way back.
|
||||
|
||||
**Two things it would have to get right:**
|
||||
|
||||
- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once
|
||||
configured it is up whether or not the far end exists. So there is no flag to read, and failing
|
||||
over means *try, fail, time out, retry elsewhere*.
|
||||
- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a
|
||||
node whose overlay is broken with nothing to say so, and it will stay broken because everything
|
||||
still works. That is this repository's recurring fault — a failure that reads as success — and a
|
||||
fallback is the easiest place in the design to reintroduce it.
|
||||
|
||||
**What stops it being an obvious win:** if the underlay path must stay available for the fallback,
|
||||
the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted
|
||||
— the gain is which network carries bytes, not what an attacker can reach.
|
||||
|
||||
**The version that would buy something is overlay-only**, with the broker firewalled to the overlay
|
||||
in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real
|
||||
trade: it exchanges the automatic way back in for a closed port.
|
||||
|
||||
Not decided either way here.
|
||||
|
||||
@@ -29,6 +29,22 @@ That is the same shape the host uses on a machine, one layer up:
|
||||
**This is not the current coordinator repaired.** That is a state machine over stages; the value
|
||||
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
|
||||
|
||||
### The module system is the CI/CD
|
||||
|
||||
*Written 2026-08-29, because this was the intention throughout and was never stated in one line.*
|
||||
|
||||
**There is no pipeline product beside the mesh, and there is not going to be one.** A module
|
||||
declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its
|
||||
source is ahead of its artifacts and builds it; the graph says what else that invalidates; the
|
||||
node that should run it is told. Build, test, publish and deploy are the same reconciliation seen
|
||||
at four points, not four stages wired together.
|
||||
|
||||
**Which is why the module system is the core of the setup rather than one component of it.** Every
|
||||
other layer is carried by it: the substrate is modules the bundle raises before there is a mesh,
|
||||
the control plane is a module, and an application is a module with a different manifest. A thing
|
||||
that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it
|
||||
is the one keeping a second delivery mechanism from growing beside this one.
|
||||
|
||||
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
|
||||
from, which is how a verify stage was built and never scheduled, and no run to lose.
|
||||
|
||||
|
||||
@@ -81,6 +81,17 @@ outbound-only and carries its own identity, so it needs nothing the overlay prov
|
||||
patched with an `/etc/hosts` floor written underneath the resolver; under this design there is
|
||||
nothing to patch.
|
||||
|
||||
**Step 1 has a precondition this document treated as a fact to record rather than a requirement:
|
||||
the broker's node must be dialable by every node, at a stable address, and so must the hub**
|
||||
([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Across the internet that means publicly
|
||||
reachable; on one network it does not. A mesh whose nodes are all behind NAT cannot be raised, and
|
||||
a broker node whose address moves invalidates every token issued for it.
|
||||
|
||||
**Whether the link should later move onto the overlay, with the underlay as fallback, is
|
||||
[open](../../02-DECISIONS/0007-connectivity.md).** It is a decision rather than a derivation: the
|
||||
gain is which network carries bytes, not what an attacker can reach, since the link is already
|
||||
encrypted against a pinned fingerprint.
|
||||
|
||||
## 1 — The overlay
|
||||
|
||||
**What is decided:** the peer graph. For every node: its overlay address, which peers it holds,
|
||||
|
||||
Reference in New Issue
Block a user