diff --git a/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md b/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md index 0e437fa..1f86095 100644 --- a/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md +++ b/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md @@ -70,6 +70,37 @@ invariants were found violated simultaneously (see Consequences). **The mesh brokers capabilities. Nodes are places where work runs. Agents are personas that think and act.** Everything else supports one of those three. +### What this is, plainly — and what "mesh" does not mean + +*Written 2026-08-29, from working through connectivity and asking whether the word still fits.* + +Four layers. Naming them honestly is worth more than the word on the tin: + +| | | +|---|---| +| **machines are linked by a private network** | and every machine reaches every other over it | +| **one node holds knowledge of all of them** | the control plane, and only it | +| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) | +| **agents are hired onto nodes and do the work** | the layer the other three exist to carry | + +**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what +machines can reach, not how they are governed: + +| | a mesh? | +|---|---| +| what a machine can reach | **yes** — genuinely any to any | +| how the traffic travels | no — anything crossing sites transits the hub | +| who decides | no. One node, declared | + +**And *master* overstates it in the other direction.** A master implies the others need it in +order to function. They do not: every node holds what it was last told and runs from that copy +*always* — not as a fallback, as the only mode it has. So the control plane being gone is every +node in the ordinary disconnected situation at once, and **what is lost is change, not +operation.** + +The accurate phrase is **one authority, no failover**, and both halves are deliberate +([ADR 0006](0006-the-substrate-and-the-control-plane.md)). + ### Nodes and agents are decoupled A node is a place where an agent can run — that is the entire relationship. There is no diff --git a/02-DECISIONS/0006-the-substrate-and-the-control-plane.md b/02-DECISIONS/0006-the-substrate-and-the-control-plane.md index ff33ec1..b24a5ba 100644 --- a/02-DECISIONS/0006-the-substrate-and-the-control-plane.md +++ b/02-DECISIONS/0006-the-substrate-and-the-control-plane.md @@ -36,6 +36,21 @@ than by being ours. **`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does not need to know a node exists. *Being ours does not make something infrastructure.* +**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and +stated nowhere.* SSH appears three times across this design and every time as something that +*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary +traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only +thing that knows which humans and agents exist and which nodes they may reach, which is +`identity`'s definition. The node end already works, since an `authorized_keys` file is a file. + +It is three questions wearing one name, and only two of them are the mesh's: + +| | | +|---|---| +| **humans** | their key, on the nodes they are allowed on | +| **agents** | the same, with a lifetime — and revocation that has to actually work | +| **node to node** | **not a mesh function.** The host has no inbound control surface by decision ([ADR 0004](0004-a-node-and-how-it-joins.md)); nodes SSHing to each other would be a second control path arriving through the back door | + **Where the record lives is deliberately open.** Contexts integrate through it, which makes it load-bearing, and putting it in the substrate risks recreating the circularity the tiers just removed. Listing it as an eighth context would settle by naming what has not been settled by @@ -46,6 +61,36 @@ arguing. **Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it built, so none of it can be subtly wrong. +#### The option that would make it a real mesh, and why not + +*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way +to reject anything.* + +A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price: + +- every node holds the **whole inventory**, so there is a replication process between them; +- replication needs a writer, so one node is elected **master**, and something promotes a new one + when it drops — Redis Sentinel and its whole family of problems; +- and it still would not deliver what the name promises, because **application databases are not + replicated.** A workload's store lives where the workload lives. + +That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have +to become **a replicated database system for everything running on it** — not for our own +inventory, for every consumer's data too. That is a product, and a much larger one than the thing +it would be supporting. + +**So there are three central roles, not one**, and it is worth seeing them separately because +only the third costs operation: + +| | its loss costs | +|---|---| +| **the control plane** | nothing can be *changed*. Nothing stops running | +| **the broker** | nothing can be told anything, or report anything | +| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) | + +**Whether these are one node is not decided here.** All three must be dialable by every node, which +pushes toward one; nothing says they must be. + **That is sound rather than merely cheap**, because the design already tolerates its absence by construction: a node reconciles from its own store and never needed to ask anybody to hold the state it was last given. **The control plane being down is not a new failure mode — it is every diff --git a/02-DECISIONS/0007-connectivity.md b/02-DECISIONS/0007-connectivity.md index 086d889..f70d784 100644 --- a/02-DECISIONS/0007-connectivity.md +++ b/02-DECISIONS/0007-connectivity.md @@ -84,6 +84,45 @@ with a declared one, the disagreement is a **reportable condition**, not a silen home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a pattern match. +### Some nodes must be reachable, and this had not been said + +*Written 2026-08-29, on being asked and finding no answer.* + +Everything above treats reachability as a **fact to record** — which node can be dialled, so that +exposure and certificate issuance can be placed. It never said the converse, and the converse is a +hard requirement: + +| role | dialled by | so it needs | +|---|---|---| +| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** | +| **the hub** | every node not co-located with its peer | the same | + +**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh +confined to one network it does not — the requirement is about the nodes that exist, not about the +public internet. + +**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there +is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose +address moves invalidates every token issued for it, and a node that was disconnected across the +change cannot get back. + +**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it +belongs with the others rather than being discovered. + +### The link stays on the underlay, and that is a repair channel + +The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH, +services, node to node — and the exception is each node's own outbound link to the broker. + +**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one. +**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A +repair channel carried over the thing being repaired is not a repair channel: a node whose only +path home was the overlay is gone the moment an overlay declaration is wrong. + +**What it does not cost is confidentiality.** The link is already authenticated and encrypted +against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the +overlay would not protect traffic that is unprotected today. + ## A filter rule names its source `scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no @@ -122,3 +161,31 @@ forbidden is a listening thing that accepts instructions and changes the machine mesh's own for internal ones. A single-CA lab would hide any bug living in the split. - **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located peers and every assigned workload keep running. + +## Open — the link over the overlay, with a fallback + +*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks +and it is a decision rather than a derivation.* + +The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay +is not working — so ordinary operation is private and the underlay stays as the way back. + +**Two things it would have to get right:** + +- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once + configured it is up whether or not the far end exists. So there is no flag to read, and failing + over means *try, fail, time out, retry elsewhere*. +- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a + node whose overlay is broken with nothing to say so, and it will stay broken because everything + still works. That is this repository's recurring fault — a failure that reads as success — and a + fallback is the easiest place in the design to reintroduce it. + +**What stops it being an obvious win:** if the underlay path must stay available for the fallback, +the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted +— the gain is which network carries bytes, not what an attacker can reach. + +**The version that would buy something is overlay-only**, with the broker firewalled to the overlay +in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real +trade: it exchanges the automatic way back in for a closed port. + +Not decided either way here. diff --git a/02-DECISIONS/0010-delivery.md b/02-DECISIONS/0010-delivery.md index 74f1b73..adf3f97 100644 --- a/02-DECISIONS/0010-delivery.md +++ b/02-DECISIONS/0010-delivery.md @@ -29,6 +29,22 @@ That is the same shape the host uses on a machine, one layer up: **This is not the current coordinator repaired.** That is a state machine over stages; the value of it here is as a catalogue of the ways this fails, and it has been used for exactly that. +### The module system is the CI/CD + +*Written 2026-08-29, because this was the intention throughout and was never stated in one line.* + +**There is no pipeline product beside the mesh, and there is not going to be one.** A module +declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its +source is ahead of its artifacts and builds it; the graph says what else that invalidates; the +node that should run it is told. Build, test, publish and deploy are the same reconciliation seen +at four points, not four stages wired together. + +**Which is why the module system is the core of the setup rather than one component of it.** Every +other layer is carried by it: the substrate is modules the bundle raises before there is a mesh, +the control plane is a module, and an application is a module with a different manifest. A thing +that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it +is the one keeping a second delivery mechanism from growing beside this one. + **What disappears is the pipeline as a state machine** — no stage list something can be omitted from, which is how a verify stage was built and never scheduled, and no run to lose. diff --git a/03-DESIGN/01-to-be/08-connectivity.md b/03-DESIGN/01-to-be/08-connectivity.md index 2258770..385cd44 100644 --- a/03-DESIGN/01-to-be/08-connectivity.md +++ b/03-DESIGN/01-to-be/08-connectivity.md @@ -81,6 +81,17 @@ outbound-only and carries its own identity, so it needs nothing the overlay prov patched with an `/etc/hosts` floor written underneath the resolver; under this design there is nothing to patch. +**Step 1 has a precondition this document treated as a fact to record rather than a requirement: +the broker's node must be dialable by every node, at a stable address, and so must the hub** +([ADR 0007](../../02-DECISIONS/0007-connectivity.md)). Across the internet that means publicly +reachable; on one network it does not. A mesh whose nodes are all behind NAT cannot be raised, and +a broker node whose address moves invalidates every token issued for it. + +**Whether the link should later move onto the overlay, with the underlay as fallback, is +[open](../../02-DECISIONS/0007-connectivity.md).** It is a decision rather than a derivation: the +gain is which network carries bytes, not what an attacker can reach, since the link is already +encrypted against a pinned fingerprint. + ## 1 — The overlay **What is decided:** the peer graph. For every node: its overlay address, which peers it holds,