What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in somebody's head and written nowhere. It is not a mesh in the peer-to-peer sense and will not become one. 0001 now says what it is instead: machines linked by a private network, one node holding knowledge of all of them, modules as the way anything is built and delivered, and agents hired onto nodes to do the work. The word describes what machines can reach, not how they are governed. "Master" overstates it the other way -- nothing needs that node to keep running, only to change. 0006 gains the option that would make it a real mesh, recorded as considered rather than rejected by silence: every node holding the whole inventory, a replication process, an elected master with promotion on failure. What settles it is not the complexity but that it still would not deliver the name, because application databases are not replicated -- so a genuine peer-to-peer mesh means becoming a replicated database system for every consumer's data too. That is a larger product than the thing it would support. Also in 0006: three central roles, not one. Losing the control plane costs change, losing the broker costs being told anything, and losing the hub costs nodes in different places reaching each other at all -- which is operation, not administration. Whether they are one node is not decided. And SSH access is identity's. It appeared three times as something that uses the overlay and never as something the mesh provides, which reads as settled when nothing decided it. Nobody else could: the mesh is the only thing that knows which humans and agents exist and which nodes they may reach. Node to node SSH stays out -- the host has no inbound control surface by decision, and nodes reaching each other that way is a second control path through the back door. 0007 gains the requirement underneath all of it. Reachability was recorded as a fact to track and never as a thing some node must have. The broker's node and the hub must be dialable by every node at a stable address, or nothing can join and a disconnected node cannot return. A mesh entirely behind NAT cannot be raised. That is a precondition and it belongs with the others. The link staying on the underlay is also argued now rather than asserted. At join time it is forced; afterwards it is a choice, and the reason is that a repair channel carried over the thing being repaired is not one. Moving it onto the overlay, with fallback, is recorded as open with what it would have to get right -- a WireGuard interface has no link state to test, and a silent fallback is this repository's recurring fault in a new place. 0010 says in one line what was the intention throughout: the module system is the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are one reconciliation seen at four points, which is why a thing that cannot be a module cannot be delivered.
This commit is contained in:
@@ -36,6 +36,21 @@ than by being ours.
|
||||
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
|
||||
not need to know a node exists. *Being ours does not make something infrastructure.*
|
||||
|
||||
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
|
||||
stated nowhere.* SSH appears three times across this design and every time as something that
|
||||
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
|
||||
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
|
||||
thing that knows which humans and agents exist and which nodes they may reach, which is
|
||||
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
|
||||
|
||||
It is three questions wearing one name, and only two of them are the mesh's:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **humans** | their key, on the nodes they are allowed on |
|
||||
| **agents** | the same, with a lifetime — and revocation that has to actually work |
|
||||
| **node to node** | **not a mesh function.** The host has no inbound control surface by decision ([ADR 0004](0004-a-node-and-how-it-joins.md)); nodes SSHing to each other would be a second control path arriving through the back door |
|
||||
|
||||
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
|
||||
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
|
||||
removed. Listing it as an eighth context would settle by naming what has not been settled by
|
||||
@@ -46,6 +61,36 @@ arguing.
|
||||
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
|
||||
built, so none of it can be subtly wrong.
|
||||
|
||||
#### The option that would make it a real mesh, and why not
|
||||
|
||||
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
|
||||
to reject anything.*
|
||||
|
||||
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
|
||||
|
||||
- every node holds the **whole inventory**, so there is a replication process between them;
|
||||
- replication needs a writer, so one node is elected **master**, and something promotes a new one
|
||||
when it drops — Redis Sentinel and its whole family of problems;
|
||||
- and it still would not deliver what the name promises, because **application databases are not
|
||||
replicated.** A workload's store lives where the workload lives.
|
||||
|
||||
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
|
||||
to become **a replicated database system for everything running on it** — not for our own
|
||||
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
|
||||
it would be supporting.
|
||||
|
||||
**So there are three central roles, not one**, and it is worth seeing them separately because
|
||||
only the third costs operation:
|
||||
|
||||
| | its loss costs |
|
||||
|---|---|
|
||||
| **the control plane** | nothing can be *changed*. Nothing stops running |
|
||||
| **the broker** | nothing can be told anything, or report anything |
|
||||
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
|
||||
|
||||
**Whether these are one node is not decided here.** All three must be dialable by every node, which
|
||||
pushes toward one; nothing says they must be.
|
||||
|
||||
**That is sound rather than merely cheap**, because the design already tolerates its absence by
|
||||
construction: a node reconciles from its own store and never needed to ask anybody to hold the
|
||||
state it was last given. **The control plane being down is not a new failure mode — it is every
|
||||
|
||||
Reference in New Issue
Block a user