Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception.
3.1 KiB
status, date, deciders, reconstructed
| status | date | deciders | reconstructed |
|---|---|---|---|
| accepted | 2026-02-25 | jochen | true |
2. Nodes communicate over a message broker, not over HTTP
Reconstructed after the fact from the evidence cited below. The decision was taken in implementation, not in a record; this document states what was decided and why, not a deliberation that happened.
Context
The mesh is a set of machines that must call each other's capabilities. On the day the repository was founded there was no inter-node transport at all — each node was configured independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
Considered options
- HTTP APIs between nodes. Rejected. Every node becomes a server that every other node must be able to reach, which the NAT case makes impossible without inbound tunnels to each participant. It also makes node liveness a caller's problem: a request to a sleeping node is an error rather than a wait.
- Polling a shared database. Rejected. Latency is the poll interval, load is constant and independent of demand, and request/reply has to be built on top of it by hand.
- A central message broker with per-node exchanges. Chosen.
Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange named for itself and a request queue; a shared mesh exchange carries commands and events that are not addressed to one node.
Three message shapes, and only three:
- RPC — request/reply, for calling a capability that lives on another node.
- Commands — instructions to do a stage of work, addressed by what is to be done.
- Events — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually right and occasionally the wrong thing entirely — a queued command for a node that never returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay where the capability is.
References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base:
mesh(transport, exchanges, queue naming),troubleshooting/amqp-credential-rotation. - The stall shape is recorded in
troubleshooting/empty-pipeline-blocks-the-queueandtroubleshooting/daemon-and-tool-server-share-a-request-queue.