Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception.
69 lines
3.1 KiB
Markdown
69 lines
3.1 KiB
Markdown
---
|
|
status: accepted
|
|
date: 2026-02-25
|
|
deciders: jochen
|
|
reconstructed: true
|
|
---
|
|
|
|
# 2. Nodes communicate over a message broker, not over HTTP
|
|
|
|
> Reconstructed after the fact from the evidence cited below. The decision was taken in
|
|
> implementation, not in a record; this document states what was decided and why, not a
|
|
> deliberation that happened.
|
|
|
|
## Context
|
|
|
|
The mesh is a set of machines that must call each other's capabilities. On the day the
|
|
repository was founded there was no inter-node transport at all — each node was configured
|
|
independently and shared nothing at runtime.
|
|
|
|
Three properties were required and are visible in everything built since:
|
|
|
|
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
|
|
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
|
|
should wait, not error.
|
|
- Adding a node must not require editing anything on the nodes that already exist.
|
|
|
|
## Considered options
|
|
|
|
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
|
|
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
|
|
participant. It also makes node liveness a caller's problem: a request to a sleeping node
|
|
is an error rather than a wait.
|
|
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
|
|
independent of demand, and request/reply has to be built on top of it by hand.
|
|
3. **A central message broker with per-node exchanges.** Chosen.
|
|
|
|
## Decision
|
|
|
|
All inter-node communication goes through a message broker. Every node owns a topic exchange
|
|
named for itself and a request queue; a shared mesh exchange carries commands and events that
|
|
are not addressed to one node.
|
|
|
|
Three message shapes, and only three:
|
|
|
|
- **RPC** — request/reply, for calling a capability that lives on another node.
|
|
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
|
|
- **Events** — statements that something happened, addressed to nobody.
|
|
|
|
Every node dials the broker outbound. Nothing dials a node.
|
|
|
|
## Consequences
|
|
|
|
- NAT stops being an architectural concern. A node's reachability is a property of the broker
|
|
connection, not of its network position.
|
|
- A call to a node that is down waits in that node's queue instead of failing. This is usually
|
|
right and occasionally the wrong thing entirely — a queued command for a node that never
|
|
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
|
|
- The broker is a single point of failure and a single point of trust. Its credential is
|
|
mesh-wide, so rotating it is a mesh-wide operation.
|
|
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
|
|
where the capability is.
|
|
|
|
## References
|
|
|
|
- The broker was stood up on 2026-02-25, the second day of the repository.
|
|
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
|
|
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
|
|
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.
|