Files
hq/02-DECISIONS/0002-nodes-communicate-over-a-broker.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

69 lines
3.1 KiB
Markdown

---
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 2. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
> deliberation that happened.
## Context
The mesh is a set of machines that must call each other's capabilities. On the day the
repository was founded there was no inter-node transport at all — each node was configured
independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
## Considered options
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
participant. It also makes node liveness a caller's problem: a request to a sleeping node
is an error rather than a wait.
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
independent of demand, and request/reply has to be built on top of it by hand.
3. **A central message broker with per-node exchanges.** Chosen.
## Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange
named for itself and a request queue; a shared mesh exchange carries commands and events that
are not addressed to one node.
Three message shapes, and only three:
- **RPC** — request/reply, for calling a capability that lives on another node.
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
- **Events** — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
## Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker
connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually
right and occasionally the wrong thing entirely — a queued command for a node that never
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
where the capability is.
## References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.