Files
hq/02-DECISIONS/0002-nodes-communicate-over-a-broker.md
jschoubben b4607dfc03 Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code
comments across two repositories, none of which would have failed to compile.
They would have pointed at the wrong reasoning, which is worse than a broken
link because nothing reports it.

So a number identifies a record and never changes. It cannot also be a
position -- a position moves when the set changes, and an identity that moves
is not one.

The reading order moves into an index generated from each record's `topic:`.
Six topics, in the order somebody learns the system.

The index is WRITTEN rather than only generated on demand, which reverses what
this repository previously said. The reason it said otherwise is that a
hand-written index drifts -- but a reader looking at the folder on a forge sees
the folder, not a command, and the drift objection is answered by checking
rather than by refusing to write one. That is §5's own rule: a rule states how
it is checked.

Two checks, both confirmed to bite. index.py fails when the written order no
longer matches the records. records.py fails when a record has no topic or one
nobody defined -- the quiet failure being a record that vanishes from the order
rather than appearing in the wrong place.
2026-08-28 23:39:18 +02:00

70 lines
3.1 KiB
Markdown

---
topic: the mesh
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 2. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
> deliberation that happened.
## Context
The mesh is a set of machines that must call each other's capabilities. On the day the
repository was founded there was no inter-node transport at all — each node was configured
independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
## Considered options
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
participant. It also makes node liveness a caller's problem: a request to a sleeping node
is an error rather than a wait.
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
independent of demand, and request/reply has to be built on top of it by hand.
3. **A central message broker with per-node exchanges.** Chosen.
## Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange
named for itself and a request queue; a shared mesh exchange carries commands and events that
are not addressed to one node.
Three message shapes, and only three:
- **RPC** — request/reply, for calling a capability that lives on another node.
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
- **Events** — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
## Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker
connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually
right and occasionally the wrong thing entirely — a queued command for a node that never
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
where the capability is.
## References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.