--- status: accepted date: 2026-02-25 deciders: jochen reconstructed: true --- # 1. Nodes communicate over a message broker, not over HTTP > Reconstructed after the fact from the evidence cited below. The decision was taken in > implementation, not in a record; this document states what was decided and why, not a > deliberation that happened. ## Context The mesh is a set of machines that must call each other's capabilities. On the day the repository was founded there was no inter-node transport at all — each node was configured independently and shared nothing at runtime. Three properties were required and are visible in everything built since: - A node behind a household NAT must participate fully. It can dial out; nothing can dial in. - A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request should wait, not error. - Adding a node must not require editing anything on the nodes that already exist. ## Considered options 1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node must be able to reach, which the NAT case makes impossible without inbound tunnels to each participant. It also makes node liveness a caller's problem: a request to a sleeping node is an error rather than a wait. 2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and independent of demand, and request/reply has to be built on top of it by hand. 3. **A central message broker with per-node exchanges.** Chosen. ## Decision All inter-node communication goes through a message broker. Every node owns a topic exchange named for itself and a request queue; a shared mesh exchange carries commands and events that are not addressed to one node. Three message shapes, and only three: - **RPC** — request/reply, for calling a capability that lives on another node. - **Commands** — instructions to do a stage of work, addressed by what is to be done. - **Events** — statements that something happened, addressed to nobody. Every node dials the broker outbound. Nothing dials a node. ## Consequences - NAT stops being an architectural concern. A node's reachability is a property of the broker connection, not of its network position. - A call to a node that is down waits in that node's queue instead of failing. This is usually right and occasionally the wrong thing entirely — a queued command for a node that never returns is a stall with no error, which is the failure shape this mesh keeps rediscovering. - The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation. - Tools never leave the host: a remote call proxies over the broker and the credentials stay where the capability is. ## References - The broker was stood up on 2026-02-25, the second day of the repository. - Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`. - The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and `troubleshooting/daemon-and-tool-server-share-a-request-queue`.