papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist.
3.1 KiB
status, date, deciders, reconstructed
| status | date | deciders | reconstructed |
|---|---|---|---|
| accepted | 2026-02-25 | jochen | true |
1. Nodes communicate over a message broker, not over HTTP
Reconstructed after the fact from the evidence cited below. The decision was taken in implementation, not in a record; this document states what was decided and why, not a deliberation that happened.
Context
The mesh is a set of machines that must call each other's capabilities. On the day the repository was founded there was no inter-node transport at all — each node was configured independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
Considered options
- HTTP APIs between nodes. Rejected. Every node becomes a server that every other node must be able to reach, which the NAT case makes impossible without inbound tunnels to each participant. It also makes node liveness a caller's problem: a request to a sleeping node is an error rather than a wait.
- Polling a shared database. Rejected. Latency is the poll interval, load is constant and independent of demand, and request/reply has to be built on top of it by hand.
- A central message broker with per-node exchanges. Chosen.
Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange named for itself and a request queue; a shared mesh exchange carries commands and events that are not addressed to one node.
Three message shapes, and only three:
- RPC — request/reply, for calling a capability that lives on another node.
- Commands — instructions to do a stage of work, addressed by what is to be done.
- Events — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually right and occasionally the wrong thing entirely — a queued command for a node that never returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay where the capability is.
References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base:
mesh(transport, exchanges, queue naming),troubleshooting/amqp-credential-rotation. - The stall shape is recorded in
troubleshooting/empty-pipeline-blocks-the-queueandtroubleshooting/daemon-and-tool-server-share-a-request-queue.