Files
hq/02-DECISIONS/0002-nodes-communicate-over-a-broker.md
jschoubben b4607dfc03 Numbers are identity; the reading order is a generated, checked index
Decided after measuring what renumbering actually costs: 96 references in code
comments across two repositories, none of which would have failed to compile.
They would have pointed at the wrong reasoning, which is worse than a broken
link because nothing reports it.

So a number identifies a record and never changes. It cannot also be a
position -- a position moves when the set changes, and an identity that moves
is not one.

The reading order moves into an index generated from each record's `topic:`.
Six topics, in the order somebody learns the system.

The index is WRITTEN rather than only generated on demand, which reverses what
this repository previously said. The reason it said otherwise is that a
hand-written index drifts -- but a reader looking at the folder on a forge sees
the folder, not a command, and the drift objection is answered by checking
rather than by refusing to write one. That is §5's own rule: a rule states how
it is checked.

Two checks, both confirmed to bite. index.py fails when the written order no
longer matches the records. records.py fails when a record has no topic or one
nobody defined -- the quiet failure being a record that vanishes from the order
rather than appearing in the wrong place.
2026-08-28 23:39:18 +02:00

3.1 KiB

topic, status, date, deciders, reconstructed
topic status date deciders reconstructed
the mesh accepted 2026-02-25 jochen true

2. Nodes communicate over a message broker, not over HTTP

Reconstructed after the fact from the evidence cited below. The decision was taken in implementation, not in a record; this document states what was decided and why, not a deliberation that happened.

Context

The mesh is a set of machines that must call each other's capabilities. On the day the repository was founded there was no inter-node transport at all — each node was configured independently and shared nothing at runtime.

Three properties were required and are visible in everything built since:

  • A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
  • A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request should wait, not error.
  • Adding a node must not require editing anything on the nodes that already exist.

Considered options

  1. HTTP APIs between nodes. Rejected. Every node becomes a server that every other node must be able to reach, which the NAT case makes impossible without inbound tunnels to each participant. It also makes node liveness a caller's problem: a request to a sleeping node is an error rather than a wait.
  2. Polling a shared database. Rejected. Latency is the poll interval, load is constant and independent of demand, and request/reply has to be built on top of it by hand.
  3. A central message broker with per-node exchanges. Chosen.

Decision

All inter-node communication goes through a message broker. Every node owns a topic exchange named for itself and a request queue; a shared mesh exchange carries commands and events that are not addressed to one node.

Three message shapes, and only three:

  • RPC — request/reply, for calling a capability that lives on another node.
  • Commands — instructions to do a stage of work, addressed by what is to be done.
  • Events — statements that something happened, addressed to nobody.

Every node dials the broker outbound. Nothing dials a node.

Consequences

  • NAT stops being an architectural concern. A node's reachability is a property of the broker connection, not of its network position.
  • A call to a node that is down waits in that node's queue instead of failing. This is usually right and occasionally the wrong thing entirely — a queued command for a node that never returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
  • The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation.
  • Tools never leave the host: a remote call proxies over the broker and the credentials stay where the capability is.

References

  • The broker was stood up on 2026-02-25, the second day of the repository.
  • Knowledge base: mesh (transport, exchanges, queue naming), troubleshooting/amqp-credential-rotation.
  • The stall shape is recorded in troubleshooting/empty-pipeline-blocks-the-queue and troubleshooting/daemon-and-tool-server-share-a-request-queue.