HQ held only the to-be. Every reader had to already know the system the decisions were about, and an as-is claim had nowhere to live except inside an intention. Adds 02-DESIGN/00-as-is — eleven documents written from the implementation and the operational record, not from intent, including the parts nobody would choose again. The two existing designs move under 01-to-be. Layers are declared in frontmatter and never mix: a design that ships does not move, its as-is counterpart is written, and both stand. Back-fills adr/0001-0014 for decisions taken in implementation and never recorded — the broker, the module abstraction, the mesh database, managed files, provisioning, migrations, the workspace removal, failing loudly, the constitution, application placement, linking, the employee model, the artifact, the three silos. Each marked reconstructed, dated from the history, and citing the evidence it was recovered from. The two existing records renumber to 0015 and 0016 so the ledger runs oldest first; 0017 extends 0015 to modules outside the core, principle only — the domain list is deliberately not invented here. how-we-build.md becomes the source of the mesh constitution, with a sync playbook, so the enforced copy stops being the only one that is true. Process becomes explicit: five playbooks, eight thin skills that defer to them, a repository map, and AGENTS.md with CLAUDE.md as its include. The five Observations become 04-ISSUES 001-005 where they can be owned and closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge base. That claim is what decision 27 rests on, it was never checked, and the README now says so instead of repeating it. Also corrects the ADR index into something generated, the "02-DESIGN is empty" claim, the VISION.md pointer that did not survive the repo split, and a note asserting the symlink rule was contradicted — it was a misreading; the rule forbids hand-made links, the installer links by design.
3.1 KiB
status, date, deciders, reconstructed
| status | date | deciders | reconstructed |
|---|---|---|---|
| accepted | 2026-02-25 | jochen | true |
1. Nodes communicate over a message broker, not over HTTP
Reconstructed after the fact from the evidence cited below. The decision was taken in implementation, not in a record; this document states what was decided and why, not a deliberation that happened.
Context
The mesh is a set of machines that must call each other's capabilities. On the day the repository was founded there was no inter-node transport at all — each node was configured independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
Considered options
- HTTP APIs between nodes. Rejected. Every node becomes a server that every other node must be able to reach, which the NAT case makes impossible without inbound tunnels to each participant. It also makes node liveness a caller's problem: a request to a sleeping node is an error rather than a wait.
- Polling a shared database. Rejected. Latency is the poll interval, load is constant and independent of demand, and request/reply has to be built on top of it by hand.
- A central message broker with per-node exchanges. Chosen.
Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange named for itself and a request queue; a shared mesh exchange carries commands and events that are not addressed to one node.
Three message shapes, and only three:
- RPC — request/reply, for calling a capability that lives on another node.
- Commands — instructions to do a stage of work, addressed by what is to be done.
- Events — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually right and occasionally the wrong thing entirely — a queued command for a node that never returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay where the capability is.
References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base:
mesh(transport, exchanges, queue naming),troubleshooting/amqp-credential-rotation. - The stall shape is recorded in
troubleshooting/empty-pipeline-blocks-the-queueandtroubleshooting/daemon-and-tool-server-share-a-request-queue.