The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist.
This commit is contained in:
@@ -0,0 +1,113 @@
|
||||
---
|
||||
layer: as-is
|
||||
status: implemented
|
||||
code: [hal]
|
||||
updated: 2026-08-23
|
||||
decisions:
|
||||
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
|
||||
- 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md
|
||||
---
|
||||
|
||||
# The mesh and its transport
|
||||
|
||||
Two things make a set of machines into a mesh: a database that holds every binding, and a
|
||||
broker that carries every message. Neither is optional and neither is replaceable at present.
|
||||
|
||||
## The mesh database
|
||||
|
||||
One relational database holds the bindings. Its content divides cleanly:
|
||||
|
||||
| Holds | Describes |
|
||||
|---|---|
|
||||
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
|
||||
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
|
||||
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
|
||||
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
|
||||
| Grants | Which consumer holds which resource from which provider, with the credential |
|
||||
|
||||
The runtime loads this at startup. If the database cannot be reached it falls back to a local
|
||||
cache and continues.
|
||||
|
||||
**The repository contains none of this.** It defines what exists; the database defines what
|
||||
runs where. This is what makes the repository node-agnostic, and it is the property that lets
|
||||
anything about the mesh be published at all.
|
||||
|
||||
### What the cache costs
|
||||
|
||||
Running from cache is the difference between a node that survives a database outage and one
|
||||
that stops. It is the right trade and it has a cost that is worth naming: a node running from
|
||||
cache looks identical to a node running from the database. There is no age on the cache and
|
||||
nothing reports divergence, so a node can be running yesterday's assignment set indefinitely
|
||||
without any signal that it is.
|
||||
|
||||
### Where the settings are
|
||||
|
||||
A mesh-level setting is a value every node needs and no node owns — the broker's location, the
|
||||
forge's, the registry's. These live in the database rather than in any node's configuration,
|
||||
so a node learns where the broker is from the mesh rather than from a file, and moving the
|
||||
broker is a database change rather than a fleet-wide edit.
|
||||
|
||||
The circularity is real: a node must reach the database to learn where the broker is, and the
|
||||
database is itself a module the mesh provisions. It is resolved by the initialisation script
|
||||
that stands up the first node, which is the reason such a script exists separately from
|
||||
everything else.
|
||||
|
||||
## The broker
|
||||
|
||||
Every node connects **outbound** to a single broker. Nothing ever connects to a node.
|
||||
|
||||
Each node declares a topic exchange named for itself and consumes from its own request queue.
|
||||
A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and
|
||||
the events they emit.
|
||||
|
||||
Three message shapes, and only three:
|
||||
|
||||
- **Requests** expect a reply. This is how a capability on another node is called.
|
||||
- **Commands** instruct that a stage of work be done. They are addressed by what is to be done
|
||||
and consumed by whichever node is meant to do it.
|
||||
- **Events** state that something happened, addressed to nobody. Anything interested subscribes.
|
||||
|
||||
### Consequences the design accepts
|
||||
|
||||
A node behind a household connection with no inbound route participates exactly as a publicly
|
||||
named one does. This is the property the transport was chosen for.
|
||||
|
||||
A call to a node that is down **waits** rather than failing. Usually this is what is wanted. It
|
||||
is also how the mesh's most confusing stalls happen: a command queued for a node that never
|
||||
returns is a stall with no error anywhere, and the pipeline has produced this shape more than
|
||||
once — an empty pipeline that never completes blocks every pipeline queued behind it.
|
||||
|
||||
Two consumers accidentally sharing one queue silently split the traffic between them, each
|
||||
receiving half of what it expects. This has happened between a module's daemon and its
|
||||
capability server.
|
||||
|
||||
The broker is a single point of failure and a single point of trust. Its credential is
|
||||
mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker
|
||||
down.
|
||||
|
||||
## Reaching a capability on another node
|
||||
|
||||
A node hosts some capabilities and can reach the rest.
|
||||
|
||||
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a
|
||||
local stand-in that forwards over the broker. For anything hosted both locally and elsewhere
|
||||
it wraps the local one so a caller can name a target.
|
||||
|
||||
The effect is that a caller does not know where a capability runs. The important half is what
|
||||
does **not** move: the work happens where the capability is, so its credentials never leave
|
||||
that host. A remote call transports a request and a reply, never a secret.
|
||||
|
||||
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent
|
||||
at that moment is simply not discovered, and the node runs without that capability until it
|
||||
restarts. Nothing re-discovers on a schedule.
|
||||
|
||||
## Names and reachability
|
||||
|
||||
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the
|
||||
underlying network provides. A node's mesh name is its overlay address; its public name, if it
|
||||
has one, is a separate fact used by things outside the mesh.
|
||||
|
||||
Two lessons are embedded in that separation, both learned the expensive way. A name resolved
|
||||
by local multicast discovery introduces a delay and a failure mode that appears on one node and
|
||||
not others, so mesh names are not multicast names. And a node must not pin its own public name
|
||||
locally: the duplicate record breaks resolution for everything else that needs it.
|
||||
Reference in New Issue
Block a user