papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist.
114 lines
5.5 KiB
Markdown
114 lines
5.5 KiB
Markdown
---
|
|
layer: as-is
|
|
status: implemented
|
|
code: [hal]
|
|
updated: 2026-08-23
|
|
decisions:
|
|
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
|
|
- 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md
|
|
---
|
|
|
|
# The mesh and its transport
|
|
|
|
Two things make a set of machines into a mesh: a database that holds every binding, and a
|
|
broker that carries every message. Neither is optional and neither is replaceable at present.
|
|
|
|
## The mesh database
|
|
|
|
One relational database holds the bindings. Its content divides cleanly:
|
|
|
|
| Holds | Describes |
|
|
|---|---|
|
|
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
|
|
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
|
|
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
|
|
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
|
|
| Grants | Which consumer holds which resource from which provider, with the credential |
|
|
|
|
The runtime loads this at startup. If the database cannot be reached it falls back to a local
|
|
cache and continues.
|
|
|
|
**The repository contains none of this.** It defines what exists; the database defines what
|
|
runs where. This is what makes the repository node-agnostic, and it is the property that lets
|
|
anything about the mesh be published at all.
|
|
|
|
### What the cache costs
|
|
|
|
Running from cache is the difference between a node that survives a database outage and one
|
|
that stops. It is the right trade and it has a cost that is worth naming: a node running from
|
|
cache looks identical to a node running from the database. There is no age on the cache and
|
|
nothing reports divergence, so a node can be running yesterday's assignment set indefinitely
|
|
without any signal that it is.
|
|
|
|
### Where the settings are
|
|
|
|
A mesh-level setting is a value every node needs and no node owns — the broker's location, the
|
|
forge's, the registry's. These live in the database rather than in any node's configuration,
|
|
so a node learns where the broker is from the mesh rather than from a file, and moving the
|
|
broker is a database change rather than a fleet-wide edit.
|
|
|
|
The circularity is real: a node must reach the database to learn where the broker is, and the
|
|
database is itself a module the mesh provisions. It is resolved by the initialisation script
|
|
that stands up the first node, which is the reason such a script exists separately from
|
|
everything else.
|
|
|
|
## The broker
|
|
|
|
Every node connects **outbound** to a single broker. Nothing ever connects to a node.
|
|
|
|
Each node declares a topic exchange named for itself and consumes from its own request queue.
|
|
A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and
|
|
the events they emit.
|
|
|
|
Three message shapes, and only three:
|
|
|
|
- **Requests** expect a reply. This is how a capability on another node is called.
|
|
- **Commands** instruct that a stage of work be done. They are addressed by what is to be done
|
|
and consumed by whichever node is meant to do it.
|
|
- **Events** state that something happened, addressed to nobody. Anything interested subscribes.
|
|
|
|
### Consequences the design accepts
|
|
|
|
A node behind a household connection with no inbound route participates exactly as a publicly
|
|
named one does. This is the property the transport was chosen for.
|
|
|
|
A call to a node that is down **waits** rather than failing. Usually this is what is wanted. It
|
|
is also how the mesh's most confusing stalls happen: a command queued for a node that never
|
|
returns is a stall with no error anywhere, and the pipeline has produced this shape more than
|
|
once — an empty pipeline that never completes blocks every pipeline queued behind it.
|
|
|
|
Two consumers accidentally sharing one queue silently split the traffic between them, each
|
|
receiving half of what it expects. This has happened between a module's daemon and its
|
|
capability server.
|
|
|
|
The broker is a single point of failure and a single point of trust. Its credential is
|
|
mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker
|
|
down.
|
|
|
|
## Reaching a capability on another node
|
|
|
|
A node hosts some capabilities and can reach the rest.
|
|
|
|
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a
|
|
local stand-in that forwards over the broker. For anything hosted both locally and elsewhere
|
|
it wraps the local one so a caller can name a target.
|
|
|
|
The effect is that a caller does not know where a capability runs. The important half is what
|
|
does **not** move: the work happens where the capability is, so its credentials never leave
|
|
that host. A remote call transports a request and a reply, never a secret.
|
|
|
|
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent
|
|
at that moment is simply not discovered, and the node runs without that capability until it
|
|
restarts. Nothing re-discovers on a schedule.
|
|
|
|
## Names and reachability
|
|
|
|
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the
|
|
underlying network provides. A node's mesh name is its overlay address; its public name, if it
|
|
has one, is a separate fact used by things outside the mesh.
|
|
|
|
Two lessons are embedded in that separation, both learned the expensive way. A name resolved
|
|
by local multicast discovery introduces a delay and a failure mode that appears on one node and
|
|
not others, so mesh names are not multicast names. And a node must not pin its own public name
|
|
locally: the duplicate record breaks resolution for everything else that needs it.
|