Files
hq/03-DESIGN/00-as-is/01-mesh-and-transport.md
T
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00

5.5 KiB

layer, status, code, updated, decisions
layer status code updated decisions
as-is implemented
hal
2026-08-23
02-DECISIONS/0001-nodes-communicate-over-a-broker.md
02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md

The mesh and its transport

Two things make a set of machines into a mesh: a database that holds every binding, and a broker that carries every message. Neither is optional and neither is replaceable at present.

The mesh database

One relational database holds the bindings. Its content divides cleanly:

Holds Describes
Node records Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text
Assignments Which node hosts which module, at which selection, and whether it starts automatically
Overrides Per-node, per-module values that take precedence over anything the manifest generates
Mesh settings Values every node reads — where the broker is, where the forge is, where the registry is
Grants Which consumer holds which resource from which provider, with the credential

The runtime loads this at startup. If the database cannot be reached it falls back to a local cache and continues.

The repository contains none of this. It defines what exists; the database defines what runs where. This is what makes the repository node-agnostic, and it is the property that lets anything about the mesh be published at all.

What the cache costs

Running from cache is the difference between a node that survives a database outage and one that stops. It is the right trade and it has a cost that is worth naming: a node running from cache looks identical to a node running from the database. There is no age on the cache and nothing reports divergence, so a node can be running yesterday's assignment set indefinitely without any signal that it is.

Where the settings are

A mesh-level setting is a value every node needs and no node owns — the broker's location, the forge's, the registry's. These live in the database rather than in any node's configuration, so a node learns where the broker is from the mesh rather than from a file, and moving the broker is a database change rather than a fleet-wide edit.

The circularity is real: a node must reach the database to learn where the broker is, and the database is itself a module the mesh provisions. It is resolved by the initialisation script that stands up the first node, which is the reason such a script exists separately from everything else.

The broker

Every node connects outbound to a single broker. Nothing ever connects to a node.

Each node declares a topic exchange named for itself and consumes from its own request queue. A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and the events they emit.

Three message shapes, and only three:

  • Requests expect a reply. This is how a capability on another node is called.
  • Commands instruct that a stage of work be done. They are addressed by what is to be done and consumed by whichever node is meant to do it.
  • Events state that something happened, addressed to nobody. Anything interested subscribes.

Consequences the design accepts

A node behind a household connection with no inbound route participates exactly as a publicly named one does. This is the property the transport was chosen for.

A call to a node that is down waits rather than failing. Usually this is what is wanted. It is also how the mesh's most confusing stalls happen: a command queued for a node that never returns is a stall with no error anywhere, and the pipeline has produced this shape more than once — an empty pipeline that never completes blocks every pipeline queued behind it.

Two consumers accidentally sharing one queue silently split the traffic between them, each receiving half of what it expects. This has happened between a module's daemon and its capability server.

The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker down.

Reaching a capability on another node

A node hosts some capabilities and can reach the rest.

At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a local stand-in that forwards over the broker. For anything hosted both locally and elsewhere it wraps the local one so a caller can name a target.

The effect is that a caller does not know where a capability runs. The important half is what does not move: the work happens where the capability is, so its credentials never leave that host. A remote call transports a request and a reply, never a secret.

Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent at that moment is simply not discovered, and the node runs without that capability until it restarts. Nothing re-discovers on a schedule.

Names and reachability

Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the underlying network provides. A node's mesh name is its overlay address; its public name, if it has one, is a separate fact used by things outside the mesh.

Two lessons are embedded in that separation, both learned the expensive way. A name resolved by local multicast discovery introduces a delay and a failure mode that appears on one node and not others, so mesh names are not multicast names. And a node must not pin its own public name locally: the duplicate record breaks resolution for everything else that needs it.