Files
hq/02-DESIGN/00-as-is/01-mesh-and-transport.md
T
jschoubben 702efca6bb Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
2026-08-23 03:08:26 +02:00

114 lines
5.5 KiB
Markdown

---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0001-nodes-communicate-over-a-broker.md
- adr/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh and its transport
Two things make a set of machines into a mesh: a database that holds every binding, and a
broker that carries every message. Neither is optional and neither is replaceable at present.
## The mesh database
One relational database holds the bindings. Its content divides cleanly:
| Holds | Describes |
|---|---|
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
| Grants | Which consumer holds which resource from which provider, with the credential |
The runtime loads this at startup. If the database cannot be reached it falls back to a local
cache and continues.
**The repository contains none of this.** It defines what exists; the database defines what
runs where. This is what makes the repository node-agnostic, and it is the property that lets
anything about the mesh be published at all.
### What the cache costs
Running from cache is the difference between a node that survives a database outage and one
that stops. It is the right trade and it has a cost that is worth naming: a node running from
cache looks identical to a node running from the database. There is no age on the cache and
nothing reports divergence, so a node can be running yesterday's assignment set indefinitely
without any signal that it is.
### Where the settings are
A mesh-level setting is a value every node needs and no node owns — the broker's location, the
forge's, the registry's. These live in the database rather than in any node's configuration,
so a node learns where the broker is from the mesh rather than from a file, and moving the
broker is a database change rather than a fleet-wide edit.
The circularity is real: a node must reach the database to learn where the broker is, and the
database is itself a module the mesh provisions. It is resolved by the initialisation script
that stands up the first node, which is the reason such a script exists separately from
everything else.
## The broker
Every node connects **outbound** to a single broker. Nothing ever connects to a node.
Each node declares a topic exchange named for itself and consumes from its own request queue.
A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and
the events they emit.
Three message shapes, and only three:
- **Requests** expect a reply. This is how a capability on another node is called.
- **Commands** instruct that a stage of work be done. They are addressed by what is to be done
and consumed by whichever node is meant to do it.
- **Events** state that something happened, addressed to nobody. Anything interested subscribes.
### Consequences the design accepts
A node behind a household connection with no inbound route participates exactly as a publicly
named one does. This is the property the transport was chosen for.
A call to a node that is down **waits** rather than failing. Usually this is what is wanted. It
is also how the mesh's most confusing stalls happen: a command queued for a node that never
returns is a stall with no error anywhere, and the pipeline has produced this shape more than
once — an empty pipeline that never completes blocks every pipeline queued behind it.
Two consumers accidentally sharing one queue silently split the traffic between them, each
receiving half of what it expects. This has happened between a module's daemon and its
capability server.
The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker
down.
## Reaching a capability on another node
A node hosts some capabilities and can reach the rest.
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a
local stand-in that forwards over the broker. For anything hosted both locally and elsewhere
it wraps the local one so a caller can name a target.
The effect is that a caller does not know where a capability runs. The important half is what
does **not** move: the work happens where the capability is, so its credentials never leave
that host. A remote call transports a request and a reply, never a secret.
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent
at that moment is simply not discovered, and the node runs without that capability until it
restarts. Nothing re-discovers on a schedule.
## Names and reachability
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the
underlying network provides. A node's mesh name is its overlay address; its public name, if it
has one, is a separate fact used by things outside the mesh.
Two lessons are embedded in that separation, both learned the expensive way. A name resolved
by local multicast discovery introduces a delay and a failure mode that appears on one node and
not others, so mesh names are not multicast names. And a node must not pin its own public name
locally: the duplicate record breaks resolution for everything else that needs it.