HQ held only the to-be. Every reader had to already know the system the decisions were about, and an as-is claim had nowhere to live except inside an intention. Adds 02-DESIGN/00-as-is — eleven documents written from the implementation and the operational record, not from intent, including the parts nobody would choose again. The two existing designs move under 01-to-be. Layers are declared in frontmatter and never mix: a design that ships does not move, its as-is counterpart is written, and both stand. Back-fills adr/0001-0014 for decisions taken in implementation and never recorded — the broker, the module abstraction, the mesh database, managed files, provisioning, migrations, the workspace removal, failing loudly, the constitution, application placement, linking, the employee model, the artifact, the three silos. Each marked reconstructed, dated from the history, and citing the evidence it was recovered from. The two existing records renumber to 0015 and 0016 so the ledger runs oldest first; 0017 extends 0015 to modules outside the core, principle only — the domain list is deliberately not invented here. how-we-build.md becomes the source of the mesh constitution, with a sync playbook, so the enforced copy stops being the only one that is true. Process becomes explicit: five playbooks, eight thin skills that defer to them, a repository map, and AGENTS.md with CLAUDE.md as its include. The five Observations become 04-ISSUES 001-005 where they can be owned and closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge base. That claim is what decision 27 rests on, it was never checked, and the README now says so instead of repeating it. Also corrects the ADR index into something generated, the "02-DESIGN is empty" claim, the VISION.md pointer that did not survive the repo split, and a note asserting the symlink rule was contradicted — it was a misreading; the rule forbids hand-made links, the installer links by design.
5.5 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | |||
|---|---|---|---|---|---|---|---|
| as-is | implemented |
|
2026-08-23 |
|
The mesh and its transport
Two things make a set of machines into a mesh: a database that holds every binding, and a broker that carries every message. Neither is optional and neither is replaceable at present.
The mesh database
One relational database holds the bindings. Its content divides cleanly:
| Holds | Describes |
|---|---|
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
| Grants | Which consumer holds which resource from which provider, with the credential |
The runtime loads this at startup. If the database cannot be reached it falls back to a local cache and continues.
The repository contains none of this. It defines what exists; the database defines what runs where. This is what makes the repository node-agnostic, and it is the property that lets anything about the mesh be published at all.
What the cache costs
Running from cache is the difference between a node that survives a database outage and one that stops. It is the right trade and it has a cost that is worth naming: a node running from cache looks identical to a node running from the database. There is no age on the cache and nothing reports divergence, so a node can be running yesterday's assignment set indefinitely without any signal that it is.
Where the settings are
A mesh-level setting is a value every node needs and no node owns — the broker's location, the forge's, the registry's. These live in the database rather than in any node's configuration, so a node learns where the broker is from the mesh rather than from a file, and moving the broker is a database change rather than a fleet-wide edit.
The circularity is real: a node must reach the database to learn where the broker is, and the database is itself a module the mesh provisions. It is resolved by the initialisation script that stands up the first node, which is the reason such a script exists separately from everything else.
The broker
Every node connects outbound to a single broker. Nothing ever connects to a node.
Each node declares a topic exchange named for itself and consumes from its own request queue. A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and the events they emit.
Three message shapes, and only three:
- Requests expect a reply. This is how a capability on another node is called.
- Commands instruct that a stage of work be done. They are addressed by what is to be done and consumed by whichever node is meant to do it.
- Events state that something happened, addressed to nobody. Anything interested subscribes.
Consequences the design accepts
A node behind a household connection with no inbound route participates exactly as a publicly named one does. This is the property the transport was chosen for.
A call to a node that is down waits rather than failing. Usually this is what is wanted. It is also how the mesh's most confusing stalls happen: a command queued for a node that never returns is a stall with no error anywhere, and the pipeline has produced this shape more than once — an empty pipeline that never completes blocks every pipeline queued behind it.
Two consumers accidentally sharing one queue silently split the traffic between them, each receiving half of what it expects. This has happened between a module's daemon and its capability server.
The broker is a single point of failure and a single point of trust. Its credential is mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker down.
Reaching a capability on another node
A node hosts some capabilities and can reach the rest.
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a local stand-in that forwards over the broker. For anything hosted both locally and elsewhere it wraps the local one so a caller can name a target.
The effect is that a caller does not know where a capability runs. The important half is what does not move: the work happens where the capability is, so its credentials never leave that host. A remote call transports a request and a reply, never a secret.
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent at that moment is simply not discovered, and the node runs without that capability until it restarts. Nothing re-discovers on a schedule.
Names and reachability
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the underlying network provides. A node's mesh name is its overlay address; its public name, if it has one, is a separate fact used by things outside the mesh.
Two lessons are embedded in that separation, both learned the expensive way. A name resolved by local multicast discovery introduces a delay and a failure mode that appears on one node and not others, so mesh names are not multicast names. And a node must not pin its own public name locally: the duplicate record breaks resolution for everything else that needs it.