Base layer: the mesh as it is, under the mesh as it should be

HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
This commit is contained in:
2026-08-23 03:08:26 +02:00
parent cf9357e8e9
commit 702efca6bb
74 changed files with 3676 additions and 138 deletions
+98
View File
@@ -0,0 +1,98 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0001-nodes-communicate-over-a-broker.md
- adr/0002-everything-is-a-module.md
- adr/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh as it stands
A set of machines, each running the same runtime, each loading only the parts of the catalogue
it has been assigned. They hold no shared filesystem and make no direct connections to one
another. What makes them a mesh is a database that knows what should run where, and a message
broker that carries everything between them.
## Three nouns
**A node** is a machine that runs the runtime. Nodes differ in what they are assigned and in
what they can reach — some carry a public name, some sit behind a household connection with no
inbound route at all — and the mesh is designed so that difference stays a property rather than
becoming a special case. A node holds no authoritative state: everything it needs is derived
onto it and can be regenerated.
**A module** is a directory with a manifest, and it is the only unit the mesh installs. A
containerised service is a module. A set of capabilities with no service behind them is a
module. A bare marker whose whole content is that a node has it is a module. The mesh's own
components are modules on exactly the same terms as everything else it carries
([ADR 0002](../../adr/0002-everything-is-a-module.md)).
**An agent** is a participant. Some agents are human. What differs is modality — how the agent
acts — and not category: both hold identity, both act, both accumulate memory
([ADR 0012](../../adr/0012-agents-are-persistent-employees.md)).
## Where truth lives
The repository defines **what exists**: the modules, what each declares, how each is built.
The mesh database defines **what runs where**: which node is assigned which module, at which
selection, with which overrides, plus the settings every node reads. No node-to-module mapping
is ever committed ([ADR 0003](../../adr/0003-the-mesh-database-is-the-source-of-truth.md)).
Everything on a node's disk is **derived** from those two, and is regenerated rather than
edited ([ADR 0004](../../adr/0004-managed-files-are-generated-never-edited.md)). A node that
loses its database keeps running from a local cache, which is deliberate and has the obvious
cost: the cache carries no indication of its own age.
## How anything moves
Nothing dials a node. Every node dials the broker outbound, owns an exchange named for itself,
and consumes from its own request queue
([ADR 0001](../../adr/0001-nodes-communicate-over-a-broker.md)). Three message shapes carry
everything: requests expecting a reply, commands instructing that a stage of work be done, and
events stating that something happened.
A capability that lives on another node is reached the same way a local one is. At startup a
node asks its peers what they host and creates a local stand-in for each remote capability, so
the caller does not know or care where the work happens. Credentials never travel: the call
goes to where the capability is.
## How change reaches a node
A push to the forge is the only trigger. What follows is three silos with deliberately
different cardinality: compile once, package and upload once, then install-configure-start-
verify **on every assigned node**
([ADR 0014](../../adr/0014-build-publish-and-deploy-are-three-silos.md)). What travels between
build and node is a self-contained build output, so a deploy is extract-and-run and touches no
network ([ADR 0013](../../adr/0013-an-artifact-is-build-output.md)).
Modules are resolved into dependency levels and a level completes before the next begins, so a
module always builds against its dependencies as they were just published.
## What the mesh does for a module
A module declares what it **provides** and what it **requires**. The mesh satisfies the
requirement: it creates the resource, generates the credential, records the grant, and writes
the values where the module will read them. The module never learns which node its database
lives on, and nobody ever writes a credential by hand
([ADR 0005](../../adr/0005-capabilities-are-provisioned-on-declaration.md)).
This is the property the mesh's whole shape rests on, and it is why provisioning is treated as
a core concern rather than as plumbing.
## The shape of its failures
Worth stating in an overview, because it is the most consistent thing about the system: the
mesh's expensive faults are almost never crashes. They are operations that reported success
and did nothing — a download that half-completed, a hook that was never called because it was
named for a feature the module does not declare, a stage that reported it had dispatched a
message rather than that the effect happened, a package that 404ed from every mirror while the
job went green.
[ADR 0008](../../adr/0008-a-failed-step-fails-the-job.md) is the response, and it is applied
instance by instance rather than enforced by a mechanism. New instances are still being found.
That is an as-is fact, not a criticism: it is the single most useful thing to know about this
system before changing it.
+113
View File
@@ -0,0 +1,113 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0001-nodes-communicate-over-a-broker.md
- adr/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh and its transport
Two things make a set of machines into a mesh: a database that holds every binding, and a
broker that carries every message. Neither is optional and neither is replaceable at present.
## The mesh database
One relational database holds the bindings. Its content divides cleanly:
| Holds | Describes |
|---|---|
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
| Grants | Which consumer holds which resource from which provider, with the credential |
The runtime loads this at startup. If the database cannot be reached it falls back to a local
cache and continues.
**The repository contains none of this.** It defines what exists; the database defines what
runs where. This is what makes the repository node-agnostic, and it is the property that lets
anything about the mesh be published at all.
### What the cache costs
Running from cache is the difference between a node that survives a database outage and one
that stops. It is the right trade and it has a cost that is worth naming: a node running from
cache looks identical to a node running from the database. There is no age on the cache and
nothing reports divergence, so a node can be running yesterday's assignment set indefinitely
without any signal that it is.
### Where the settings are
A mesh-level setting is a value every node needs and no node owns — the broker's location, the
forge's, the registry's. These live in the database rather than in any node's configuration,
so a node learns where the broker is from the mesh rather than from a file, and moving the
broker is a database change rather than a fleet-wide edit.
The circularity is real: a node must reach the database to learn where the broker is, and the
database is itself a module the mesh provisions. It is resolved by the initialisation script
that stands up the first node, which is the reason such a script exists separately from
everything else.
## The broker
Every node connects **outbound** to a single broker. Nothing ever connects to a node.
Each node declares a topic exchange named for itself and consumes from its own request queue.
A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and
the events they emit.
Three message shapes, and only three:
- **Requests** expect a reply. This is how a capability on another node is called.
- **Commands** instruct that a stage of work be done. They are addressed by what is to be done
and consumed by whichever node is meant to do it.
- **Events** state that something happened, addressed to nobody. Anything interested subscribes.
### Consequences the design accepts
A node behind a household connection with no inbound route participates exactly as a publicly
named one does. This is the property the transport was chosen for.
A call to a node that is down **waits** rather than failing. Usually this is what is wanted. It
is also how the mesh's most confusing stalls happen: a command queued for a node that never
returns is a stall with no error anywhere, and the pipeline has produced this shape more than
once — an empty pipeline that never completes blocks every pipeline queued behind it.
Two consumers accidentally sharing one queue silently split the traffic between them, each
receiving half of what it expects. This has happened between a module's daemon and its
capability server.
The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker
down.
## Reaching a capability on another node
A node hosts some capabilities and can reach the rest.
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a
local stand-in that forwards over the broker. For anything hosted both locally and elsewhere
it wraps the local one so a caller can name a target.
The effect is that a caller does not know where a capability runs. The important half is what
does **not** move: the work happens where the capability is, so its credentials never leave
that host. A remote call transports a request and a reply, never a secret.
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent
at that moment is simply not discovered, and the node runs without that capability until it
restarts. Nothing re-discovers on a schedule.
## Names and reachability
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the
underlying network provides. A node's mesh name is its overlay address; its public name, if it
has one, is a separate fact used by things outside the mesh.
Two lessons are embedded in that separation, both learned the expensive way. A name resolved
by local multicast discovery introduces a delay and a failure mode that appears on one node and
not others, so mesh names are not multicast names. And a node must not pin its own public name
locally: the duplicate record breaks resolution for everything else that needs it.
@@ -0,0 +1,113 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0002-everything-is-a-module.md
- adr/0006-schema-changes-are-numbered-migrations.md
- adr/0007-no-npm-workspace.md
---
# Modules, manifests and features
Everything the mesh installs is a module: a directory with a manifest. There is no second
mechanism.
## What a manifest declares
| Declares | Meaning |
|---|---|
| Identity | Name and version. **Version is owned by the builder** — a hand-edited version is a defect, and reviewers revert it. |
| Environment | Every variable the module reads, with how each is produced: a static default, a generated secret, a value pulled from the node's own record, or a template composed from the others. A variable not declared here is invisible to the mesh and will not be generated, injected or audited. |
| What it provides | The resource type this module can provision for others, and the network on which it is reachable. |
| What it requires | Resources it needs from other modules, and the mapping from each resource's connection fields onto its own environment variables. |
| Service shape | The primary container, and the data directories that must exist with the right ownership before it starts. |
| Exposure | The public names this module's interfaces answer on, declared portably so the reverse proxy configuration can be generated rather than written. |
| Images | Container images this module builds, so the pipeline builds and publishes them before publishing the module. |
## Features are the unit of work
A module is not the unit the pipeline addresses. A **feature** is.
A feature is a kind of content a module can carry: a service, a set of capabilities, a
long-running process, managed configuration files, migrations, firewall rules, an installable
application. One module can carry several.
Features are **detected from directory contents**, not declared. A module with a capabilities
directory has that feature; a module with a daemon directory has that one. An explicit
declaration was supported and is now discouraged, because a declared list and the directory it
describes drift, and the directory is the one that is true.
Every pipeline command and event names a feature. There is no per-module build.
### What detection costs
Detection makes the manifest shorter and the truth singular, and it makes the directory
structure load-bearing in a way that is not obvious from reading a manifest. Renaming a
directory changes what a module *is*, silently. The recurring failure is a hook named for a
feature the module does not carry: it is skipped without complaint, and the change it was
supposed to make simply never happens.
An unknown key in a manifest is likewise accepted in silence — which is how a firewall rule
can appear to restrict a port and restrict nothing (see
[`04-ISSUES/003`](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)).
## Kinds of module
The kinds are not a type system — they are what the detected features add up to.
- **A service module** carries a container definition. It gets a runtime directory, generated
environment, data directories, and is started under supervision.
- **A capability module** carries capabilities and no service. It contributes what a node can
do, locally and to its peers.
- **A flag module** carries nothing but a manifest. Its presence in a node's assignment is the
entire content: it gates behaviour elsewhere.
- **Combinations** are ordinary. A database module is a service *and* a capability provider
*and* a provisioner.
## Selections
A module can ship variants of the same feature and a node takes the one that fits it — a build
for one accelerator or another, a configuration for a public node or a private one. The
artifact stays selection-blind; the choice is a property of the assignment, held in the mesh
database.
This is one of the areas where behaviour has repeatedly diverged from intent, in both
directions: selection files that were never packaged into the artifact at all, and a stale
staged override on a node that silently won over the newly selected one. Both classes are
recorded in the knowledge base; both presented as "the change did not apply" with no error.
## Dependencies between modules
Modules depend on each other, above all on the shared library they all build against. There is
**no workspace** ([ADR 0007](../../adr/0007-no-npm-workspace.md)): each module is a standalone
package consuming published dependencies, including the mesh's own.
The pipeline resolves modules into dependency **levels** and completes a level before starting
the next, so a module always builds against its dependencies as just published.
The cost is a publish-and-consume round trip for every cross-package change, and the absence of
any repository-wide build. One thing that assumed a repository-wide build has stayed broken
since (see
[`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)).
## Persistent state
A module that owns state owns its migrations: numbered, written in the module's own language,
compiled with it, frozen once they have run anywhere, and idempotent so that re-running is safe
([ADR 0006](../../adr/0006-schema-changes-are-numbered-migrations.md)).
Two kinds exist and the distinction matters: migrations against the module's **own** local
state, and migrations against a **provisioned** resource, which run on the node that consumes
the resource rather than on the node that built the module.
## A documented rule with no enforcement
Every module exposing capabilities is documented as required to declare the mesh's core runtime
as a dependency. **Zero of the catalogue's modules do.**
This is recorded here rather than quietly corrected, because it is the clearest instance of the
rule this repository states about itself: a rule whose enforcement does not exist is
indistinguishable from a wrong one, and costs more, because people believe it. Whether the rule
or the catalogue is wrong has not been decided.
+88
View File
@@ -0,0 +1,88 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0005-capabilities-are-provisioned-on-declaration.md
- adr/0004-managed-files-are-generated-never-edited.md
---
# Provisioning
A module states what it needs. The mesh makes it exist, generates the credential, records the
grant, and puts the values where the module will read them. Nobody writes a credential and
nobody writes a topology.
This is the mesh's core concern rather than its plumbing — the property everything else is
built on.
## The declaration
A **provider** declares the resource type it can create and the network on which that resource
is reachable from a container.
A **consumer** declares, per requirement: the provider module, the resource type, optionally a
name and a target node, and a mapping from the resource's connection fields onto its own
environment variables.
The consumer must also declare those variables as empty in its environment section. A mapped
field with no declared variable is dropped — the value is produced and then discarded, because
the generator only emits variables the manifest knows about.
## What happens
Inside the pipeline the sequence is orchestrated by the coordinator and is not optional:
1. The provisioner creates the resource and its credential, records the grant, and writes the
mapped values as database overrides.
2. The synchroniser regenerates the module's environment from those overrides.
3. Only then is the service started.
Each step waits for the previous one's event. A module that declares no requirements skips the
first step entirely — which is correct, and means the absence of provisioning is
indistinguishable from provisioning that did not run.
Outside the pipeline the same two steps exist as direct operations, for debugging. They are
not the normal path.
## What can be provisioned
Providers exist for relational databases of two kinds, an object store, a cache, message-broker
virtual hosts and access, an identity provider's clients, and a download category. Each returns
the connection fields a consumer maps from — host, port, user, password, and whatever else the
resource type implies.
A provider with no registered provisioner still records a grant, with an empty connection. This
is deliberate and easy to misread: the grant exists, so the requirement looks satisfied, and
nothing was created.
## Cross-node grants
A requirement may name the node whose provider should satisfy it. The grant records consumer
node and provider node separately, so a module on one node holding a database on another is
the ordinary case rather than a special one.
For a containerised consumer, the provisioner returns the provider's routable name rather than
a loopback address — the value has to be correct from **inside** a container on another
machine.
## Where this hurts
**Rotation has no fan-out.** A resource whose credential is shared by several consumers can be
rotated by provisioning a new one, and the peers holding the old credential are not told. This
has locked the mesh out of its own broker, and has caused a node's adoption to rotate a live
shared password without informing anything that held it. Granting is easy; regranting is not
modelled.
**A frozen password outlives its generation.** A generated password is written once. If the
resource's persistent data directory already exists from an earlier initialisation, the stored
credential and the generated one diverge, and the symptom is an authentication failure that
looks like a configuration error.
**Migrations against a provisioned resource run on the consumer's node**, not the build node,
and read the deployed artifact rather than the source tree. Both facts were wrong in the
implementation for a period during which new provisioning migrations silently never ran.
**A grant is not a check.** The record says a resource was provisioned. Nothing verifies it
still exists, still has the recorded credential, or is reachable from where the consumer runs.
+100
View File
@@ -0,0 +1,100 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0014-build-publish-and-deploy-are-three-silos.md
- adr/0013-an-artifact-is-build-output.md
- adr/0008-a-failed-step-fails-the-job.md
---
# Delivery — from a push to a running node
One trigger, three silos, and a fan-out point that is the most consequential boundary in the
mesh.
## The trigger
A push to the forge. The forge calls a webhook; the receiving module verifies its signature
and emits an event. The coordinator resolves which modules the pushed commits affect, orders
them into dependency levels, and creates a pipeline per level.
There is no second path. A manual trigger exists for a module the push detection missed, and
using it routinely is a sign the detection is wrong rather than a workflow.
Detection is the pipeline's most fragile input. It has failed for reasons that have nothing to
do with the change: a webhook truncating its commit list on a large merge, a forge address
whose port broke the module-path matching. The characteristic outcome is the bad one — **a
merge that created no pipeline, and nothing said so**.
## Three silos
Cardinality is the whole point, and the three differ
([ADR 0014](../../adr/0014-build-publish-and-deploy-are-three-silos.md)):
| Silo | Runs | Where | Does |
|---|---|---|---|
| **build** | once per module feature | the build node | compile and bundle into a self-contained output |
| **publish** | once per module feature | the build node | package that output and upload it; a package-registry feature publishes here |
| **deploy** | once per module feature **per node** | every assigned node | install, configure, start, verify |
Commands and events are addressed per feature, not per module.
Build hands over a **staged tree**, not a package. Packaging belongs to publish, so a failed
upload retries by re-packaging rather than by re-sending something stale, and build never needs
to know how each module composes its artifact. The handover is a tree in a known location,
because the build's own working directory is reference-counted and may be gone by the time a
later stage runs.
## The artifact
The artifact is **build output** — compiled and bundled with its dependency graph inlined —
never a filtered copy of source ([ADR 0013](../../adr/0013-an-artifact-is-build-output.md)).
A deploy is extract-and-run and touches no network.
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after deploy — migrations reading a source layout, provisioning
scripts reading a source layout, selection files never packaged at all.
## The fan-out, and the build node
After publish, work fans out to every assigned node. The build node is the only node that has
already passed through two silos when this happens, and that asymmetry has its own defect
class: anything advancing a node's stage must account for **both** pre-fan-out stages. Code
that knew only about the first parked the build node forever while every other node deployed
cleanly — and the recovery sweeper, which knew the same subset, reported nothing to recover.
A recovery mechanism that knows less than the thing it guards is worse than none, because it
reports success over a stall it cannot see.
## Levels
A level completes before the next begins, so a module builds against its dependencies as they
were just published. The shared library is at level zero, which means anything that breaks it
breaks the first module of every cascade.
## What green proves
**A green pipeline proves transport, not effect.** The stages report that a message was
dispatched and accepted, which is not the same as the thing being running, correct, or present.
This is the mesh's most consistent failure shape and it is not incidental to the design — it is
what the stage reporting currently measures. Documented instances include a service reported
started when the container command merely returned, an image pull failure that did not fail the
deploy, a package install that 404ed from every mirror while the job went green (see
[`04-ISSUES/001`](../../04-ISSUES/001-failed-package-install-reports-success/00-report.md)),
and a node left on old code after a failed artifact download with a version marker that had
already advanced.
A verify stage exists to close this gap. It has been built and, for a period, was never
scheduled, because the coordinator's stage list did not include it.
## What is not covered
The end-to-end harness for this pipeline has not built since the workspace was removed, and
nothing reports that nothing runs it
([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). The
pipeline's coverage is currently assumed rather than checked — which is the same class of claim
this repository exists to make people stop making.
@@ -0,0 +1,102 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0002-everything-is-a-module.md
- adr/0011-the-installer-owns-linking.md
---
# The node runtime, and how a node comes into being
Every node runs the same runtime. What differs is its assignment.
## Two modes, coexisting
The runtime runs in two modes at once on any node that needs both.
**Daemon** mode is a headless consumer: it connects to the broker, consumes the node's request
queue, routes each request to a local capability, and emits the node's lifecycle events. This
is what makes a node a participant — it is reachable whether or not anyone is logged in.
**Interactive** mode exposes the node's capabilities to a session on that machine over a local
protocol. Its capability surface is larger, because it includes stand-ins for every capability
discovered on peers.
The two are the same code with the same catalogue. A capability is written once and is
available to both.
## Anatomy naming, and what it obscures
The runtime's components are named after brain anatomy: an entry point that bootstraps, a
headless listener, an interactive surface, an installer daemon, and a provisioner daemon.
This is the mesh's most-cited naming problem and it belongs in the as-is layer because it is
what a reader will actually encounter. The names are evocative and describe nothing: the most
suggestive word in the system names the node runtime, and the component whose manifest says
"mesh messaging" is documented elsewhere as the interactive runtime. Anatomy makes attractive
names and poor boundaries.
[ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md) replaces this with names
taken from what each part owns. Until then, this is the vocabulary in the code.
## Starting a module
For each assigned module the runtime, at startup:
1. Reads the manifest, if there is one — a flag module has nothing to read.
2. Skips a module's capabilities if a variable they require is unset. This is quiet by design
and hard to distinguish from a module that has no capabilities.
3. Loads the capabilities the module carries.
4. For a service module, ensures it is installed and running.
Installation is idempotent and does the unglamorous work: ensure the runtime directory,
reconcile the link from the catalogue's definition into it, generate the environment, run any
outstanding local migrations, create data directories with the right ownership, then start the
service under supervision.
**The installer is the only thing that creates a link** ([ADR
0011](../../adr/0011-the-installer-owns-linking.md)). It reconciles rather than assumes: a
missing link is created, a stale one repointed, and a real file found where a link belongs is
adopted into the node's override area and replaced. Nothing else — not a hook, not a fix, not a
person debugging — creates one.
## Supervision
Services run under the host's init system via a templated unit, one instance per module. It is
a thin layer: the unit starts and stops a container group.
Whether the mesh keeps this, drops the per-module layer, containerises the daemons, or writes
its own supervisor is **open** — costed in research effort 003 and deliberately undecided. It
no longer gates anything.
Two failures worth knowing about, both in the shape of "the change did not apply". A
per-instance copy of the unit template shadows the template, so edits to the template do
nothing. And a session-scoped one-shot job loses the environment it was given, because the
import is one-time and not persisted.
## How a node comes into being
Three bootstrap scripts, and which one runs depends on the situation:
- **First node.** Nothing exists yet, so the script stands up the database the rest of the
mesh reads from, publishes the catalogue, and starts the mesh. This resolves the
circularity of a mesh whose source of truth is itself a provisioned module.
- **Joining.** The node registers, takes its assignment from the database, and syncs.
- **Rescue.** A node that cannot reach the mesh is brought back far enough to.
Bringing a node into being is therefore a **database operation with a script attached**, not a
checkout. There is no per-node content in the repository to copy.
Adoption of a pre-existing machine's configuration was the original path and is now a legacy
one, explicitly out of scope for the lab
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)).
## Node identity
Each node carries an identity text in its own record, which the daemon writes onto the node at
startup so that a session on that machine knows which node it is on and how that node
presents itself. It carries identity only; shared rules live separately.
Like everything else derived onto a node, it is generated and not edited there.
@@ -0,0 +1,86 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0004-managed-files-are-generated-never-edited.md
- adr/0005-capabilities-are-provisioned-on-declaration.md
---
# Configuration and secrets
Every value a module reads at runtime comes from a file on a node's disk. Every one of those
files is **generated**.
## The rule
A managed file is derived from the mesh database. A synchroniser rewrites it when the values
behind it change. The write path is the mesh operation that owns the value; the file is an
output ([ADR 0004](../../adr/0004-managed-files-are-generated-never-edited.md)).
An edit to a managed file survives until the next synchronisation and is then overwritten
silently, taking whatever it was fixing with it — bringing back the bug the edit had removed,
with a delay, and with no error to connect the two events.
There is a way to ask whether a given file is managed. That question has to be asked, because
the answer is not visible from the file.
## What is managed
Generated environment files for each module, service definitions in each module's runtime
directory, managed configuration files a module declares, and node-level settings. The list is
not a category — it is whatever a synchroniser claims, which is why the question is asked of
the tooling rather than answered from a rule.
## How a value is decided
Values resolve by precedence, highest first:
1. **A database override** — set deliberately, or written by a provisioner.
2. **The value already in the generated file** — preserved for anything without an override.
3. **A generated value** — a random secret, a template composed from other variables, or a
value pulled from the node's own record.
4. **The manifest default.**
Two consequences follow, and both are subtle enough to have caused confusion.
The second rule is what keeps a generated password **stable** across regenerations. It is not
an oversight; without it every regeneration would issue a new secret and break whatever holds
the old one.
The same rule means that **changing a manifest's default does not change anything already
using it**. The existing file's value wins. The new default reaches only installations that
never had one.
Removing a declaration is worse than changing it: the old override row and the file it produced
are both left behind. Configuration is additive in practice, whatever the manifest says.
## Secrets
Generated secrets are produced by the mesh, never authored. Provisioned credentials arrive as
database overrides written by the provisioner and are marked as such, so they can be
distinguished from a deliberate override and cleaned up when the grant is removed
([ADR 0005](../../adr/0005-capabilities-are-provisioned-on-declaration.md)).
Nothing in the repository contains a credential. The repository has no per-node content at all,
which is what makes that guarantee structural rather than a matter of care.
Two known weaknesses, both recorded rather than resolved:
**Generated environment files were world-readable.** The mode passed at write time only applies
when the file is created, so regeneration left the previous mode in place. The class of bug is
worth remembering beyond the instance: a permission set at creation is not a permission
maintained.
**Rotation is not a mesh operation.** Secrets can be generated and granted; there is no
mechanism that rotates one and informs everything holding it. Where a rotation has been done,
it has been done by hand, and doing it wrong has taken services down.
## Node-level and mesh-level values
A node-level value applies to everything on one node. A mesh-level value applies everywhere and
is read by every node — the broker's location is the canonical example, and a wrong one is how
a mesh fails to form.
Neither is a file that anyone edits. Both are database records that produce files.
+78
View File
@@ -0,0 +1,78 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions: []
---
# Knowledge
The mesh keeps two knowledge stores. They are not redundant, and knowing which is which is the
difference between finding an answer in one search and rediscovering it over several hours.
## The operational memory
A store of operational notes, written and read by whoever — human or agent — is working. Each
note is a slug and a body: how something works, what went wrong, what the fix was, what
assumption turned out to be false.
It is indexed on **symptoms**. The entry someone needs is usually titled after the error they
are staring at, which is why the standing instruction is to search the literal error text
before forming a hypothesis rather than after one fails.
Its content is overwhelmingly the record of previous debugging: a large body of
troubleshooting entries, module conventions, and standing notes about work that is open. It is
the mesh's institutional memory of *what has already gone wrong*.
The cost of skipping it is documented in the mesh's own record: entries have been rediscovered
from scratch, over hours, in sessions where the search was skipped because the trail felt
confident. It fires hardest on familiar ground, not unfamiliar ground.
## The structured archive
A second store, structured rather than flat: spaces, pages, revisions, tiers, and full-text
search. Where the operational memory is a note, this is a document with an owner and a
lifecycle.
Content is promoted through tiers — private, then team, then platform — with a librarian agent
owning approval and promotion at the boundary. Proposals to edit are reviewed rather than
applied.
This is where the mesh's **governed** documents live, including the constitution injected into
design sessions ([ADR 0009](../../adr/0009-the-mesh-is-governed-by-a-constitution.md)).
## Why both
The distinction is by lifecycle, not by subject.
| Operational memory | Structured archive |
|---|---|
| Written the moment something is learned | Written deliberately, reviewed |
| Flat, symptom-indexed | Structured, tiered, owned |
| Anyone writes; nothing approves | Promotion is approved |
| Truth is "this happened" | Truth is "this is agreed" |
Collapsing them would cost one of the two properties: either every hard-won note waits for
review, or governed documents can be changed by anyone mid-incident.
## Where this repository sits
This repository is a third thing, and the objection was raised when it was created: a fourth
knowledge system repeats the mistake the split was made to fix.
The answer given was **indexing, not location** — that these documents are indexed into the
knowledge base so that a symptom search returns them alongside everything else. One source,
many surfaces.
**That indexing does not currently exist.** A search for this repository's content returns
nothing. The claim is load-bearing for the decision to separate the repository at all, and
until it is true, this repository is exactly the fourth knowledge system the objection
described. Recorded here rather than in the ledger, because it is a statement about how the
mesh's knowledge actually works today.
## The librarian
A single agent owns the archive's approvals and promotions. Its approval capabilities have at
times not been reachable as tools, which does not affect the operational memory but does mean
promotion stops silently — the store keeps accepting proposals that nothing can approve.
+88
View File
@@ -0,0 +1,88 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0012-agents-are-persistent-employees.md
- adr/0009-the-mesh-is-governed-by-a-constitution.md
---
# Agents and work
The mesh does a large share of its own design and implementation. Agents are how, and the
model they run under is the employee model, not a worker pool.
## An agent is an employee
An agent is a singular named identity with a home node, a workspace on that node, accumulating
memory, and an explicit lifecycle
([ADR 0012](../../adr/0012-agents-are-persistent-employees.md)).
| Property | Meaning |
|---|---|
| Lifecycle state | Active, draining, or retired. Retired agents are kept. |
| Home node | Where its workspace lives. One node per agent. |
| Session cap | How much work it may hold at once. **Concurrency is a property of the agent, not a count of copies.** |
| Kind | Whether it is hirable, or is a node's own agent and exempt from hiring |
The verbs are explicit: an agent is **hired** onto a node, **reassigned** only while idle, and
**retired** by draining first — forcing it is a deliberate act that aborts work in flight.
Surge capacity lives inside the model rather than against it. A template agent is a blueprint
with no life of its own; when a queue grows past a threshold it is cloned into a real agent
with a lifetime, which drains and retires when that expires. A temporary employee is still an
employee.
Because there is a continuing subject, **policy becomes possible**: an agent that violates a
rule can be warned, and a warned agent can be dismissed. A pool cannot be warned.
## Some agents are human
There is one kind of participant. What differs is **modality** — a non-human agent acts through
a spawned session and the record; a human agent acts through a shell, a desktop, or a message.
Both hold identity, both act, both accumulate memory.
The mesh does not currently record modality completely. Which user, on which node, a human
agent acts as is **required by the model and not stored** — an open question carried over from
[ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md).
## Work
Work is expressed as tasks moving through workflows. A workflow names the states a kind of work
passes through and what must be true to leave each one; a task carries its acceptance criteria
and its trail.
Several workflow shapes exist for different sizes of work — a single implementation, a larger
container of related work, and shapes that add analysis or design stages ahead of
implementation.
The area's characteristic defects are **transition** defects rather than logic defects: a task
bouncing between review and implementation because a guard was evaluated on stale state, a
result that cannot be recorded in the same act as the transition it justifies. The workflow
engine's correctness is about atomicity, and that is where it has been wrong.
## Meetings
Some work is decided in a **meeting**: several agents in turns, with distinct roles, over a
template that names the phases.
This is where governance meets execution. The constitution is injected into every eligible
meeting turn — agents do not fetch it, it arrives — and a check phase verifies the meeting's
output against it before the meeting may proceed
([ADR 0009](../../adr/0009-the-mesh-is-governed-by-a-constitution.md)). A named violation
blocks progress.
Meeting turns run on the orchestrator's node regardless of where the participating agents are
pinned. That is a known divergence between the model and its execution, not a design intent.
## What this rests on that is not built
The work domain shares one large schema with several other domains. That is the concrete
instance of a rule stated in [`how-we-build.md`](../../00-GENESIS/how-we-build.md) — *contexts
integrate through the record, never through a shared schema* — being violated by the mesh's
own largest component, and it is the reason work that belongs to one domain keeps having to be
implemented in another.
[ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md) dissolves that arrangement.
Until it does, this is the shape.
@@ -0,0 +1,78 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0001-nodes-communicate-over-a-broker.md
- adr/0008-a-failed-step-fails-the-job.md
---
# Interfaces and observability
How the mesh is reached, and how anyone can tell what it is doing.
## Capabilities are the primary interface
The mesh's primary interface is not a web console. It is a set of **capabilities**, exposed to
a session and callable in language.
A capability is contributed by a module and is available on any node, wherever it actually
runs: local ones directly, remote ones through a stand-in created at startup that forwards over
the broker. The caller does not know the difference, and the credentials never move.
This is the mesh's stated vision made concrete — an agent states an intent and the mesh works
out which node holds the thing. It is also why a capability's **schema** is load-bearing in a
way that is easy to underestimate: a parameter name that collides with the transport's own
reserved names breaks the call, and a validation-library version mismatch has silently dropped
every argument while the call still appeared to succeed.
## The board
A web interface presents the mesh — nodes, modules, pipelines, agents, work. It is a **view**.
Its own guidance is that shared logic belongs in the mesh's library rather than inline in the
board, precisely so the board does not quietly become a second implementation of the mesh's
rules.
## Public exposure
Nodes carrying a public name run a reverse proxy as the sole entry point. A module declares the
names its interfaces answer on, portably, and the proxy's configuration is **generated** from
those declarations rather than written — generated files are marked as such and anything
hand-written beside them is left alone.
Nodes without a public role use a local equivalent with locally-trusted certificates. The
generation step degrades quietly on a node with no proxy, which is intended and is one more
place where "nothing happened" is the correct outcome and looks identical to a failure.
Certificate issuance currently always targets the public authority's production endpoint,
which consumes real quota for every experiment
([`04-ISSUES/004`](../../04-ISSUES/004-certificate-issuance-targets-production/00-report.md)).
## Health
Nodes run checks and report. The mesh's health surface answers whether things are up.
What it does **not** answer is whether they are correct, and that gap is the recurring theme of
this whole system: the deploy path reports transport rather than effect, so absence reads as
success. A check that confirms a service is running does not confirm the service is running the
code that was just deployed, and a node has been left on old code with a version marker that
had already advanced.
## Thoughts
Every node's daemon runs a periodic loop that surfaces observations from that node's own
context. They are stored in the mesh and can inform a session or trigger action.
It is the one part of the mesh that is not request-driven — the mesh noticing things rather
than being asked.
## The honest summary
Observability tells you the mesh is **up**. Establishing that it is **right** currently means
reading the operational record and checking by hand.
That is the gap the lab is designed to close
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)): a place where a
change can be run end to end and a verdict produced, cheaply enough that producing one is
routine.
+90
View File
@@ -0,0 +1,90 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- adr/0002-everything-is-a-module.md
- adr/0010-applications-live-in-their-own-repository.md
- adr/0017-modules-outside-the-core-are-grouped-by-domain.md
---
# The catalogue, and what its shape says
The catalogue holds **124 modules**. Thirty-three belong to the mesh's own domain; the other
ninety-one run *on* the mesh rather than being *of* it
([ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md)).
The count is not the finding. The **shape** is.
## How it is organised today
By namespace, and the namespace records origin rather than purpose:
- **The mesh's own namespace** holds the platform: the runtime and its daemons, the shared
library, delivery, provisioning, configuration synchronisation, knowledge, identity, the
board, developer tooling, and node presentation.
- **A second namespace** holds the work domain — tasks, workflows, agents, meetings — split
across a handful of packages that share one schema.
- **Everything else sits flat at the top level**, one directory per piece of software.
## What the flat level actually contains
Grouped by what they are *for* — a grouping the catalogue itself does not express:
| Purpose | Roughly |
|---|---|
| Data and storage services the mesh provisions against | Relational and document databases, a cache, an object store, a package registry, a time-series store |
| Messaging and identity | A message broker, an identity provider |
| Reachability | A VPN, a firewall, an intrusion filter, an SSH daemon, a resolver, a certificate authority, a reverse proxy, network equipment control |
| Forge and container plumbing | Forge integrations, an image registry, container lifecycle and retention |
| Media libraries | Acquisition, organisation, playback, transcoding, streaming |
| Workstation and desktop | Browser, file manager, monitors, session management, audio, package management, runtime managers |
| Hardware-specific support | Power and firmware control for particular hardware, filesystem management |
| Collaboration and productivity | File sync, office tooling, boards, automation, chat and messaging bridges, mail, analytics, dashboards, home automation, issue trackers and wikis |
| Third-party organisation integrations | Systems belonging to organisations outside the mesh |
Every row is several modules, and **no row is a thing the mesh can see**. Four modules
together constitute "how a node is reachable", and they have no relationship the mesh can
assign, version, reason about or replace as one unit. A change to how the mesh handles
connectivity is made four times.
## What the shape records
**The catalogue's shape records what was installed, not what anything is for.** One module is
the unit of one piece of software, because that is the only granularity the module system
offers.
This is the same failure [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md)
names for the platform core — *boundaries drawn by deployment accident rather than by domain* —
appearing outside it, at four times the scale. The core is being recomposed; the flat level is
addressed in principle by
[ADR 0017](../../adr/0017-modules-outside-the-core-are-grouped-by-domain.md), which
deliberately does not yet settle the domain list.
## Two properties worth keeping
Whatever replaces the shape, two things about it are right.
**Uniformity.** A media server and the mesh's own coordinator are installed, provisioned,
delivered and verified by identical machinery. The mesh's own components hold no privilege —
which is what makes dogfooding structural rather than a discipline, and what makes moving a
module out of the repository safe.
**Placement is already decided.** A standalone application belongs in its own repository
([ADR 0010](../../adr/0010-applications-live-in-their-own-repository.md)), and reviewers reject
it in the monorepo. The catalogue's flat level is not a dumping ground by policy; it is one by
history.
## Known inconsistencies in the catalogue itself
Recorded because a reader will meet them:
- A documented requirement that every capability-exposing module declare the core runtime as a
dependency is met by **zero** modules.
- A firewall-scoping key is declared by five manifests and read by none
([`04-ISSUES/003`](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)).
- A connections block in the manifest is metadata: it describes a module's reachability and
wires nothing.
- At least one module deliberately runs outside the standard per-module supervision, for
reasons recorded in the operational memory. The standard path is not universal.
+30
View File
@@ -0,0 +1,30 @@
# 02-DESIGN / 00-as-is
The mesh as it stands. These documents describe what runs, including the parts nobody would
choose again — an as-is layer that only records the good decisions is a brochure.
They are written from the implementation and from the operational record, not from intent.
Where the two disagree, the implementation wins and the disagreement is stated.
| Document | Covers |
|---|---|
| [`00-overview.md`](00-overview.md) | The whole in one pass — what a node is, what a module is, how work reaches it |
| [`01-mesh-and-transport.md`](01-mesh-and-transport.md) | The mesh database, the broker, discovery, and how a call reaches another node |
| [`02-modules-and-manifests.md`](02-modules-and-manifests.md) | The module, the manifest, and features as the unit of work |
| [`03-provisioning.md`](03-provisioning.md) | Declared requirements, provisioners, credentials, and cross-node grants |
| [`04-delivery.md`](04-delivery.md) | Push to running: the three silos, levels, and what a green pipeline proves |
| [`05-runtime-and-installation.md`](05-runtime-and-installation.md) | The node runtime, its modes, and how a node comes into being |
| [`06-configuration-and-secrets.md`](06-configuration-and-secrets.md) | Managed files, value resolution, and where secrets live |
| [`07-knowledge.md`](07-knowledge.md) | The two knowledge stores, and what each is for |
| [`08-agents-and-work.md`](08-agents-and-work.md) | Agents as employees, tasks, workflows, and the meeting model |
| [`09-interfaces-and-observability.md`](09-interfaces-and-observability.md) | How the mesh is reached and watched — tools, board, proxy, health, thoughts |
| [`10-module-catalogue.md`](10-module-catalogue.md) | The catalogue's shape, and what its shape says |
## What these documents are not
They are not a runbook. Operational procedure — how to fix one occurrence of something — lives
in the knowledge base, which is indexed on symptoms and is the right place to search when
something is broken.
They are not exhaustive. A subsystem is described to the depth at which its **design** is
visible; below that is code.
@@ -1,6 +1,14 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [adr/0015-mesh-brokers-nodes-host-agents-think.md]
---
# Work breakdown — the decomposition
How ADR 0001 gets built, in what order, and where a human must look.
How [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md) gets built, in what order, and where a human must look.
Ordering is not preference. Each phase removes a constraint the next one needs gone.
@@ -74,7 +82,7 @@ The decomposition is impossible while a feature is a singleton per module.
| # | task | done when |
|---|---|---|
| 1.1 | ADR 0002 — named features, per-node opt-in | accepted |
| 1.1 | Decision record — named features, per-node opt-in (next free number) | accepted |
| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build |
| 1.3 | Selection: `always` / flavor-selected / `optional` | a node installs a subset; artifacts stay flavor-blind |
| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed |
@@ -1,3 +1,11 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [adr/0016-a-lab-node-is-a-virtual-machine.md]
---
# End-to-end testing
**What is under test is a module.** The mesh is the harness.
@@ -350,6 +358,6 @@ reads credentials from host paths, because there is nowhere else to put a mesh.
is, a workstation stops being collateral.
**The supervision question stops gating anything.**
[`01-RESEARCH/003-service-supervision`](../01-RESEARCH/003-service-supervision/analysis.md)
[`01-RESEARCH/003-service-supervision`](../../01-RESEARCH/003-service-supervision/analysis.md)
remains open on its own merits — and once this exists, its options are cheap to try rather
than expensive to argue about.
+22
View File
@@ -0,0 +1,22 @@
# 02-DESIGN / 01-to-be
The mesh being built toward. Every statement here traces to a record in
[`adr/`](../../adr/); nothing arrives by drafting.
A document here describes an intention. What currently runs is in
[`00-as-is/`](../00-as-is/), and the two are never merged — when something ships, the as-is
document is written and this one's status becomes `implemented`.
| Document | Covers | Rests on |
|---|---|---|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../adr/0016-a-lab-node-is-a-virtual-machine.md) |
## Not yet written
- **The eight bounded contexts.** [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md)
decides the decomposition; the per-context specifications do not exist yet. The work
breakdown says in what order they are needed.
- **Domain grouping outside the core.** [ADR 0017](../../adr/0017-modules-outside-the-core-are-grouped-by-domain.md)
settles the principle and explicitly does not settle the domain list. That is a research
effort, not a design document, until it concludes.
+46 -5
View File
@@ -2,9 +2,50 @@
The authoritative specification. Implementation is built against what is written here.
A document enters this folder only after the decision behind it is recorded in
[`adr/`](../adr/) and the research that produced it is marked `GRADUATED`.
## Two layers
Empty for now: the decomposition in ADR 0001 is decided but not yet specified. The first
entries will be the per-context designs — `hal/mesh` brokering, the `hal/stream` record,
and the feature model that `hal/delivery` owns.
| Folder | What it is |
|---|---|
| [`00-as-is/`](00-as-is/) | **The mesh that exists today.** Shipped behaviour, described as it is — including behaviour nobody would choose again. |
| [`01-to-be/`](01-to-be/) | **The mesh being built toward.** Every statement traceable to a record in [`adr/`](../adr/). |
They are never mixed. A statement about the future does not belong in an as-is document, and
an as-is document is never edited to describe an intention.
When a to-be design ships, it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand — one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning, which is
the expensive half.
## Frontmatter
Every design document (not the READMEs) carries:
```yaml
---
layer: as-is | to-be
status: designed | in-progress | implemented | abandoned
code: [] # owning code repo(s), from 00-GENESIS/repos.md
updated: YYYY-MM-DD # date of the last status change, not of text edits
decisions: [] # adr/ records this document rests on
---
```
For an as-is document, `status: implemented` is the normal state — it describes something that
runs — and `code:` names where that implementation lives.
Status changes when **implementation state** changes, never because design text was edited. An
`implemented` claim must be defensible from the owning repository's main branch, not from
intent. If it cannot be checked, it is `in-progress`.
Cross-cutting views are generated from this frontmatter by the `hal-status` skill and never
written to disk.
## What belongs here
Functional analysis, architectural description, and specification — **prose and diagrams
only, no code**. A manifest field may be named; a manifest may not be pasted. A document
enters the to-be layer only after the decision behind it is recorded in [`adr/`](../adr/)
and the research that produced it is closed.
Subfolders are encouraged where a layer grows enough to need them.