The numbering is the flow: decisions are 02, design is 03

papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
This commit is contained in:
2026-08-23 18:05:11 +02:00
parent f05e4a0dce
commit c0b35652d0
72 changed files with 217 additions and 204 deletions
+98
View File
@@ -0,0 +1,98 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh as it stands
A set of machines, each running the same runtime, each loading only the parts of the catalogue
it has been assigned. They hold no shared filesystem and make no direct connections to one
another. What makes them a mesh is a database that knows what should run where, and a message
broker that carries everything between them.
## Three nouns
**A node** is a machine that runs the runtime. Nodes differ in what they are assigned and in
what they can reach — some carry a public name, some sit behind a household connection with no
inbound route at all — and the mesh is designed so that difference stays a property rather than
becoming a special case. A node holds no authoritative state: everything it needs is derived
onto it and can be regenerated.
**A module** is a directory with a manifest, and it is the only unit the mesh installs. A
containerised service is a module. A set of capabilities with no service behind them is a
module. A bare marker whose whole content is that a node has it is a module. The mesh's own
components are modules on exactly the same terms as everything else it carries
([ADR 0002](../../02-DECISIONS/0002-everything-is-a-module.md)).
**An agent** is a participant. Some agents are human. What differs is modality — how the agent
acts — and not category: both hold identity, both act, both accumulate memory
([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)).
## Where truth lives
The repository defines **what exists**: the modules, what each declares, how each is built.
The mesh database defines **what runs where**: which node is assigned which module, at which
selection, with which overrides, plus the settings every node reads. No node-to-module mapping
is ever committed ([ADR 0003](../../02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md)).
Everything on a node's disk is **derived** from those two, and is regenerated rather than
edited ([ADR 0004](../../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)). A node that
loses its database keeps running from a local cache, which is deliberate and has the obvious
cost: the cache carries no indication of its own age.
## How anything moves
Nothing dials a node. Every node dials the broker outbound, owns an exchange named for itself,
and consumes from its own request queue
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)). Three message shapes carry
everything: requests expecting a reply, commands instructing that a stage of work be done, and
events stating that something happened.
A capability that lives on another node is reached the same way a local one is. At startup a
node asks its peers what they host and creates a local stand-in for each remote capability, so
the caller does not know or care where the work happens. Credentials never travel: the call
goes to where the capability is.
## How change reaches a node
A push to the forge is the only trigger. What follows is three silos with deliberately
different cardinality: compile once, package and upload once, then install-configure-start-
verify **on every assigned node**
([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)). What travels between
build and node is a self-contained build output, so a deploy is extract-and-run and touches no
network ([ADR 0013](../../02-DECISIONS/0013-an-artifact-is-build-output.md)).
Modules are resolved into dependency levels and a level completes before the next begins, so a
module always builds against its dependencies as they were just published.
## What the mesh does for a module
A module declares what it **provides** and what it **requires**. The mesh satisfies the
requirement: it creates the resource, generates the credential, records the grant, and writes
the values where the module will read them. The module never learns which node its database
lives on, and nobody ever writes a credential by hand
([ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)).
This is the property the mesh's whole shape rests on, and it is why provisioning is treated as
a core concern rather than as plumbing.
## The shape of its failures
Worth stating in an overview, because it is the most consistent thing about the system: the
mesh's expensive faults are almost never crashes. They are operations that reported success
and did nothing — a download that half-completed, a hook that was never called because it was
named for a feature the module does not declare, a stage that reported it had dispatched a
message rather than that the effect happened, a package that 404ed from every mirror while the
job went green.
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) is the response, and it is applied
instance by instance rather than enforced by a mechanism. New instances are still being found.
That is an as-is fact, not a criticism: it is the single most useful thing to know about this
system before changing it.
+113
View File
@@ -0,0 +1,113 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh and its transport
Two things make a set of machines into a mesh: a database that holds every binding, and a
broker that carries every message. Neither is optional and neither is replaceable at present.
## The mesh database
One relational database holds the bindings. Its content divides cleanly:
| Holds | Describes |
|---|---|
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
| Grants | Which consumer holds which resource from which provider, with the credential |
The runtime loads this at startup. If the database cannot be reached it falls back to a local
cache and continues.
**The repository contains none of this.** It defines what exists; the database defines what
runs where. This is what makes the repository node-agnostic, and it is the property that lets
anything about the mesh be published at all.
### What the cache costs
Running from cache is the difference between a node that survives a database outage and one
that stops. It is the right trade and it has a cost that is worth naming: a node running from
cache looks identical to a node running from the database. There is no age on the cache and
nothing reports divergence, so a node can be running yesterday's assignment set indefinitely
without any signal that it is.
### Where the settings are
A mesh-level setting is a value every node needs and no node owns — the broker's location, the
forge's, the registry's. These live in the database rather than in any node's configuration,
so a node learns where the broker is from the mesh rather than from a file, and moving the
broker is a database change rather than a fleet-wide edit.
The circularity is real: a node must reach the database to learn where the broker is, and the
database is itself a module the mesh provisions. It is resolved by the initialisation script
that stands up the first node, which is the reason such a script exists separately from
everything else.
## The broker
Every node connects **outbound** to a single broker. Nothing ever connects to a node.
Each node declares a topic exchange named for itself and consumes from its own request queue.
A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and
the events they emit.
Three message shapes, and only three:
- **Requests** expect a reply. This is how a capability on another node is called.
- **Commands** instruct that a stage of work be done. They are addressed by what is to be done
and consumed by whichever node is meant to do it.
- **Events** state that something happened, addressed to nobody. Anything interested subscribes.
### Consequences the design accepts
A node behind a household connection with no inbound route participates exactly as a publicly
named one does. This is the property the transport was chosen for.
A call to a node that is down **waits** rather than failing. Usually this is what is wanted. It
is also how the mesh's most confusing stalls happen: a command queued for a node that never
returns is a stall with no error anywhere, and the pipeline has produced this shape more than
once — an empty pipeline that never completes blocks every pipeline queued behind it.
Two consumers accidentally sharing one queue silently split the traffic between them, each
receiving half of what it expects. This has happened between a module's daemon and its
capability server.
The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker
down.
## Reaching a capability on another node
A node hosts some capabilities and can reach the rest.
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a
local stand-in that forwards over the broker. For anything hosted both locally and elsewhere
it wraps the local one so a caller can name a target.
The effect is that a caller does not know where a capability runs. The important half is what
does **not** move: the work happens where the capability is, so its credentials never leave
that host. A remote call transports a request and a reply, never a secret.
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent
at that moment is simply not discovered, and the node runs without that capability until it
restarts. Nothing re-discovers on a schedule.
## Names and reachability
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the
underlying network provides. A node's mesh name is its overlay address; its public name, if it
has one, is a separate fact used by things outside the mesh.
Two lessons are embedded in that separation, both learned the expensive way. A name resolved
by local multicast discovery introduces a delay and a failure mode that appears on one node and
not others, so mesh names are not multicast names. And a node must not pin its own public name
locally: the duplicate record breaks resolution for everything else that needs it.
@@ -0,0 +1,113 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0006-schema-changes-are-numbered-migrations.md
- 02-DECISIONS/0007-no-npm-workspace.md
---
# Modules, manifests and features
Everything the mesh installs is a module: a directory with a manifest. There is no second
mechanism.
## What a manifest declares
| Declares | Meaning |
|---|---|
| Identity | Name and version. **Version is owned by the builder** — a hand-edited version is a defect, and reviewers revert it. |
| Environment | Every variable the module reads, with how each is produced: a static default, a generated secret, a value pulled from the node's own record, or a template composed from the others. A variable not declared here is invisible to the mesh and will not be generated, injected or audited. |
| What it provides | The resource type this module can provision for others, and the network on which it is reachable. |
| What it requires | Resources it needs from other modules, and the mapping from each resource's connection fields onto its own environment variables. |
| Service shape | The primary container, and the data directories that must exist with the right ownership before it starts. |
| Exposure | The public names this module's interfaces answer on, declared portably so the reverse proxy configuration can be generated rather than written. |
| Images | Container images this module builds, so the pipeline builds and publishes them before publishing the module. |
## Features are the unit of work
A module is not the unit the pipeline addresses. A **feature** is.
A feature is a kind of content a module can carry: a service, a set of capabilities, a
long-running process, managed configuration files, migrations, firewall rules, an installable
application. One module can carry several.
Features are **detected from directory contents**, not declared. A module with a capabilities
directory has that feature; a module with a daemon directory has that one. An explicit
declaration was supported and is now discouraged, because a declared list and the directory it
describes drift, and the directory is the one that is true.
Every pipeline command and event names a feature. There is no per-module build.
### What detection costs
Detection makes the manifest shorter and the truth singular, and it makes the directory
structure load-bearing in a way that is not obvious from reading a manifest. Renaming a
directory changes what a module *is*, silently. The recurring failure is a hook named for a
feature the module does not carry: it is skipped without complaint, and the change it was
supposed to make simply never happens.
An unknown key in a manifest is likewise accepted in silence — which is how a firewall rule
can appear to restrict a port and restrict nothing (see
[`04-ISSUES/003`](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)).
## Kinds of module
The kinds are not a type system — they are what the detected features add up to.
- **A service module** carries a container definition. It gets a runtime directory, generated
environment, data directories, and is started under supervision.
- **A capability module** carries capabilities and no service. It contributes what a node can
do, locally and to its peers.
- **A flag module** carries nothing but a manifest. Its presence in a node's assignment is the
entire content: it gates behaviour elsewhere.
- **Combinations** are ordinary. A database module is a service *and* a capability provider
*and* a provisioner.
## Selections
A module can ship variants of the same feature and a node takes the one that fits it — a build
for one accelerator or another, a configuration for a public node or a private one. The
artifact stays selection-blind; the choice is a property of the assignment, held in the mesh
database.
This is one of the areas where behaviour has repeatedly diverged from intent, in both
directions: selection files that were never packaged into the artifact at all, and a stale
staged override on a node that silently won over the newly selected one. Both classes are
recorded in the knowledge base; both presented as "the change did not apply" with no error.
## Dependencies between modules
Modules depend on each other, above all on the shared library they all build against. There is
**no workspace** ([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)): each module is a standalone
package consuming published dependencies, including the mesh's own.
The pipeline resolves modules into dependency **levels** and completes a level before starting
the next, so a module always builds against its dependencies as just published.
The cost is a publish-and-consume round trip for every cross-package change, and the absence of
any repository-wide build. One thing that assumed a repository-wide build has stayed broken
since (see
[`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)).
## Persistent state
A module that owns state owns its migrations: numbered, written in the module's own language,
compiled with it, frozen once they have run anywhere, and idempotent so that re-running is safe
([ADR 0006](../../02-DECISIONS/0006-schema-changes-are-numbered-migrations.md)).
Two kinds exist and the distinction matters: migrations against the module's **own** local
state, and migrations against a **provisioned** resource, which run on the node that consumes
the resource rather than on the node that built the module.
## A documented rule with no enforcement
Every module exposing capabilities is documented as required to declare the mesh's core runtime
as a dependency. **Zero of the catalogue's modules do.**
This is recorded here rather than quietly corrected, because it is the clearest instance of the
rule this repository states about itself: a rule whose enforcement does not exist is
indistinguishable from a wrong one, and costs more, because people believe it. Whether the rule
or the catalogue is wrong has not been decided.
+88
View File
@@ -0,0 +1,88 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md
- 02-DECISIONS/0004-managed-files-are-generated-never-edited.md
---
# Provisioning
A module states what it needs. The mesh makes it exist, generates the credential, records the
grant, and puts the values where the module will read them. Nobody writes a credential and
nobody writes a topology.
This is the mesh's core concern rather than its plumbing — the property everything else is
built on.
## The declaration
A **provider** declares the resource type it can create and the network on which that resource
is reachable from a container.
A **consumer** declares, per requirement: the provider module, the resource type, optionally a
name and a target node, and a mapping from the resource's connection fields onto its own
environment variables.
The consumer must also declare those variables as empty in its environment section. A mapped
field with no declared variable is dropped — the value is produced and then discarded, because
the generator only emits variables the manifest knows about.
## What happens
Inside the pipeline the sequence is orchestrated by the coordinator and is not optional:
1. The provisioner creates the resource and its credential, records the grant, and writes the
mapped values as database overrides.
2. The synchroniser regenerates the module's environment from those overrides.
3. Only then is the service started.
Each step waits for the previous one's event. A module that declares no requirements skips the
first step entirely — which is correct, and means the absence of provisioning is
indistinguishable from provisioning that did not run.
Outside the pipeline the same two steps exist as direct operations, for debugging. They are
not the normal path.
## What can be provisioned
Providers exist for relational databases of two kinds, an object store, a cache, message-broker
virtual hosts and access, an identity provider's clients, and a download category. Each returns
the connection fields a consumer maps from — host, port, user, password, and whatever else the
resource type implies.
A provider with no registered provisioner still records a grant, with an empty connection. This
is deliberate and easy to misread: the grant exists, so the requirement looks satisfied, and
nothing was created.
## Cross-node grants
A requirement may name the node whose provider should satisfy it. The grant records consumer
node and provider node separately, so a module on one node holding a database on another is
the ordinary case rather than a special one.
For a containerised consumer, the provisioner returns the provider's routable name rather than
a loopback address — the value has to be correct from **inside** a container on another
machine.
## Where this hurts
**Rotation has no fan-out.** A resource whose credential is shared by several consumers can be
rotated by provisioning a new one, and the peers holding the old credential are not told. This
has locked the mesh out of its own broker, and has caused a node's adoption to rotate a live
shared password without informing anything that held it. Granting is easy; regranting is not
modelled.
**A frozen password outlives its generation.** A generated password is written once. If the
resource's persistent data directory already exists from an earlier initialisation, the stored
credential and the generated one diverge, and the symptom is an authentication failure that
looks like a configuration error.
**Migrations against a provisioned resource run on the consumer's node**, not the build node,
and read the deployed artifact rather than the source tree. Both facts were wrong in the
implementation for a period during which new provisioning migrations silently never ran.
**A grant is not a check.** The record says a resource was provisioned. Nothing verifies it
still exists, still has the recorded credential, or is reachable from where the consumer runs.
+100
View File
@@ -0,0 +1,100 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md
- 02-DECISIONS/0013-an-artifact-is-build-output.md
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
---
# Delivery — from a push to a running node
One trigger, three silos, and a fan-out point that is the most consequential boundary in the
mesh.
## The trigger
A push to the forge. The forge calls a webhook; the receiving module verifies its signature
and emits an event. The coordinator resolves which modules the pushed commits affect, orders
them into dependency levels, and creates a pipeline per level.
There is no second path. A manual trigger exists for a module the push detection missed, and
using it routinely is a sign the detection is wrong rather than a workflow.
Detection is the pipeline's most fragile input. It has failed for reasons that have nothing to
do with the change: a webhook truncating its commit list on a large merge, a forge address
whose port broke the module-path matching. The characteristic outcome is the bad one — **a
merge that created no pipeline, and nothing said so**.
## Three silos
Cardinality is the whole point, and the three differ
([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)):
| Silo | Runs | Where | Does |
|---|---|---|---|
| **build** | once per module feature | the build node | compile and bundle into a self-contained output |
| **publish** | once per module feature | the build node | package that output and upload it; a package-registry feature publishes here |
| **deploy** | once per module feature **per node** | every assigned node | install, configure, start, verify |
Commands and events are addressed per feature, not per module.
Build hands over a **staged tree**, not a package. Packaging belongs to publish, so a failed
upload retries by re-packaging rather than by re-sending something stale, and build never needs
to know how each module composes its artifact. The handover is a tree in a known location,
because the build's own working directory is reference-counted and may be gone by the time a
later stage runs.
## The artifact
The artifact is **build output** — compiled and bundled with its dependency graph inlined —
never a filtered copy of source ([ADR 0013](../../02-DECISIONS/0013-an-artifact-is-build-output.md)).
A deploy is extract-and-run and touches no network.
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after deploy — migrations reading a source layout, provisioning
scripts reading a source layout, selection files never packaged at all.
## The fan-out, and the build node
After publish, work fans out to every assigned node. The build node is the only node that has
already passed through two silos when this happens, and that asymmetry has its own defect
class: anything advancing a node's stage must account for **both** pre-fan-out stages. Code
that knew only about the first parked the build node forever while every other node deployed
cleanly — and the recovery sweeper, which knew the same subset, reported nothing to recover.
A recovery mechanism that knows less than the thing it guards is worse than none, because it
reports success over a stall it cannot see.
## Levels
A level completes before the next begins, so a module builds against its dependencies as they
were just published. The shared library is at level zero, which means anything that breaks it
breaks the first module of every cascade.
## What green proves
**A green pipeline proves transport, not effect.** The stages report that a message was
dispatched and accepted, which is not the same as the thing being running, correct, or present.
This is the mesh's most consistent failure shape and it is not incidental to the design — it is
what the stage reporting currently measures. Documented instances include a service reported
started when the container command merely returned, an image pull failure that did not fail the
deploy, a package install that 404ed from every mirror while the job went green (see
[`04-ISSUES/001`](../../04-ISSUES/001-failed-package-install-reports-success/00-report.md)),
and a node left on old code after a failed artifact download with a version marker that had
already advanced.
A verify stage exists to close this gap. It has been built and, for a period, was never
scheduled, because the coordinator's stage list did not include it.
## What is not covered
The end-to-end harness for this pipeline has not built since the workspace was removed, and
nothing reports that nothing runs it
([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). The
pipeline's coverage is currently assumed rather than checked — which is the same class of claim
this repository exists to make people stop making.
@@ -0,0 +1,107 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0011-the-installer-owns-linking.md
---
# The node runtime, and how a node comes into being
Every node runs the same runtime. What differs is its assignment.
## Two modes, coexisting
The runtime runs in two modes at once on any node that needs both.
**Daemon** mode is a headless consumer: it connects to the broker, consumes the node's request
queue, routes each request to a local capability, and emits the node's lifecycle events. This
is what makes a node a participant — it is reachable whether or not anyone is logged in.
**Interactive** mode exposes the node's capabilities to a session on that machine over a local
protocol. Its capability surface is larger, because it includes stand-ins for every capability
discovered on peers.
The two are the same code with the same catalogue. A capability is written once and is
available to both.
## Anatomy naming, and what it obscures
The runtime's components are named after brain anatomy: an entry point that bootstraps, a
headless listener, an interactive surface, an installer daemon, and a provisioner daemon.
This is the mesh's most-cited naming problem and it belongs in the as-is layer because it is
what a reader will actually encounter. The names are evocative and describe nothing: the most
suggestive word in the system names the node runtime, and the component whose manifest says
"mesh messaging" is documented elsewhere as the interactive runtime. Anatomy makes attractive
names and poor boundaries.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) replaces this with names
taken from what each part owns. Until then, this is the vocabulary in the code.
## Starting a module
For each assigned module the runtime, at startup:
1. Reads the manifest, if there is one — a flag module has nothing to read.
2. Skips a module's capabilities if a variable they require is unset. This is quiet by design
and hard to distinguish from a module that has no capabilities.
3. Loads the capabilities the module carries.
4. For a service module, ensures it is installed and running.
Installation is idempotent and does the unglamorous work: ensure the runtime directory,
reconcile the link from the catalogue's definition into it, generate the environment, run any
outstanding local migrations, create data directories with the right ownership, then start the
service under supervision.
**The installer is the only thing that creates a link** ([ADR
0011](../../02-DECISIONS/0011-the-installer-owns-linking.md)). It reconciles rather than assumes: a
missing link is created, a stale one repointed, and a real file found where a link belongs is
adopted into the node's override area and replaced. Nothing else — not a hook, not a fix, not a
person debugging — creates one.
That is the as-is. The intent is to remove linking altogether and derive a real file instead,
which the reconciliation machinery already makes possible
([ADR 0018](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), proposed). What is described above
is what runs today.
## Supervision
Services run under the host's init system via a templated unit, one instance per module. It is
a thin layer: the unit starts and stops a container group.
Whether the mesh keeps this, drops the per-module layer, containerises the daemons, or writes
its own supervisor is **open** — costed in research effort 003 and deliberately undecided. It
no longer gates anything.
Two failures worth knowing about, both in the shape of "the change did not apply". A
per-instance copy of the unit template shadows the template, so edits to the template do
nothing. And a session-scoped one-shot job loses the environment it was given, because the
import is one-time and not persisted.
## How a node comes into being
Three bootstrap scripts, and which one runs depends on the situation:
- **First node.** Nothing exists yet, so the script stands up the database the rest of the
mesh reads from, publishes the catalogue, and starts the mesh. This resolves the
circularity of a mesh whose source of truth is itself a provisioned module.
- **Joining.** The node registers, takes its assignment from the database, and syncs.
- **Rescue.** A node that cannot reach the mesh is brought back far enough to.
Bringing a node into being is therefore a **database operation with a script attached**, not a
checkout. There is no per-node content in the repository to copy.
Adoption of a pre-existing machine's configuration was the original path and is now a legacy
one, explicitly out of scope for the lab
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)).
## Node identity
Each node carries an identity text in its own record, which the daemon writes onto the node at
startup so that a session on that machine knows which node it is on and how that node
presents itself. It carries identity only; shared rules live separately.
Like everything else derived onto a node, it is generated and not edited there.
@@ -0,0 +1,86 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0004-managed-files-are-generated-never-edited.md
- 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md
---
# Configuration and secrets
Every value a module reads at runtime comes from a file on a node's disk. Every one of those
files is **generated**.
## The rule
A managed file is derived from the mesh database. A synchroniser rewrites it when the values
behind it change. The write path is the mesh operation that owns the value; the file is an
output ([ADR 0004](../../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)).
An edit to a managed file survives until the next synchronisation and is then overwritten
silently, taking whatever it was fixing with it — bringing back the bug the edit had removed,
with a delay, and with no error to connect the two events.
There is a way to ask whether a given file is managed. That question has to be asked, because
the answer is not visible from the file.
## What is managed
Generated environment files for each module, service definitions in each module's runtime
directory, managed configuration files a module declares, and node-level settings. The list is
not a category — it is whatever a synchroniser claims, which is why the question is asked of
the tooling rather than answered from a rule.
## How a value is decided
Values resolve by precedence, highest first:
1. **A database override** — set deliberately, or written by a provisioner.
2. **The value already in the generated file** — preserved for anything without an override.
3. **A generated value** — a random secret, a template composed from other variables, or a
value pulled from the node's own record.
4. **The manifest default.**
Two consequences follow, and both are subtle enough to have caused confusion.
The second rule is what keeps a generated password **stable** across regenerations. It is not
an oversight; without it every regeneration would issue a new secret and break whatever holds
the old one.
The same rule means that **changing a manifest's default does not change anything already
using it**. The existing file's value wins. The new default reaches only installations that
never had one.
Removing a declaration is worse than changing it: the old override row and the file it produced
are both left behind. Configuration is additive in practice, whatever the manifest says.
## Secrets
Generated secrets are produced by the mesh, never authored. Provisioned credentials arrive as
database overrides written by the provisioner and are marked as such, so they can be
distinguished from a deliberate override and cleaned up when the grant is removed
([ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)).
Nothing in the repository contains a credential. The repository has no per-node content at all,
which is what makes that guarantee structural rather than a matter of care.
Two known weaknesses, both recorded rather than resolved:
**Generated environment files were world-readable.** The mode passed at write time only applies
when the file is created, so regeneration left the previous mode in place. The class of bug is
worth remembering beyond the instance: a permission set at creation is not a permission
maintained.
**Rotation is not a mesh operation.** Secrets can be generated and granted; there is no
mechanism that rotates one and informs everything holding it. Where a rotation has been done,
it has been done by hand, and doing it wrong has taken services down.
## Node-level and mesh-level values
A node-level value applies to everything on one node. A mesh-level value applies everywhere and
is read by every node — the broker's location is the canonical example, and a wrong one is how
a mesh fails to form.
Neither is a file that anyone edits. Both are database records that produce files.
+78
View File
@@ -0,0 +1,78 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions: []
---
# Knowledge
The mesh keeps two knowledge stores. They are not redundant, and knowing which is which is the
difference between finding an answer in one search and rediscovering it over several hours.
## The operational memory
A store of operational notes, written and read by whoever — human or agent — is working. Each
note is a slug and a body: how something works, what went wrong, what the fix was, what
assumption turned out to be false.
It is indexed on **symptoms**. The entry someone needs is usually titled after the error they
are staring at, which is why the standing instruction is to search the literal error text
before forming a hypothesis rather than after one fails.
Its content is overwhelmingly the record of previous debugging: a large body of
troubleshooting entries, module conventions, and standing notes about work that is open. It is
the mesh's institutional memory of *what has already gone wrong*.
The cost of skipping it is documented in the mesh's own record: entries have been rediscovered
from scratch, over hours, in sessions where the search was skipped because the trail felt
confident. It fires hardest on familiar ground, not unfamiliar ground.
## The structured archive
A second store, structured rather than flat: spaces, pages, revisions, tiers, and full-text
search. Where the operational memory is a note, this is a document with an owner and a
lifecycle.
Content is promoted through tiers — private, then team, then platform — with a librarian agent
owning approval and promotion at the boundary. Proposals to edit are reviewed rather than
applied.
This is where the mesh's **governed** documents live, including the constitution injected into
design sessions ([ADR 0009](../../02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md)).
## Why both
The distinction is by lifecycle, not by subject.
| Operational memory | Structured archive |
|---|---|
| Written the moment something is learned | Written deliberately, reviewed |
| Flat, symptom-indexed | Structured, tiered, owned |
| Anyone writes; nothing approves | Promotion is approved |
| Truth is "this happened" | Truth is "this is agreed" |
Collapsing them would cost one of the two properties: either every hard-won note waits for
review, or governed documents can be changed by anyone mid-incident.
## Where this repository sits
This repository is a third thing, and the objection was raised when it was created: a fourth
knowledge system repeats the mistake the split was made to fix.
The answer given was **indexing, not location** — that these documents are indexed into the
knowledge base so that a symptom search returns them alongside everything else. One source,
many surfaces.
**That indexing does not currently exist.** A search for this repository's content returns
nothing. The claim is load-bearing for the decision to separate the repository at all, and
until it is true, this repository is exactly the fourth knowledge system the objection
described. Recorded here rather than in the ledger, because it is a statement about how the
mesh's knowledge actually works today.
## The librarian
A single agent owns the archive's approvals and promotions. Its approval capabilities have at
times not been reachable as tools, which does not affect the operational memory but does mean
promotion stops silently — the store keeps accepting proposals that nothing can approve.
+88
View File
@@ -0,0 +1,88 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0012-agents-are-persistent-employees.md
- 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md
---
# Agents and work
The mesh does a large share of its own design and implementation. Agents are how, and the
model they run under is the employee model, not a worker pool.
## An agent is an employee
An agent is a singular named identity with a home node, a workspace on that node, accumulating
memory, and an explicit lifecycle
([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)).
| Property | Meaning |
|---|---|
| Lifecycle state | Active, draining, or retired. Retired agents are kept. |
| Home node | Where its workspace lives. One node per agent. |
| Session cap | How much work it may hold at once. **Concurrency is a property of the agent, not a count of copies.** |
| Kind | Whether it is hirable, or is a node's own agent and exempt from hiring |
The verbs are explicit: an agent is **hired** onto a node, **reassigned** only while idle, and
**retired** by draining first — forcing it is a deliberate act that aborts work in flight.
Surge capacity lives inside the model rather than against it. A template agent is a blueprint
with no life of its own; when a queue grows past a threshold it is cloned into a real agent
with a lifetime, which drains and retires when that expires. A temporary employee is still an
employee.
Because there is a continuing subject, **policy becomes possible**: an agent that violates a
rule can be warned, and a warned agent can be dismissed. A pool cannot be warned.
## Some agents are human
There is one kind of participant. What differs is **modality** — a non-human agent acts through
a spawned session and the record; a human agent acts through a shell, a desktop, or a message.
Both hold identity, both act, both accumulate memory.
The mesh does not currently record modality completely. Which user, on which node, a human
agent acts as is **required by the model and not stored** — an open question carried over from
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md).
## Work
Work is expressed as tasks moving through workflows. A workflow names the states a kind of work
passes through and what must be true to leave each one; a task carries its acceptance criteria
and its trail.
Several workflow shapes exist for different sizes of work — a single implementation, a larger
container of related work, and shapes that add analysis or design stages ahead of
implementation.
The area's characteristic defects are **transition** defects rather than logic defects: a task
bouncing between review and implementation because a guard was evaluated on stale state, a
result that cannot be recorded in the same act as the transition it justifies. The workflow
engine's correctness is about atomicity, and that is where it has been wrong.
## Meetings
Some work is decided in a **meeting**: several agents in turns, with distinct roles, over a
template that names the phases.
This is where governance meets execution. The constitution is injected into every eligible
meeting turn — agents do not fetch it, it arrives — and a check phase verifies the meeting's
output against it before the meeting may proceed
([ADR 0009](../../02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md)). A named violation
blocks progress.
Meeting turns run on the orchestrator's node regardless of where the participating agents are
pinned. That is a known divergence between the model and its execution, not a design intent.
## What this rests on that is not built
The work domain shares one large schema with several other domains. That is the concrete
instance of a rule stated in [`how-we-build.md`](../../00-META/how-we-build.md) — *contexts
integrate through the record, never through a shared schema* — being violated by the mesh's
own largest component, and it is the reason work that belongs to one domain keeps having to be
implemented in another.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) dissolves that arrangement.
Until it does, this is the shape.
@@ -0,0 +1,78 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
---
# Interfaces and observability
How the mesh is reached, and how anyone can tell what it is doing.
## Capabilities are the primary interface
The mesh's primary interface is not a web console. It is a set of **capabilities**, exposed to
a session and callable in language.
A capability is contributed by a module and is available on any node, wherever it actually
runs: local ones directly, remote ones through a stand-in created at startup that forwards over
the broker. The caller does not know the difference, and the credentials never move.
This is the mesh's stated vision made concrete — an agent states an intent and the mesh works
out which node holds the thing. It is also why a capability's **schema** is load-bearing in a
way that is easy to underestimate: a parameter name that collides with the transport's own
reserved names breaks the call, and a validation-library version mismatch has silently dropped
every argument while the call still appeared to succeed.
## The board
A web interface presents the mesh — nodes, modules, pipelines, agents, work. It is a **view**.
Its own guidance is that shared logic belongs in the mesh's library rather than inline in the
board, precisely so the board does not quietly become a second implementation of the mesh's
rules.
## Public exposure
Nodes carrying a public name run a reverse proxy as the sole entry point. A module declares the
names its interfaces answer on, portably, and the proxy's configuration is **generated** from
those declarations rather than written — generated files are marked as such and anything
hand-written beside them is left alone.
Nodes without a public role use a local equivalent with locally-trusted certificates. The
generation step degrades quietly on a node with no proxy, which is intended and is one more
place where "nothing happened" is the correct outcome and looks identical to a failure.
Certificate issuance currently always targets the public authority's production endpoint,
which consumes real quota for every experiment
([`04-ISSUES/004`](../../04-ISSUES/004-certificate-issuance-targets-production/00-report.md)).
## Health
Nodes run checks and report. The mesh's health surface answers whether things are up.
What it does **not** answer is whether they are correct, and that gap is the recurring theme of
this whole system: the deploy path reports transport rather than effect, so absence reads as
success. A check that confirms a service is running does not confirm the service is running the
code that was just deployed, and a node has been left on old code with a version marker that
had already advanced.
## Thoughts
Every node's daemon runs a periodic loop that surfaces observations from that node's own
context. They are stored in the mesh and can inform a session or trigger action.
It is the one part of the mesh that is not request-driven — the mesh noticing things rather
than being asked.
## The honest summary
Observability tells you the mesh is **up**. Establishing that it is **right** currently means
reading the operational record and checking by hand.
That is the gap the lab is designed to close
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)): a place where a
change can be run end to end and a verdict produced, cheaply enough that producing one is
routine.
+90
View File
@@ -0,0 +1,90 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0010-applications-live-in-their-own-repository.md
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
---
# The catalogue, and what its shape says
The catalogue holds **124 modules**. Thirty-three belong to the mesh's own domain; the other
ninety-one run *on* the mesh rather than being *of* it
([ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)).
The count is not the finding. The **shape** is.
## How it is organised today
By namespace, and the namespace records origin rather than purpose:
- **The mesh's own namespace** holds the platform: the runtime and its daemons, the shared
library, delivery, provisioning, configuration synchronisation, knowledge, identity, the
board, developer tooling, and node presentation.
- **A second namespace** holds the work domain — tasks, workflows, agents, meetings — split
across a handful of packages that share one schema.
- **Everything else sits flat at the top level**, one directory per piece of software.
## What the flat level actually contains
Grouped by what they are *for* — a grouping the catalogue itself does not express:
| Purpose | Roughly |
|---|---|
| Data and storage services the mesh provisions against | Relational and document databases, a cache, an object store, a package registry, a time-series store |
| Messaging and identity | A message broker, an identity provider |
| Reachability | A VPN, a firewall, an intrusion filter, an SSH daemon, a resolver, a certificate authority, a reverse proxy, network equipment control |
| Forge and container plumbing | Forge integrations, an image registry, container lifecycle and retention |
| Media libraries | Acquisition, organisation, playback, transcoding, streaming |
| Workstation and desktop | Browser, file manager, monitors, session management, audio, package management, runtime managers |
| Hardware-specific support | Power and firmware control for particular hardware, filesystem management |
| Collaboration and productivity | File sync, office tooling, boards, automation, chat and messaging bridges, mail, analytics, dashboards, home automation, issue trackers and wikis |
| Third-party organisation integrations | Systems belonging to organisations outside the mesh |
Every row is several modules, and **no row is a thing the mesh can see**. Four modules
together constitute "how a node is reachable", and they have no relationship the mesh can
assign, version, reason about or replace as one unit. A change to how the mesh handles
connectivity is made four times.
## What the shape records
**The catalogue's shape records what was installed, not what anything is for.** One module is
the unit of one piece of software, because that is the only granularity the module system
offers.
This is the same failure [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)
names for the platform core — *boundaries drawn by deployment accident rather than by domain* —
appearing outside it, at four times the scale. The core is being recomposed; the flat level is
addressed in principle by
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md), which
deliberately does not yet settle the domain list.
## Two properties worth keeping
Whatever replaces the shape, two things about it are right.
**Uniformity.** A media server and the mesh's own coordinator are installed, provisioned,
delivered and verified by identical machinery. The mesh's own components hold no privilege —
which is what makes dogfooding structural rather than a discipline, and what makes moving a
module out of the repository safe.
**Placement is already decided.** A standalone application belongs in its own repository
([ADR 0010](../../02-DECISIONS/0010-applications-live-in-their-own-repository.md)), and reviewers reject
it in the monorepo. The catalogue's flat level is not a dumping ground by policy; it is one by
history.
## Known inconsistencies in the catalogue itself
Recorded because a reader will meet them:
- A documented requirement that every capability-exposing module declare the core runtime as a
dependency is met by **zero** modules.
- A firewall-scoping key is declared by five manifests and read by none
([`04-ISSUES/003`](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)).
- A connections block in the manifest is metadata: it describes a module's reachability and
wires nothing.
- At least one module deliberately runs outside the standard per-module supervision, for
reasons recorded in the operational memory. The standard path is not universal.
+30
View File
@@ -0,0 +1,30 @@
# 03-DESIGN / 00-as-is
The mesh as it stands. These documents describe what runs, including the parts nobody would
choose again — an as-is layer that only records the good decisions is a brochure.
They are written from the implementation and from the operational record, not from intent.
Where the two disagree, the implementation wins and the disagreement is stated.
| Document | Covers |
|---|---|
| [`00-overview.md`](00-overview.md) | The whole in one pass — what a node is, what a module is, how work reaches it |
| [`01-mesh-and-transport.md`](01-mesh-and-transport.md) | The mesh database, the broker, discovery, and how a call reaches another node |
| [`02-modules-and-manifests.md`](02-modules-and-manifests.md) | The module, the manifest, and features as the unit of work |
| [`03-provisioning.md`](03-provisioning.md) | Declared requirements, provisioners, credentials, and cross-node grants |
| [`04-delivery.md`](04-delivery.md) | Push to running: the three silos, levels, and what a green pipeline proves |
| [`05-runtime-and-installation.md`](05-runtime-and-installation.md) | The node runtime, its modes, and how a node comes into being |
| [`06-configuration-and-secrets.md`](06-configuration-and-secrets.md) | Managed files, value resolution, and where secrets live |
| [`07-knowledge.md`](07-knowledge.md) | The two knowledge stores, and what each is for |
| [`08-agents-and-work.md`](08-agents-and-work.md) | Agents as employees, tasks, workflows, and the meeting model |
| [`09-interfaces-and-observability.md`](09-interfaces-and-observability.md) | How the mesh is reached and watched — tools, board, proxy, health, thoughts |
| [`10-module-catalogue.md`](10-module-catalogue.md) | The catalogue's shape, and what its shape says |
## What these documents are not
They are not a runbook. Operational procedure — how to fix one occurrence of something — lives
in the knowledge base, which is indexed on symptoms and is the right place to search when
something is broken.
They are not exhaustive. A subsystem is described to the depth at which its **design** is
visible; below that is code.
+158
View File
@@ -0,0 +1,158 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md]
---
# Work breakdown — the decomposition
How [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) gets built, in what order, and where a human must look.
Ordering is not preference. Each phase removes a constraint the next one needs gone.
---
## Rules of engagement
These exist so the work can run largely unattended without accumulating the kind of
damage this refactor is meant to remove.
### Autonomous by default
An agent may, without asking:
- read anything, measure anything, query any database read-only
- create branches, write code and tests, open pull requests
- run the test suite and typechecks
- write and update `hq/` documents
### Always stop and ask
- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a
live credential, removing a module from a node
- **merging anything** — every merge is a human checkpoint, without exception
- **a decision the ADRs do not already answer** — record the question in the relevant
research effort rather than picking and moving on
- **any change to `hq/00-META`** — it is stable by nature
### Definition of done for every task
1. tests written **and failing first**, then passing
2. typecheck clean in every package the change touches
3. the local mesh (Phase 0) comes up, and the behaviour is demonstrated in it
4. `hq/` updated if the task changed or answered anything documented
5. deployed, and **delivery verified on every node** — not "the pipeline was green"
### Non-negotiables carried from the current system
- **Never edit mesh-managed files on disk.** Use the owning tool.
- **Never write to production databases directly.** Migrations for schema, application
code for data.
- **Every schema change ships twice** — consolidated schema *and* an incremental
migration.
- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the
old one — never in a single step.
- **A green pipeline proves transport, not effect.** Verify the effect.
---
## Phase 0 — A mesh that runs locally *(prerequisite)*
Nothing else starts until this exists. Every fault this refactor addresses was found in
production because there was nowhere else to find it.
| # | task | done when |
|---|---|---|
| 0.1 | Container image for a node runtime | a node process starts in a container and registers |
| 0.2 | Compose topology: broker, registry DB, object store, *n* nodes | `up` yields a mesh that elects a provider node and settles |
| 0.3 | Seed a minimal mesh: nodes, one module, one provision | a module deploys end-to-end with no external service |
| 0.4 | Run the pipeline inside it | a push-equivalent produces a cascade and a deployed artifact |
| 0.5 | Fixtures for the failure modes already known | credential rotation reaching a running session; a provider deploy rotating a shared credential; a migration that ships nothing — each reproducible on demand |
**Checkpoint:** a human confirms the local mesh reproduces at least one bug from
2026-08-22 before any decomposition begins.
---
## Phase 1 — Make the model expressible
The decomposition is impossible while a feature is a singleton per module.
| # | task | done when |
|---|---|---|
| 1.1 | Decision record — named features, per-node opt-in (next free number) | accepted |
| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build |
| 1.3 | Selection: `always` / flavor-selected / `optional` | a node installs a subset; artifacts stay flavor-blind |
| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed |
| 1.5 | Assignment carries the opted-in feature set | opting a node in requires no rebuild |
**Checkpoint:** one existing module converted to declared features, deployed, verified —
before any others follow.
---
## Phase 2 — Draw the boundary the domain already has
Cheapest first, and each one proves the extraction pattern before the expensive ones.
| # | task | extracted from | risk |
|---|---|---|---|
| 2.1 | `hal/knowledge` — one store, review workflow ported | hippocampus + noxflow `knowledge_*` | low — additive |
| 2.2 | `hal/stream` — the record; notifications and messaging as views | axon, synapse, notifications, meetings, conversations | medium |
| 2.3 | `hal/agents` — identity, licence, runs, memory, thoughts | noxflow agents, `hal/thoughts` | **high** — touches credentials |
| 2.4 | `hal/work` — what remains of noxflow | noxflow tasks | medium |
| 2.5 | `hal/ai` — provider integration, flavored | `hal/claude*` | medium |
Each extraction is expand-then-contract: new context alongside, dual-write, verify, cut
over, remove. **Never a move commit.**
**Checkpoint:** after 2.1, a human confirms the extraction pattern before 2.2 begins.
After 2.3, a human confirms credentials still reach every agent on every node.
---
## Phase 3 — Reclaim the kernel
Only possible once domains have modules to own their code.
| # | task | done when |
|---|---|---|
| 3.1 | Move work-domain code out of `hal/sdk` | `workflow-engine.ts`, `task-commands.ts` live in `hal/work` |
| 3.2 | Move provider code out | `claude-credentials.ts` lives in `hal/ai` |
| 3.3 | Move delivery code out | feature handlers, artifact manager, build executor live in `hal/delivery` |
| 3.4 | Decide the residue | ADR: what `hal/sdk` keeps (open question 4) |
**Measure:** `hal/sdk` line count, tracked per task. Today: **34,636** across **155**
files.
---
## Phase 4 — Separate what the mesh runs from the mesh
| # | task | done when |
|---|---|---|
| 4.1 | Decide the destination (open question 3) | ADR accepted |
| 4.2 | Cross-repository dependency resolution proven | a catalogue module builds against a published `@hal/*` |
| 4.3 | Move the 91 catalogue modules | this repository contains only mesh contexts |
**Checkpoint:** move one application first and run it for a week before the rest follow.
---
## Sequencing constraints
- **0 before everything.** Unverifiable refactors are how this list got long.
- **1 before 2.** Extracting into contexts without per-node features recreates the module
count inside the new names.
- **2 before 3.** A domain can only own its shared code once the domain has a module.
- **2.3 after 2.1 and 2.2.** Agents touch credentials; do it once the pattern is proven on
cheaper contexts.
- **4 last.** It is the only phase that is pure movement, so it is the only one safe to
defer indefinitely.
## What "done" looks like
Eight contexts. `hal/sdk` holding only what is genuinely cross-cutting. A mesh that stands
up on a laptop. A module count that grows only when the domain does.
+363
View File
@@ -0,0 +1,363 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
---
# End-to-end testing
**What is under test is a module.** The mesh is the harness.
You change a module or write a new one, run it end to end, and get a verdict before it goes
anywhere near production. That loop is the product of this design; everything else exists to
make it fast and honest.
This is the Phase 0 prerequisite from [`00-work-breakdown.md`](00-work-breakdown.md).
---
## Where this sits in the way work happens
Work reaches the mesh along one path today:
```
a change is made a human in a session, or an agent given work
│
▼
a pull request appears
│
▼
a human reads the diff and merges ← the gate
│
▼
the coordinator delivers to the real nodes
│
▼
production reports whether it worked ← the test
```
**The gate is a human reading a diff, and the test is production.** That is workable at a
change a day and it is the constraint at ten. For autonomous work it is worse than a
constraint: an agent's output arrives as a diff that *looks* right, carrying no evidence
that it runs, and the only reviewer is the condition `00-META/context.md` calls mandatory
— *human agents are few, often one, and usually asleep.*
The missing step goes between the pull request and the merge:
```
a pull request appears
│
▼
the coordinator delivers the branch to a SCENARIO mesh ← the missing step
and runs the same stages, ending in verify
│
▼
the verdict is attached to the pull request
│
▼
a human merges evidence rather than hope
│
▼
the coordinator delivers to the real nodes — same verify, now loud
```
### One pipeline, two targets
| target | triggered by | what a failure means |
|---|---|---|
| a scenario mesh | a branch, or a pull request | the change is not finished; it should not merge |
| the real mesh | a merge | a red delivery, loudly, before anything is built on it |
Same coordinator, same cascade, same stages, same verification. **Only the target mesh
differs.** This is the through-line of the whole design — one verification statement with two
jobs, one runner with two callers, one pipeline with two targets. Nothing forks, so nothing
drifts.
### What it demands
- **The coordinator must accept a target mesh.** Today a pipeline's targets are derived: the
nodes that have the module assigned. It needs to be able to run the same pipeline against
a mesh named by the request instead.
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
at once, each needing its own network and nodes. This is affordable with system containers
and would not be with virtual machines — the unit choice is what makes the gate possible
at all.
- **The gate is only as good as the verification behind it.** A module with no assertions
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
the number that matters**, and it starts at approximately zero.
- **It has to be fast enough to wait for.** A gate an agent cannot wait on is a report nobody
reads.
---
## The coordinator drives it
The temptation is to build a framework that delivers a module and checks it. That would be a
**second delivery path**, and a second delivery path is worthless — the faults worth catching
live in the real one.
So the rule is not "no new components". It is: **nothing new drives delivery.** A scenario is
a complete mesh with its own coordinator. Push the working tree to that mesh's forge; its
coordinator does exactly what a coordinator does — works out the cascade, dispatches build,
install, configure, start, and then **verify** — and its meshware executes on its nodes. The
result of that pipeline *is* the verdict.
A **test runner** is a legitimate component within that, used *by* the coordinator rather
than instead of it. Two jobs plausibly belong to it, and their boundary is worth settling
before either is built:
- **scenario lifecycle** — materialise the mesh a test needs, restore it to a snapshot, tear
it down. Something must do this before a coordinator exists to drive anything.
- **assertion execution** — give verification more than "run a script and check the exit
code": setup and teardown, timeouts, retry-until-true for things that settle, and results
structured enough to report rather than grep.
The line to hold is the pipeline itself. A runner that stands up a mesh and executes
assertions is a component. A runner that decides what to build, in what order, and ships it
to a node is a fork of the coordinator.
### It has two callers, and they want different things
The runner serves **the coordinator** and **a person developing the mesh**, and its interface
has to suit both:
| caller | wants |
|---|---|
| the coordinator | non-interactive, structured results it can record against a pipeline, a clean teardown, no prompts and no colour |
| someone working on HAL | readable output, the failing mesh **left standing** to open a shell into, and a way to re-run one assertion without repeating the whole delivery |
Hence at least two verbs: one that runs to a verdict and tears down, and one that stands a
scenario up and leaves it there. The second is how a developer works *inside* a mesh —
which is the thing the current host-borrowing tooling is really for, and the reason this
replaces it rather than sitting beside it.
That the same runner serves both is deliberate, and it is the same argument as the module's
own assertions serving both development and delivery: **one statement, two jobs.** Anything
that only the developer path can do is a divergence, and it will drift.
```
edit a module in the working tree
│
▼
push to the scenario's forge
│
▼
the scenario's COORDINATOR runs a pipeline ← existing machinery, unchanged
│
├── cascade: which modules are affected
├── build → install → configure → start
└── verify: the module's own assertions ← existing stage, dispatched today
│
▼
the pipeline result is the verdict
```
This is the same property `00-META/mission.md` asks for: *the mesh's own components ship
through the same machinery as anything else it carries — if they need an exception, the
machinery is not finished.* A test that needed its own delivery path would be that exception.
**It tests the working tree**, because the forge is inside the scenario. A loop that requires
pushing to production and waiting is not a loop. Real path, local code, nothing shared with
production.
---
## What it catches
This list is the specification. These are the ways a module change fails today, and each one
currently reaches production or wastes a pipeline run:
| failure | why it survives today |
|---|---|
| a file never reached the artifact | absence and "declares nothing here" are indistinguishable, so it ships green |
| a migration compiled to nothing, or never ran | the stage reports success when there is nothing to run |
| the manifest is wrong — bad package list, wrong paths | validated shallowly, if at all |
| a capability was never provisioned, or its credential never arrived | delivery reports transport, not effect |
| an environment value was not generated | the module starts and reads a default |
| the service did not come up, or came up and crashed | nothing asserts it is still running a minute later |
| the dependency cascade did not include the module | a green pipeline that rebuilt the wrong set |
| it works on a fresh install but breaks on upgrade | almost never exercised — see below |
| it works on one node and not another | only one node is ever tried |
A test that only proves "the pipeline went green" reproduces the exact blindness this is
meant to remove.
---
## Fresh install and upgrade are different tests
The most common shape of a module bug is: works from scratch, breaks on the machine that
already had the previous version. Existing state outranks new state, a file is added but
never removed, a migration assumes a column that an older node lacks.
Snapshots make both cheap, so both are default:
- **fresh** — restore a mesh that has never seen the module, deliver, assert
- **upgrade** — restore a mesh running the *previous released version*, deliver the working
tree over it, assert
Same assertions, different starting state. A module that passes one and fails the other is
the normal case, not an edge case.
---
## A module carries its own assertions
**A module states what must be true about it, and it states it once.** That statement is the
module's verification — the stage the coordinator already dispatches at the end of every
delivery, with a working handler, implemented today by essentially nothing.
It asserts **outcomes**, never that a step ran: the unit is active and still active shortly
after, the schema has the column, the name resolves, the credential authenticates, the file
on the node holds what the mesh believes it holds, the endpoint answers.
The same statement serves both places, which is the point:
| where it runs | what a failure means |
|---|---|
| in a scenario, during development | your change is not finished — cheap, fast, nobody affected |
| on delivery to production | the deploy is red, loudly, before anyone builds on it |
**This is what finally makes verification worth writing.** Today it can only ever cost you a
deploy, which is precisely why almost no module has one. Give it a second job — telling a
developer whether their change works — and writing it stops being an act of discipline and
starts being the fastest way to get an answer.
Where a module has no verification yet, the coordinator asserts the generic invariants it can
know on its own: the artifact contained what the manifest declared, the migrations that were
pending ran, the declared capabilities were provisioned, the declared services are up.
---
## The mesh shape is a parameter
A test declares the mesh it needs, and the default is the smallest one that can exercise the
module:
```yaml
module: a-web-service
mesh:
my-cool-node: { role: published, publishes: my-cool-node.com }
assert:
- https://my-cool-node.com answers 200
- the certificate presented is valid for that name
```
One node, because one node is enough to answer that question. Standing up a mesh to test one
module is the same mistake as starting the application to test a function.
More nodes when the module's behaviour is *between* nodes:
```yaml
module: a-module-requiring-a-database
mesh:
store: { role: anchor }
consumer-a: { role: resident }
consumer-b: { role: mobile }
assert:
- both consumers authenticate against the database
- after redeploying the provider, both still do
```
Node names and domains in a test are **invented**. The mesh under test is whatever the test
says it is — which is also how this document stays free of any particular installation.
### Scale, when scale is the question
Size is chosen by what is being asked, and ranges from one node to twenty or more.
| size | what only this size answers |
|---|---|
| one | does it install, migrate, provision and run at all — the fastest loop |
| two–three | anything *between* nodes: delivery, provisioning, rotation, absence |
| ten–twenty | whether a fan-out reaches *every* node, whether the cascade converges, whether something is quietly quadratic |
A fan-out reaching three of four nodes reads as a flake; at twenty it is a diagnosis, and
which nodes were missed tells you why. **This is what the unit choice below bought** — twenty
virtual machines do not fit on a workstation, and twenty system containers do.
---
## Two kinds of test
**Module tests** are the daily case and the reason this exists: does my change work, end to
end.
**Mesh tests** use the same machinery to ask whether the mesh itself behaves — that a
credential rotation reaches every consumer, that delivery to an absent node is reported as
pending rather than done, that a returning node catches up. These are fewer and change
rarely, but they are where the known production faults get encoded so they stay fixed.
The known faults become mesh tests that fail today. That is the Phase 0 checkpoint.
---
## The substrate
### A node is a system container
An OS userspace with its own init, its own network interface, its own filesystem, sharing
the host kernel.
**A node's job is to run containers**, so modelling a node *as* an application container
inverts the thing being modelled: it forces nested containers through a privileged daemon or
a shared socket, and a shared socket makes isolation between nodes cosmetic. A system
container has no such problem — init runs as PID 1, so units and timers work as written;
containers nest properly, so module stacks run as they do anywhere. **Module code and mesh
code both run unmodified**, which is the property that makes a verdict trustworthy. Anything
needing a special case locally is a divergence that will hide a fault.
Boot is around a second and snapshots are cheap, which is what makes this an inner loop
rather than an errand.
**Any node can be a full virtual machine instead**, through the same tooling and the same
test file — when a question needs a kernel to answer it, or when the point is that nodes are
*not* identical.
### The network
Two segments and an overlay, because some module behaviour is only visible across a real
network boundary:
- **wan** — a published node holds an address here, and an authoritative resolver maps its
name to it, so a public touchpoint is real enough to exercise routing, virtual hosts and
certificates
- **local** — behind translation, as a home network is
- **the overlay** — the mechanism production uses; the mesh addresses peers by mesh name and
never learns which segment anyone is on
A node can be moved between segments or detached entirely, mid-test.
### What is not real
- **the model provider** — thinking is stubbed, so tests cost nothing to run
- **the public internet** — a bridge, with an authoritative resolver rather than delegation
- **the public certificate authority** — the lab runs **its own ACME issuer** on the public
segment, so issuance, challenge and renewal are genuinely exercised rather than stubbed.
Internal names keep the mesh CA, so the lab preserves production's two-authority split
rather than collapsing it into one.
Everything a node itself does is real, because a node is a real machine.
---
## Consequences
**Bringing a node into being is part of the framework.** A test creates its own nodes — one
for the smallest, twenty for the largest — repeatably, unattended, and cheaply enough to do
it twenty times in a row. Those are the constraints a real mesh wants, so the mechanism the
runner needs is the one the mesh should keep.
**Module verification becomes worth writing**, because it is the thing that gives a
developer a verdict, not just a stricter deploy.
**Host-borrowing ends.** Today's tooling starts providers on the host's own init system and
reads credentials from host paths, because there is nowhere else to put a mesh. Once there
is, a workstation stops being collateral.
**The supervision question stops gating anything.**
[`01-RESEARCH/003-service-supervision`](../../01-RESEARCH/003-service-supervision/analysis.md)
remains open on its own merits — and once this exists, its options are cheap to try rather
than expensive to argue about.
+22
View File
@@ -0,0 +1,22 @@
# 03-DESIGN / 01-to-be
The mesh being built toward. Every statement here traces to a record in
[`02-DECISIONS/`](../../02-DECISIONS/); nothing arrives by drafting.
A document here describes an intention. What currently runs is in
[`00-as-is/`](../00-as-is/), and the two are never merged — when something ships, the as-is
document is written and this one's status becomes `implemented`.
| Document | Covers | Rests on |
|---|---|---|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) |
## Not yet written
- **The eight bounded contexts.** [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)
decides the decomposition; the per-context specifications do not exist yet. The work
breakdown says in what order they are needed.
- **Domain grouping outside the core.** [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)
settles the principle and explicitly does not settle the domain list. That is a research
effort, not a design document, until it concludes.
+51
View File
@@ -0,0 +1,51 @@
# 03-DESIGN
The authoritative specification. Implementation is built against what is written here.
## Two layers
| Folder | What it is |
|---|---|
| [`00-as-is/`](00-as-is/) | **The mesh that exists today.** Shipped behaviour, described as it is — including behaviour nobody would choose again. |
| [`01-to-be/`](01-to-be/) | **The mesh being built toward.** Every statement traceable to a record in [`02-DECISIONS/`](../02-DECISIONS/). |
They are never mixed. A statement about the future does not belong in an as-is document, and
an as-is document is never edited to describe an intention.
When a to-be design ships, it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand — one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning, which is
the expensive half.
## Frontmatter
Every design document (not the READMEs) carries:
```yaml
---
layer: as-is | to-be
status: designed | in-progress | implemented | abandoned
code: [] # owning code repo(s), from 00-META/repos.md
updated: YYYY-MM-DD # date of the last status change, not of text edits
decisions: [] # 02-DECISIONS/ records this document rests on
---
```
For an as-is document, `status: implemented` is the normal state — it describes something that
runs — and `code:` names where that implementation lives.
Status changes when **implementation state** changes, never because design text was edited. An
`implemented` claim must be defensible from the owning repository's main branch, not from
intent. If it cannot be checked, it is `in-progress`.
Cross-cutting views are generated from this frontmatter by the `hal-status` skill and never
written to disk.
## What belongs here
Functional analysis, architectural description, and specification — **prose and diagrams
only, no code**. A manifest field may be named; a manifest may not be pasted. A document
enters the to-be layer only after the decision behind it is recorded in [`02-DECISIONS/`](../02-DECISIONS/)
and the research that produced it is closed.
Subfolders are encouraged where a layer grows enough to need them.