Base layer: the mesh as it is, under the mesh as it should be

HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
This commit is contained in:
2026-08-23 03:08:26 +02:00
parent cf9357e8e9
commit 702efca6bb
74 changed files with 3676 additions and 138 deletions
+158
View File
@@ -0,0 +1,158 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [adr/0015-mesh-brokers-nodes-host-agents-think.md]
---
# Work breakdown — the decomposition
How [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md) gets built, in what order, and where a human must look.
Ordering is not preference. Each phase removes a constraint the next one needs gone.
---
## Rules of engagement
These exist so the work can run largely unattended without accumulating the kind of
damage this refactor is meant to remove.
### Autonomous by default
An agent may, without asking:
- read anything, measure anything, query any database read-only
- create branches, write code and tests, open pull requests
- run the test suite and typechecks
- write and update `hq/` documents
### Always stop and ask
- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a
live credential, removing a module from a node
- **merging anything** — every merge is a human checkpoint, without exception
- **a decision the ADRs do not already answer** — record the question in the relevant
research effort rather than picking and moving on
- **any change to `hq/00-GENESIS`** — it is stable by nature
### Definition of done for every task
1. tests written **and failing first**, then passing
2. typecheck clean in every package the change touches
3. the local mesh (Phase 0) comes up, and the behaviour is demonstrated in it
4. `hq/` updated if the task changed or answered anything documented
5. deployed, and **delivery verified on every node** — not "the pipeline was green"
### Non-negotiables carried from the current system
- **Never edit mesh-managed files on disk.** Use the owning tool.
- **Never write to production databases directly.** Migrations for schema, application
code for data.
- **Every schema change ships twice** — consolidated schema *and* an incremental
migration.
- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the
old one — never in a single step.
- **A green pipeline proves transport, not effect.** Verify the effect.
---
## Phase 0 — A mesh that runs locally *(prerequisite)*
Nothing else starts until this exists. Every fault this refactor addresses was found in
production because there was nowhere else to find it.
| # | task | done when |
|---|---|---|
| 0.1 | Container image for a node runtime | a node process starts in a container and registers |
| 0.2 | Compose topology: broker, registry DB, object store, *n* nodes | `up` yields a mesh that elects a provider node and settles |
| 0.3 | Seed a minimal mesh: nodes, one module, one provision | a module deploys end-to-end with no external service |
| 0.4 | Run the pipeline inside it | a push-equivalent produces a cascade and a deployed artifact |
| 0.5 | Fixtures for the failure modes already known | credential rotation reaching a running session; a provider deploy rotating a shared credential; a migration that ships nothing — each reproducible on demand |
**Checkpoint:** a human confirms the local mesh reproduces at least one bug from
2026-08-22 before any decomposition begins.
---
## Phase 1 — Make the model expressible
The decomposition is impossible while a feature is a singleton per module.
| # | task | done when |
|---|---|---|
| 1.1 | Decision record — named features, per-node opt-in (next free number) | accepted |
| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build |
| 1.3 | Selection: `always` / flavor-selected / `optional` | a node installs a subset; artifacts stay flavor-blind |
| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed |
| 1.5 | Assignment carries the opted-in feature set | opting a node in requires no rebuild |
**Checkpoint:** one existing module converted to declared features, deployed, verified —
before any others follow.
---
## Phase 2 — Draw the boundary the domain already has
Cheapest first, and each one proves the extraction pattern before the expensive ones.
| # | task | extracted from | risk |
|---|---|---|---|
| 2.1 | `hal/knowledge` — one store, review workflow ported | hippocampus + noxflow `knowledge_*` | low — additive |
| 2.2 | `hal/stream` — the record; notifications and messaging as views | axon, synapse, notifications, meetings, conversations | medium |
| 2.3 | `hal/agents` — identity, licence, runs, memory, thoughts | noxflow agents, `hal/thoughts` | **high** — touches credentials |
| 2.4 | `hal/work` — what remains of noxflow | noxflow tasks | medium |
| 2.5 | `hal/ai` — provider integration, flavored | `hal/claude*` | medium |
Each extraction is expand-then-contract: new context alongside, dual-write, verify, cut
over, remove. **Never a move commit.**
**Checkpoint:** after 2.1, a human confirms the extraction pattern before 2.2 begins.
After 2.3, a human confirms credentials still reach every agent on every node.
---
## Phase 3 — Reclaim the kernel
Only possible once domains have modules to own their code.
| # | task | done when |
|---|---|---|
| 3.1 | Move work-domain code out of `hal/sdk` | `workflow-engine.ts`, `task-commands.ts` live in `hal/work` |
| 3.2 | Move provider code out | `claude-credentials.ts` lives in `hal/ai` |
| 3.3 | Move delivery code out | feature handlers, artifact manager, build executor live in `hal/delivery` |
| 3.4 | Decide the residue | ADR: what `hal/sdk` keeps (open question 4) |
**Measure:** `hal/sdk` line count, tracked per task. Today: **34,636** across **155**
files.
---
## Phase 4 — Separate what the mesh runs from the mesh
| # | task | done when |
|---|---|---|
| 4.1 | Decide the destination (open question 3) | ADR accepted |
| 4.2 | Cross-repository dependency resolution proven | a catalogue module builds against a published `@hal/*` |
| 4.3 | Move the 91 catalogue modules | this repository contains only mesh contexts |
**Checkpoint:** move one application first and run it for a week before the rest follow.
---
## Sequencing constraints
- **0 before everything.** Unverifiable refactors are how this list got long.
- **1 before 2.** Extracting into contexts without per-node features recreates the module
count inside the new names.
- **2 before 3.** A domain can only own its shared code once the domain has a module.
- **2.3 after 2.1 and 2.2.** Agents touch credentials; do it once the pattern is proven on
cheaper contexts.
- **4 last.** It is the only phase that is pure movement, so it is the only one safe to
defer indefinitely.
## What "done" looks like
Eight contexts. `hal/sdk` holding only what is genuinely cross-cutting. A mesh that stands
up on a laptop. A module count that grows only when the domain does.
+363
View File
@@ -0,0 +1,363 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [adr/0016-a-lab-node-is-a-virtual-machine.md]
---
# End-to-end testing
**What is under test is a module.** The mesh is the harness.
You change a module or write a new one, run it end to end, and get a verdict before it goes
anywhere near production. That loop is the product of this design; everything else exists to
make it fast and honest.
This is the Phase 0 prerequisite from [`00-work-breakdown.md`](00-work-breakdown.md).
---
## Where this sits in the way work happens
Work reaches the mesh along one path today:
```
a change is made a human in a session, or an agent given work
│
▼
a pull request appears
│
▼
a human reads the diff and merges ← the gate
│
▼
the coordinator delivers to the real nodes
│
▼
production reports whether it worked ← the test
```
**The gate is a human reading a diff, and the test is production.** That is workable at a
change a day and it is the constraint at ten. For autonomous work it is worse than a
constraint: an agent's output arrives as a diff that *looks* right, carrying no evidence
that it runs, and the only reviewer is the condition `00-GENESIS/context.md` calls mandatory
— *human agents are few, often one, and usually asleep.*
The missing step goes between the pull request and the merge:
```
a pull request appears
│
▼
the coordinator delivers the branch to a SCENARIO mesh ← the missing step
and runs the same stages, ending in verify
│
▼
the verdict is attached to the pull request
│
▼
a human merges evidence rather than hope
│
▼
the coordinator delivers to the real nodes — same verify, now loud
```
### One pipeline, two targets
| target | triggered by | what a failure means |
|---|---|---|
| a scenario mesh | a branch, or a pull request | the change is not finished; it should not merge |
| the real mesh | a merge | a red delivery, loudly, before anything is built on it |
Same coordinator, same cascade, same stages, same verification. **Only the target mesh
differs.** This is the through-line of the whole design — one verification statement with two
jobs, one runner with two callers, one pipeline with two targets. Nothing forks, so nothing
drifts.
### What it demands
- **The coordinator must accept a target mesh.** Today a pipeline's targets are derived: the
nodes that have the module assigned. It needs to be able to run the same pipeline against
a mesh named by the request instead.
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
at once, each needing its own network and nodes. This is affordable with system containers
and would not be with virtual machines — the unit choice is what makes the gate possible
at all.
- **The gate is only as good as the verification behind it.** A module with no assertions
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
the number that matters**, and it starts at approximately zero.
- **It has to be fast enough to wait for.** A gate an agent cannot wait on is a report nobody
reads.
---
## The coordinator drives it
The temptation is to build a framework that delivers a module and checks it. That would be a
**second delivery path**, and a second delivery path is worthless — the faults worth catching
live in the real one.
So the rule is not "no new components". It is: **nothing new drives delivery.** A scenario is
a complete mesh with its own coordinator. Push the working tree to that mesh's forge; its
coordinator does exactly what a coordinator does — works out the cascade, dispatches build,
install, configure, start, and then **verify** — and its meshware executes on its nodes. The
result of that pipeline *is* the verdict.
A **test runner** is a legitimate component within that, used *by* the coordinator rather
than instead of it. Two jobs plausibly belong to it, and their boundary is worth settling
before either is built:
- **scenario lifecycle** — materialise the mesh a test needs, restore it to a snapshot, tear
it down. Something must do this before a coordinator exists to drive anything.
- **assertion execution** — give verification more than "run a script and check the exit
code": setup and teardown, timeouts, retry-until-true for things that settle, and results
structured enough to report rather than grep.
The line to hold is the pipeline itself. A runner that stands up a mesh and executes
assertions is a component. A runner that decides what to build, in what order, and ships it
to a node is a fork of the coordinator.
### It has two callers, and they want different things
The runner serves **the coordinator** and **a person developing the mesh**, and its interface
has to suit both:
| caller | wants |
|---|---|
| the coordinator | non-interactive, structured results it can record against a pipeline, a clean teardown, no prompts and no colour |
| someone working on HAL | readable output, the failing mesh **left standing** to open a shell into, and a way to re-run one assertion without repeating the whole delivery |
Hence at least two verbs: one that runs to a verdict and tears down, and one that stands a
scenario up and leaves it there. The second is how a developer works *inside* a mesh —
which is the thing the current host-borrowing tooling is really for, and the reason this
replaces it rather than sitting beside it.
That the same runner serves both is deliberate, and it is the same argument as the module's
own assertions serving both development and delivery: **one statement, two jobs.** Anything
that only the developer path can do is a divergence, and it will drift.
```
edit a module in the working tree
│
▼
push to the scenario's forge
│
▼
the scenario's COORDINATOR runs a pipeline ← existing machinery, unchanged
│
├── cascade: which modules are affected
├── build → install → configure → start
└── verify: the module's own assertions ← existing stage, dispatched today
│
▼
the pipeline result is the verdict
```
This is the same property `00-GENESIS/mission.md` asks for: *the mesh's own components ship
through the same machinery as anything else it carries — if they need an exception, the
machinery is not finished.* A test that needed its own delivery path would be that exception.
**It tests the working tree**, because the forge is inside the scenario. A loop that requires
pushing to production and waiting is not a loop. Real path, local code, nothing shared with
production.
---
## What it catches
This list is the specification. These are the ways a module change fails today, and each one
currently reaches production or wastes a pipeline run:
| failure | why it survives today |
|---|---|
| a file never reached the artifact | absence and "declares nothing here" are indistinguishable, so it ships green |
| a migration compiled to nothing, or never ran | the stage reports success when there is nothing to run |
| the manifest is wrong — bad package list, wrong paths | validated shallowly, if at all |
| a capability was never provisioned, or its credential never arrived | delivery reports transport, not effect |
| an environment value was not generated | the module starts and reads a default |
| the service did not come up, or came up and crashed | nothing asserts it is still running a minute later |
| the dependency cascade did not include the module | a green pipeline that rebuilt the wrong set |
| it works on a fresh install but breaks on upgrade | almost never exercised — see below |
| it works on one node and not another | only one node is ever tried |
A test that only proves "the pipeline went green" reproduces the exact blindness this is
meant to remove.
---
## Fresh install and upgrade are different tests
The most common shape of a module bug is: works from scratch, breaks on the machine that
already had the previous version. Existing state outranks new state, a file is added but
never removed, a migration assumes a column that an older node lacks.
Snapshots make both cheap, so both are default:
- **fresh** — restore a mesh that has never seen the module, deliver, assert
- **upgrade** — restore a mesh running the *previous released version*, deliver the working
tree over it, assert
Same assertions, different starting state. A module that passes one and fails the other is
the normal case, not an edge case.
---
## A module carries its own assertions
**A module states what must be true about it, and it states it once.** That statement is the
module's verification — the stage the coordinator already dispatches at the end of every
delivery, with a working handler, implemented today by essentially nothing.
It asserts **outcomes**, never that a step ran: the unit is active and still active shortly
after, the schema has the column, the name resolves, the credential authenticates, the file
on the node holds what the mesh believes it holds, the endpoint answers.
The same statement serves both places, which is the point:
| where it runs | what a failure means |
|---|---|
| in a scenario, during development | your change is not finished — cheap, fast, nobody affected |
| on delivery to production | the deploy is red, loudly, before anyone builds on it |
**This is what finally makes verification worth writing.** Today it can only ever cost you a
deploy, which is precisely why almost no module has one. Give it a second job — telling a
developer whether their change works — and writing it stops being an act of discipline and
starts being the fastest way to get an answer.
Where a module has no verification yet, the coordinator asserts the generic invariants it can
know on its own: the artifact contained what the manifest declared, the migrations that were
pending ran, the declared capabilities were provisioned, the declared services are up.
---
## The mesh shape is a parameter
A test declares the mesh it needs, and the default is the smallest one that can exercise the
module:
```yaml
module: a-web-service
mesh:
my-cool-node: { role: published, publishes: my-cool-node.com }
assert:
- https://my-cool-node.com answers 200
- the certificate presented is valid for that name
```
One node, because one node is enough to answer that question. Standing up a mesh to test one
module is the same mistake as starting the application to test a function.
More nodes when the module's behaviour is *between* nodes:
```yaml
module: a-module-requiring-a-database
mesh:
store: { role: anchor }
consumer-a: { role: resident }
consumer-b: { role: mobile }
assert:
- both consumers authenticate against the database
- after redeploying the provider, both still do
```
Node names and domains in a test are **invented**. The mesh under test is whatever the test
says it is — which is also how this document stays free of any particular installation.
### Scale, when scale is the question
Size is chosen by what is being asked, and ranges from one node to twenty or more.
| size | what only this size answers |
|---|---|
| one | does it install, migrate, provision and run at all — the fastest loop |
| two–three | anything *between* nodes: delivery, provisioning, rotation, absence |
| ten–twenty | whether a fan-out reaches *every* node, whether the cascade converges, whether something is quietly quadratic |
A fan-out reaching three of four nodes reads as a flake; at twenty it is a diagnosis, and
which nodes were missed tells you why. **This is what the unit choice below bought** — twenty
virtual machines do not fit on a workstation, and twenty system containers do.
---
## Two kinds of test
**Module tests** are the daily case and the reason this exists: does my change work, end to
end.
**Mesh tests** use the same machinery to ask whether the mesh itself behaves — that a
credential rotation reaches every consumer, that delivery to an absent node is reported as
pending rather than done, that a returning node catches up. These are fewer and change
rarely, but they are where the known production faults get encoded so they stay fixed.
The known faults become mesh tests that fail today. That is the Phase 0 checkpoint.
---
## The substrate
### A node is a system container
An OS userspace with its own init, its own network interface, its own filesystem, sharing
the host kernel.
**A node's job is to run containers**, so modelling a node *as* an application container
inverts the thing being modelled: it forces nested containers through a privileged daemon or
a shared socket, and a shared socket makes isolation between nodes cosmetic. A system
container has no such problem — init runs as PID 1, so units and timers work as written;
containers nest properly, so module stacks run as they do anywhere. **Module code and mesh
code both run unmodified**, which is the property that makes a verdict trustworthy. Anything
needing a special case locally is a divergence that will hide a fault.
Boot is around a second and snapshots are cheap, which is what makes this an inner loop
rather than an errand.
**Any node can be a full virtual machine instead**, through the same tooling and the same
test file — when a question needs a kernel to answer it, or when the point is that nodes are
*not* identical.
### The network
Two segments and an overlay, because some module behaviour is only visible across a real
network boundary:
- **wan** — a published node holds an address here, and an authoritative resolver maps its
name to it, so a public touchpoint is real enough to exercise routing, virtual hosts and
certificates
- **local** — behind translation, as a home network is
- **the overlay** — the mechanism production uses; the mesh addresses peers by mesh name and
never learns which segment anyone is on
A node can be moved between segments or detached entirely, mid-test.
### What is not real
- **the model provider** — thinking is stubbed, so tests cost nothing to run
- **the public internet** — a bridge, with an authoritative resolver rather than delegation
- **the public certificate authority** — the lab runs **its own ACME issuer** on the public
segment, so issuance, challenge and renewal are genuinely exercised rather than stubbed.
Internal names keep the mesh CA, so the lab preserves production's two-authority split
rather than collapsing it into one.
Everything a node itself does is real, because a node is a real machine.
---
## Consequences
**Bringing a node into being is part of the framework.** A test creates its own nodes — one
for the smallest, twenty for the largest — repeatably, unattended, and cheaply enough to do
it twenty times in a row. Those are the constraints a real mesh wants, so the mechanism the
runner needs is the one the mesh should keep.
**Module verification becomes worth writing**, because it is the thing that gives a
developer a verdict, not just a stricter deploy.
**Host-borrowing ends.** Today's tooling starts providers on the host's own init system and
reads credentials from host paths, because there is nowhere else to put a mesh. Once there
is, a workstation stops being collateral.
**The supervision question stops gating anything.**
[`01-RESEARCH/003-service-supervision`](../../01-RESEARCH/003-service-supervision/analysis.md)
remains open on its own merits — and once this exists, its options are cheap to try rather
than expensive to argue about.
+22
View File
@@ -0,0 +1,22 @@
# 02-DESIGN / 01-to-be
The mesh being built toward. Every statement here traces to a record in
[`adr/`](../../adr/); nothing arrives by drafting.
A document here describes an intention. What currently runs is in
[`00-as-is/`](../00-as-is/), and the two are never merged — when something ships, the as-is
document is written and this one's status becomes `implemented`.
| Document | Covers | Rests on |
|---|---|---|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../adr/0016-a-lab-node-is-a-virtual-machine.md) |
## Not yet written
- **The eight bounded contexts.** [ADR 0015](../../adr/0015-mesh-brokers-nodes-host-agents-think.md)
decides the decomposition; the per-context specifications do not exist yet. The work
breakdown says in what order they are needed.
- **Domain grouping outside the core.** [ADR 0017](../../adr/0017-modules-outside-the-core-are-grouped-by-domain.md)
settles the principle and explicitly does not settle the domain list. That is a research
effort, not a design document, until it concludes.