HQ: the as-is base layer, the process, and the names #1

Merged
jschoubben merged 10 commits from docs/as-is-base-layer-and-process into main 2026-08-23 19:21:14 +00:00
97 changed files with 5678 additions and 259 deletions
+45
View File
@@ -0,0 +1,45 @@
---
name: hal-amend-design
description: Use when an HQ design document must change, or when an as-is document is found to be wrong about what the mesh actually does. Triggers on "the design changed", "that's not how it works any more", "update the as-is", "this shipped differently".
---
# hal-amend-design
Changes a design document. **Authoritative playbook:**
[`00-META/process/02-graduation.md`](../../../00-META/process/02-graduation.md).
## Which layer is being changed
**Establish this first — the two amend for opposite reasons.**
| Layer | Amend when | Requires |
|---|---|---|
| `01-to-be/` | The intention changed | A decision record. Always. |
| `00-as-is/` | The mesh changed, or the document was wrong about it | Evidence from the implementation. Never a decision. |
An as-is document is corrected against **what the code does**, not against what anyone meant.
If the implementation and the intention disagree, the as-is document records the implementation
and says they disagree.
## Amending the to-be layer
1. Write the decision record. If it reverses an earlier one, that record gets
`status: superseded` and `superseded-by:` — its text is never edited.
2. Edit the design; set `updated:` to today.
3. If this came from an issue, set that issue's `amended-design:`.
## Amending the as-is layer
1. Establish what is actually true — from the code, a pipeline log, a live query, the
operational record. Not from a design document.
2. Edit the document. Set `updated:` only if the **implementation state** changed; a
correction to text that was always wrong does not change it.
3. If the correction reveals that something shipped differently from its design, say so in the
as-is document and leave the to-be document alone. That divergence is a finding, and may
deserve an issue.
## Do not
- Do not put a future intention in an as-is document.
- Do not quietly fix a wrong claim that held up a decision. Record that it was wrong and where
it was relied on — that is the failure this repository exists to name.
+40
View File
@@ -0,0 +1,40 @@
---
name: hal-diagnose
description: Use when investigating an open HQ issue — finding which component owns a symptom, and why. Triggers on "diagnose issue N", "where does this live", "who owns this bug", "why does this happen".
---
# hal-diagnose
Investigates an open issue to the point where its owner is known. **Authoritative playbook:**
[`00-META/process/03-issues.md`](../../../00-META/process/03-issues.md).
## Before forming a hypothesis
**Search the operational memory for the literal symptom text.** Not after a hypothesis fails —
before forming one. The knowledge base is indexed on symptoms, and the entry needed is usually
titled after the error being stared at.
This fires hardest on familiar ground, where a confident trail feels like progress. Two entries
have been rediscovered from scratch over several hours in one session because the search was
skipped. Both were already written down.
If the search returns nothing and the problem is then solved, write the finding back. An empty
result is not "nothing to learn" — it is the reason the next person repeats the work.
## Steps
1. Read `00-report.md`. Move `status:` to `diagnosing`.
2. Investigate. Write `01-diagnosis.md` in the same folder: the trail, **dated**, including
what was ruled out and how. Archaeology — which commit, which pull request, which date a
behaviour changed — is the most valuable content here.
3. When the owner is known, set `status: located` and fill `located-in:` with repositories or
modules from [`00-META/repos.md`](../../../00-META/repos.md).
4. On resolution: `status: resolved`, fill `fixed-by:`. If the root cause was a design gap, run
`hal-graduate` for the amendment and fill `amended-design:`.
## Rules
- State evidence, not assertion. A date and a reference outrank a conclusion.
- Record what was ruled out. The next person needs to know where not to look.
- Closed issues are never deleted.
- `wontfix` is legitimate and requires a sentence saying why.
+44
View File
@@ -0,0 +1,44 @@
---
name: hal-graduate
description: Use when a research effort in HQ concludes and becomes design, or when a design must change. Triggers on "graduate this research", "this is decided", "write the ADR", "close the effort", "the design changed".
---
# hal-graduate
Closes a research effort into a decision and a design, or amends an existing design.
**Authoritative playbook:**
[`00-META/process/02-graduation.md`](../../../00-META/process/02-graduation.md) — read
it; this skill adds only the scaffolding.
## Steps
1. **Check against GENESIS** — `mission.md`, `context.md`, `effect.md`. If the conclusion
conflicts, say which is wrong, in writing, before proceeding.
2. **Write the decision record** in `02-DECISIONS/`, next free number, frontmatter per
[`02-DECISIONS/README.md`](../../../02-DECISIONS/README.md). `reconstructed: false` — this one is being taken,
not recovered. Record the rejected options.
3. **Write the design** under `03-DESIGN/01-to-be/` with `layer: to-be`, `status: designed`,
empty `code:`, today's `updated:`, and `decisions:` pointing at the record.
4. **Close the effort** — the research `00-overview.md` gets `status: graduated` and `became:`
pointing at both.
## Amending an existing design
A design changes only through a decision. Write the record first; if it reverses an earlier
one, set that record's `status: superseded` and `superseded-by:` — **never edit its text**.
Then edit the design and set `updated:`.
## When something ships
Write or update the matching `03-DESIGN/00-as-is/` document so it describes what now runs,
including anything that shipped differently from the intent. Set the to-be document's
`status: implemented` and its `code:`. **The to-be document does not move** — both stand.
`status: implemented` must be defensible from the owning repository's main branch. If it
cannot be checked, it is `in-progress`.
## Do not
- Do not move a to-be document into `00-as-is/`.
- Do not edit an as-is document to describe an intention.
- Do not change a decision record's meaning. Supersede it.
+32
View File
@@ -0,0 +1,32 @@
---
name: hal-handoff
description: Use when an HQ design is settled and implementation is about to start in a code repository. Triggers on "start building this", "hand this off", "ready to implement", "who owns this now".
---
# hal-handoff
Hands a settled design to a code repository. **Authoritative playbook:**
[`00-META/process/04-build-handoff.md`](../../../00-META/process/04-build-handoff.md).
## Steps
1. **Confirm it is settled.** `status: designed`, and every claim traceable to a record in
`02-DECISIONS/`. An open question in the text is a reason to run `hal-new-research`, not to build
around it.
2. **Name the owner.** Set `code:` from
[`00-META/repos.md`](../../../00-META/repos.md). If the repository does not exist yet,
add it to `repos.md` in the same change.
3. **Read the as-is counterpart.** What is being replaced must be written down in
`03-DESIGN/00-as-is/` before it is changed. Building against an undocumented as-is is how
shipped behaviour gets lost.
4. **Flip the status** to `in-progress`, `updated:` today.
5. Build in the code repository. HQ never carries implementation.
6. On completion, run `hal-graduate`'s "when something ships" section.
## Non-negotiable
- Every merge is a human checkpoint, without exception.
- **Never open a pull request unprompted.** A permissions list saying it is allowed is not a
request.
- Work in an isolated worktree, never a shared checkout.
- Never push to the main branch.
+47
View File
@@ -0,0 +1,47 @@
---
name: hal-new-issue
description: Use when something is wrong with the HAL mesh at the level of design or governance — a rule enforced by nothing, a stated behaviour that does not happen, a failure the design lets pass silently. Triggers on "this is broken", "open an issue", "that rule isn't enforced", "this reports success and does nothing".
---
# hal-new-issue
Opens a numbered issue. **Authoritative playbook:**
[`00-META/process/03-issues.md`](../../../00-META/process/03-issues.md).
## First, decide it belongs here
| Belongs in `04-ISSUES` | Belongs in the knowledge base |
|---|---|
| The design permits a failure to be silent | How to fix one occurrence of it |
| A documented rule is enforced by nothing | A command that works around it |
| A stated invariant is false in practice | A node-specific quirk |
| Finding the owner needs the whole mesh in view | Symptom → fix, once the answer is known |
**Search the knowledge base first** for the literal symptom text. If the answer is already
there, this is not an issue — it is a lookup. If the answer is a general lesson, it belongs in
both.
## Steps
1. Next free number. Create `04-ISSUES/NNN-short-name/00-report.md` with:
```yaml
---
status: open
opened: YYYY-MM-DD
located-in: []
fixed-by:
amended-design:
---
```
2. Write the symptom **as observed**, in plain terms, with the evidence that it happened —
what was run, what came back, when.
3. Say why it matters beyond the instance. An issue that is only one occurrence is a knowledge
base entry.
4. End with open questions rather than a proposed fix. Diagnosis is a separate step.
## Do not
- Do not name nodes, domains, addresses or paths. This repository is public.
- Do not guess the owner — `located-in:` is filled by diagnosis, not by opening.
+42
View File
@@ -0,0 +1,42 @@
---
name: hal-new-research
description: Use when starting a new research effort in HQ — an idea, technology or approach worth investigating before it is committed to design. Triggers on "research X", "investigate X", "spike X", "should we use X", "is X worth doing".
---
# hal-new-research
Scaffolds a new research effort. **Authoritative playbook:**
[`00-META/process/01-research.md`](../../../00-META/process/01-research.md) — read it;
this skill only does the mechanical setup.
## Steps
1. Find the next free number: list `01-RESEARCH/NNN-*`, take the highest plus one, zero-padded
to three digits.
2. Create `01-RESEARCH/NNN-descriptive-name/` (kebab-case from the topic).
3. Create `00-overview.md` in it with exactly this frontmatter, then a prose summary of **what** is
being investigated, **why**, and **what it touches**:
```yaml
---
status: active
initiated: YYYY-MM-DD
touches: []
became: []
---
```
4. Do the research in further documents in the same folder. Keep the summary current.
## Do not
- Do not restate the status in prose. It lives in frontmatter, in one place.
- Do not fill `became:` while the effort is open — `hal-graduate` sets it at closure.
- Do not skip or reuse a sequence number.
- Do not write into `03-DESIGN` from an open effort.
- Do not name the mesh being observed. Evidence is required; identification is forbidden.
## Closing
An effort never just stops. It closes through `hal-graduate` as `graduated` or `abandoned`,
always with `became:` pointing at what it turned into. Nothing is deleted.
+41
View File
@@ -0,0 +1,41 @@
---
name: hal-status
description: Use when you need a cross-cutting view of where HQ stands — research state, design implementation state, open issues, or the decision-record index. Triggers on "what's the status", "show the ADR index", "where do things stand", "what's in progress", "what's open".
---
# hal-status
Generates a cross-cutting view **from frontmatter**. This is a read-and-render skill, not a
workflow — HQ has **no central status file by design** (decision 35). Every view is
generated on demand and never written back to disk.
## What to read
| Section | Files | Frontmatter |
|---|---|---|
| Research | `01-RESEARCH/NNN-*/00-overview.md` | `status` (`active` / `graduated` / `abandoned`), `initiated`, `touches`, `became` |
| Design | `03-DESIGN/**/*.md` (not READMEs) | `layer` (`as-is` / `to-be`), `status` (`designed` / `in-progress` / `implemented` / `abandoned`), `code`, `updated`, `decisions` |
| Decisions | `02-DECISIONS/NNNN-*.md` | `status` (`proposed` / `accepted` / `superseded`), `date`, `deciders`, `reconstructed`, `superseded-by`, `extends` |
| Issues | `04-ISSUES/NNN-*/00-report.md` | `status` (`open` / `diagnosing` / `located` / `resolved` / `wontfix`), `opened`, `located-in`, `fixed-by`, `amended-design` |
## Steps
1. Read the frontmatter block from every file above — a grep across each tree is enough, no
need to load bodies.
2. Render what was asked as Markdown tables. Group and sort sensibly. The **ADR index** is one
of these views: records ordered by number, with title, date and status, and reconstructed
ones marked.
3. Flag anything inconsistent at the end, as flags — do not silently correct the render:
- a `graduated` or `abandoned` effort with an empty `became:`
- an `implemented` design with an empty `code:`
- a `superseded` record with no `superseded-by:`
- an `open` or `diagnosing` issue with no activity
- a to-be design whose `decisions:` points at a record that does not exist
- an as-is design whose claims are older than its `updated:` date suggests
## Do not
- **Do not write a status file.** Central status files are explicitly rejected. The view is
always generated, always ephemeral. This includes the ADR index — the hand-written one had
already drifted after a single addition, which is why it was removed.
- Do not infer status from prose. Trust only the frontmatter; if it is wrong, flag it.
@@ -0,0 +1,38 @@
---
name: hal-sync-constitution
description: Use after changing a rule in HQ's 00-META/how-we-build.md, to publish the derived constitution page the mesh injects into design sessions. Triggers on "sync the constitution", "publish the rules", "I changed how-we-build", "update the governed page".
---
# hal-sync-constitution
Publishes the derived constitution from its source. **Authoritative playbook:**
[`00-META/process/05-constitution-sync.md`](../../../00-META/process/05-constitution-sync.md).
## What this is
[`00-META/how-we-build.md`](../../../00-META/how-we-build.md) is the **source**. The
knowledge base carries a **derived** page that the mesh injects into every eligible design
session, and against which a check phase can block work.
Two texts stating the same rules drift, and the enforced copy wins by default — so the reasoned
copy quietly stops being true. This skill is the mechanism that prevents that, and it is how
the claim "HQ is the source" is checked.
## Steps
1. Confirm the source change is recorded as a decision. A rule the mesh enforces is
architecturally significant; if there is no record, run `hal-graduate` first.
2. Derive the page: the **rules without the reasoning**. Section numbering is stable — the
orchestrator and the review fragments cite sections by number, so never renumber to tidy.
3. Publish it to the knowledge base under the constitution slug, replacing the body.
4. **Read it back and verify the change is present.** A publish that reported success and did
nothing is exactly the failure class this repository exists to name — do not skip this.
5. Note the sync in the decision record's Consequences.
## Do not
- **Never edit the derived page directly.** An edit there survives until the next sync and then
vanishes, taking its reasoning with it.
- Do not relax a rule in the derived page. Overrides may only tighten.
- Do not report the sync as done without the read-back. If it could not be performed, say so —
an unsynced rule is a rule the mesh does not enforce, whatever the source says.
-28
View File
@@ -1,28 +0,0 @@
# 00-GENESIS
The **northern star**. What HAL is, the environment it runs in, and what changes when it
works. Every research effort and design decision is checked against this folder.
| File | Purpose |
|------|---------|
| [`mission.md`](mission.md) | Vision, mission, and the values that decide arguments |
| [`context.md`](context.md) | The environment — conditions, not aspirations |
| [`effect.md`](effect.md) | What is different when the work is done |
| [`how-we-build.md`](how-we-build.md) | Rules that hold across the mesh, each one earned |
## Rules
- Markdown only.
- **Stable by nature.** Changes here reflect a genuine shift in intent, not iteration.
- Research and design must be traceable back to what is written here.
## Note on `VISION.md`
The repository root carries `VISION.md`, an architecture overview predating this folder.
It is a useful description of *how* the mesh works and should be folded into
[`02-DESIGN`](../02-DESIGN/), not here — GENESIS answers *why*.
It has also drifted: it lists "Symlinks, not copies" as a key design principle, while the
operating rules forbid creating symlinks at all after one caused production data loss.
A founding document contradicting a hard rule is precisely the failure this folder exists
to prevent.
-42
View File
@@ -1,42 +0,0 @@
# How we build
Working notes on the rules that hold across the mesh. Short, and each one earned.
## Name a context after its aggregate, not after a metaphor
`hal/agents` owns **Agent**. A *brain* — memory, thoughts, cognition — is something an
agent **has**, a concept inside the aggregate. It is not a module.
The cost of getting this wrong is visible today: `hal/brain` names the node runtime, so
the most evocative word in the system points at infrastructure, and `hal/cortex`
describes itself as "mesh messaging" in its manifest while the anatomy documentation
calls it the interactive runtime — and it runs on no node at all.
Anatomy makes attractive names and poor boundaries. Name the thing the domain calls it.
## Ubiquitous language is checked, not assumed
If a document states a rule about the mesh, say how the rule is verified. This repository
has a documented requirement that every module exposing tools declares `brain` as a
dependency. Zero modules do. An unenforced rule is indistinguishable from a wrong one,
and costs more, because people believe it.
## Contexts integrate through the record, never through a shared schema
Publish to the stream; do not join across a boundary. Today five domains share one
45-table schema, which is why work that belongs to one context keeps having to be
implemented in another.
## A failed step must stop the steps after it
Scripted work runs as a sequence, and a sequence that continues past a failure does the
next thing in the wrong place. Gate each step on the last: `cd X || exit`, not `cd X`
followed by a newline.
Earned the obvious way. A `git worktree add` failed because the branch name collided with
an existing namespace; the `cd` into that worktree failed too; and the `cp`, `git add` and
`git commit` that followed ran in the shared checkout and committed to local `main`. The
error was printed and scrolled past.
This is the same shape as the faults this refactor exists to remove — a step reported
failure, nothing stopped, and the damage happened somewhere nobody was looking.
+53
View File
@@ -0,0 +1,53 @@
# 00-META
The **northern star**. What HAL is, the environment it runs in, and what changes when it
works. Every research effort and design decision is checked against this folder.
| File / folder | Purpose |
|------|---------|
| [`mission.md`](mission.md) | Vision, mission, and the values that decide arguments |
| [`context.md`](context.md) | The environment — conditions, not aspirations |
| [`effect.md`](effect.md) | What is different when the work is done |
| [`how-we-build.md`](how-we-build.md) | The rules that hold across the mesh, each one earned. **The source of the mesh constitution** — the governed page the mesh injects into design sessions is derived from it. |
| [`repos.md`](repos.md) | Where implementation lives, and what each repository owns |
| [`process/`](process/) | The playbooks — how work moves through this repository, for engineers and agents alike |
## Rules
- Markdown only.
- **Stable by nature.** Changes here reflect a genuine shift in intent, not iteration. The one
exception is `how-we-build.md`, which changes whenever a rule is earned — and only through
its amendment process.
- Research and design must be traceable back to what is written here.
- **Instance-agnostic.** These documents describe the mesh as a concept. No machine names, no
counts, no topology.
## On the architecture overview in the code repository
The code repository carries an architecture overview predating this folder. It is a useful
description of *how* the mesh works, and its content now lives — anonymised and checked against
the implementation — in [`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/). GENESIS answers *why*;
that document answered *how*, which is the design layer's job.
It had also drifted from the implementation in ways worth recording, since both were found by
comparing it against the code rather than by anyone noticing:
- It described the pipeline as having a separate builder process and a build stage that
packages. Neither was true after 2026-08-04; the documents stayed stale until 2026-08-06
([ADR 0014](../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)).
- It listed the mesh as spanning a fixed number of named machines, which is exactly the
content this repository cannot carry.
It also lists **"symlinks, not copies" as a key design principle**, and that is a genuine
contradiction rather than a stale detail. The mesh's stated intent is that it creates no
symlinks at all — the rule is not merely "only the installer may link", and a founding document
elevating linking to a principle points the opposite way from where this is going.
What exists today is that the installer owns and reconciles every link
([ADR 0011](../02-DECISIONS/0011-the-installer-owns-linking.md)) — an as-is fact, recorded in
[`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md).
Centralising who may link narrowed the incident class; it did not close it. The intent is to
remove the mechanism, recorded as [ADR 0018](../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md).
A founding document contradicting the direction of travel is precisely the failure this folder
exists to prevent.
@@ -1,3 +1,8 @@
---
status: canonical
updated: 2026-08-22
---
# Engineering Context
The conditions the mesh is built for. Properties, not an inventory — no node here is
@@ -1,3 +1,8 @@
---
status: canonical
updated: 2026-08-22
---
# Effect
Imagine the mesh works as intended. What is different?
+190
View File
@@ -0,0 +1,190 @@
---
status: canonical
updated: 2026-08-23
derives: knowledge-base constitution page
decisions:
- 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md
---
# How we build
The rules that hold across the mesh. Short, and each one earned.
**This document is the source of the mesh constitution.** The governed page the mesh injects
into design sessions is *derived* from it, section for section, and carries the rules without
the reasoning. Never edit that page directly — an edit there survives until the next sync and
then vanishes, taking its reasoning with it. The sync is playbook
[`process/05-constitution-sync.md`](process/05-constitution-sync.md), and it is how the claim
"HQ is the source" is checked.
Section numbers are stable. The orchestrator and the review fragments cite them.
---
## 1. Purpose
These are the principles and guardrails every piece of work in the mesh is checked against —
design sessions, analysis gates, reviews, and agents acting on their own. A rule stated here is
non-negotiable unless amended per §6.
They exist because each was violated first. Where a rule reads as arbitrary, that is a sign the
incident behind it is not written down, and the fix is to write it down, not to relax the rule.
---
## 2. Non-negotiables
| Rule | What it means |
|---|---|
| **Never write to a production database directly** | No insert, update, delete or schema statement executed against production by hand. Schema changes go through numbered migrations; data changes go through application code or the module's own capabilities. Raw statements skip every side effect the proper path has — events, audit, cache invalidation, fan-out. |
| **Every schema change is a migration** | Numbered, in the module's own language, compiled with it. Both a baseline for a fresh installation *and* an incremental migration for installations that already exist. If code references a column, the migration creating it must exist. [ADR 0006](../02-DECISIONS/0006-schema-changes-are-numbered-migrations.md) |
| **Never bypass the pipeline** | No manual database edit, no manual restart as a workaround. Fix the cause and deploy. A workaround that works is a workaround that is never removed, and the next person cannot tell the node from its declaration. |
| **Never create a symlink** | A hand-made link caused production data loss through container volume resolution, and the judgement needed to make a safe exception is exactly the judgement unavailable at the moment it matters. Today the installer owns and reconciles the links the mesh still uses [ADR 0011](../02-DECISIONS/0011-the-installer-owns-linking.md); the intent is that the mesh creates none at all [ADR 0018](../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md). Neither reading permits you to make one. |
| **Never push directly to the main branch** | Branch, push, review, merge. Every merge is a human checkpoint, without exception — **including in this repository**. A documentation repository is not a lower tier of care; a decision record lands the same way a service does. |
| **One change per pull request, and never merge your own** | Unrelated improvements bundled together cannot be reviewed or reverted separately. Self-merging removes the checkpoint that is the entire point. |
| **Never open a pull request unprompted** | A permissions list saying it is allowed is not a request. |
| **A failed step fails the job** | A sequence that continues past a failure does the next thing in the wrong place. Gate each step on the last. [ADR 0008](../02-DECISIONS/0008-a-failed-step-fails-the-job.md), and §5. |
### A failed step must stop the steps after it — how it was earned
A worktree creation failed because the branch name collided with an existing namespace. The
change into that worktree failed too. The copy, the staging and the commit that followed all
ran in the shared checkout and committed to a local main. The error was printed and scrolled
past.
This is the same shape as the faults the mesh's whole refactor exists to remove: a step
reported failure, nothing stopped, and the damage happened somewhere nobody was looking.
---
## 3. Module and infrastructure rules
### Manifests
- **Never bump a version by hand.** The builder owns versioning. A version change in a diff is
a defect; revert it.
- **Features are detected, not declared.** The installer discovers what a module carries from
what its directory contains. A declared list and the directory it describes drift, and the
directory is the one that is true.
- **Every runtime variable is declared.** A variable the module reads and the manifest does not
declare is invisible to the mesh: it will not be generated, injected, or audited.
- **Provisioned credentials arrive through declared requirements**, never hardcoded in code,
compose files or scripts. [ADR 0005](../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)
- **Never install a package by hand.** A package is declared in the manifest and arrives the
way every other package does. A hand-installed package is invisible to the mesh: it is not
declared, not reproduced on the next node, and not present after a rebuild — and the node
works until it doesn't.
### Placement
- **Core modules belong to the monorepo** — the runtime, delivery, provisioning, configuration,
knowledge, and the shared infrastructure the mesh provisions against.
- **Every standalone application gets its own repository**, with a manifest at its root,
registered as a build source. Creating an application directory in the monorepo is a
convention violation and reviewers reject it.
[ADR 0010](../02-DECISIONS/0010-applications-live-in-their-own-repository.md)
### Migrations
- Numbered, idempotent, and safe to re-run. Guard every statement.
- The initial migration is **frozen** once it has run anywhere. Change is a new number.
- Numbers are unique. A duplicate prefix is a defect, not a style question.
### Managed files
**A file edited on a node is a bug with a delay on it.** Everything under the mesh's managed
surface is regenerated from the mesh database; a local edit survives one synchronisation and is
then silently overwritten, bringing back whatever it fixed. Use the mesh operation that owns
the value. If unsure whether a file is managed, ask the tooling — the answer is not visible
from the file. [ADR 0004](../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)
---
## 4. Naming and boundaries
### Name a context after its aggregate, not after a metaphor
A context named `agents` owns **Agent**. A *brain* — memory, thoughts, cognition — is something
an agent **has**: a concept inside the aggregate, not a module.
The cost of getting this wrong is visible today. An anatomy name points at the node runtime, so
the most evocative word in the system names infrastructure; another describes itself as "mesh
messaging" in its manifest while the anatomy documentation calls it the interactive runtime —
and it runs on no node at all.
Anatomy makes attractive names and poor boundaries. Name the thing the domain calls it.
### Group by domain, not by single function
A module is a purpose, not a piece of software. Four modules that together constitute "how a
node is reachable" and cannot be assigned, versioned or replaced as one thing are four
accidents, not four boundaries.
[ADR 0017](../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)
### Contexts integrate through the record, never through a shared schema
Publish to the stream; do not join across a boundary. Today several domains share one
forty-five-table schema, which is why work belonging to one context keeps having to be
implemented in another.
---
## 5. Evidence and verification
### Ubiquitous language is checked, not assumed
**If a document states a rule about the mesh, it says how the rule is verified.** This
repository has a documented requirement that every capability-exposing module declare the core
runtime as a dependency. Zero modules do.
An unenforced rule is indistinguishable from a wrong one, and costs more, because people
believe it.
### Behavioural criteria require runtime evidence
A criterion of the form *"the script runs"*, *"the endpoint answers"*, or *"the migration
applied"* is satisfied only when the change has actually been exercised: a real run, a real
request, a pipeline log, a live query showing the expected result.
**Marking a runtime criterion verified from a diff is itself a violation.** A reviewer who
finds one names the evidence required and returns the work.
The reason is the mesh's most consistent failure shape: a green result proves transport, not
effect. Absence reads as success unless something looked.
### Search the record before forming a hypothesis
The first action on any error message, failing service or unexpected behaviour is to search the
operational memory for the literal error text — before a hypothesis, not after one fails. The
knowledge base is indexed on symptoms.
This fires hardest on *familiar* ground, where a confident trail feels like progress. Two
entries were each rediscovered from scratch over several hours in a single session because the
search was skipped. Both were already written down.
---
## 6. Amendment
This document is governed. It does not change by commit message or unilateral decision.
1. **Propose** — a change stating what rule is changing, why the current wording is
inadequate, and what reviewed it.
2. **Review** — sign-off by reviewers who are not the proposer.
3. **Record** — the change is a decision and gets a record in [`02-DECISIONS/`](../02-DECISIONS/), because a
rule the mesh enforces is architecturally significant.
4. **Sync** — playbook [`process/05-constitution-sync.md`](process/05-constitution-sync.md)
publishes the derived page. An unsynced rule is a rule the mesh does not enforce, whatever
this document says.
No drive-by edits. Every change traces to a recorded decision.
---
## 7. Overrides
A team or product may define additional constraints that **narrow or tighten** these rules.
They may never relax them.
An override says which rule it tightens, or which gap it fills, and follows the same amendment
process. Absence of an override means these rules apply unmodified.
@@ -1,3 +1,8 @@
---
status: canonical
updated: 2026-08-22
---
# Mission
## Vision
+60
View File
@@ -0,0 +1,60 @@
# Process — overview
How work moves through HQ, and who may do what. Every other document in this folder is a
playbook: trigger, who runs it, steps, outputs. Engineers and agents follow the same
playbooks; agents must not act outside them.
## The audiences
| Audience | Contract |
|---|---|
| **Engineers** | Read and write everything. HQ is the single source of truth for mission, research, design, decisions and issue diagnosis. |
| **Agents** | The same rights as engineers, exercised through these playbooks. |
| **Anyone else** | This repository is public and written for them, but it is not a support channel. Nothing here identifies the mesh it describes. |
## The knowledge flow
```
idea ──► 01-RESEARCH ──► decision (02-DECISIONS/) ──► 03-DESIGN/01-to-be ──► built (code repo)
│ │ │
│ │ └─► 03-DESIGN/00-as-is once shipped
│ └────► abandoned (recorded, kept)
└─(small/obvious, still recorded in 02-DECISIONS)──────────► 03-DESIGN directly
symptom ──► 04-ISSUES ──► diagnosis ──► code-repo fix and/or design amendment
how-we-build.md ──► constitution sync ──► knowledge base ──► injected into design meetings
```
## The two design layers
`03-DESIGN` holds two layers that are never mixed:
| Layer | What it is | Changes when |
|---|---|---|
| `00-as-is/` | The mesh that exists today. Describes shipped behaviour, including behaviour nobody would choose again. | Something ships, or an as-is claim is found to be wrong. |
| `01-to-be/` | The mesh being built toward. Every statement traceable to a record in `02-DECISIONS/`. | A decision is taken or amended. |
A to-be document that ships does not move. Its as-is counterpart is written or updated, the
to-be document's frontmatter goes to `implemented`, and both stand — one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning, which is
the expensive half.
## The playbooks
| # | Playbook | Trigger |
|---|---|---|
| [01](01-research.md) | Research | An idea worth investigating before committing to design |
| [02](02-graduation.md) | Graduation & design change | Research concludes, or a design must change |
| [03](03-issues.md) | Issues | Something is wrong — often with the owner unknown |
| [04](04-build-handoff.md) | Build handoff | A design is ready to be built |
| [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces |
## Status lives in frontmatter
Research overviews, design docs, issue reports and decision records each carry their status as
YAML frontmatter (schemas in the section READMEs and playbooks). There are **no central status
files** and no decision ledger. **Every decision is a record in
[`02-DECISIONS`](../../02-DECISIONS/)** — if it is worth recording it is worth a record, and if
it is not worth a record it is not recorded. Cross-cutting views, the decision index included,
are generated on demand by the `hal-status` skill and never written to disk.
+48
View File
@@ -0,0 +1,48 @@
# Playbook 01 — Research
**Trigger.** An idea, technology or approach worth investigating before it is committed to
design. Also: an as-is document that raises a question nobody can answer.
**Who runs it.** Anyone.
## Steps
1. Take the next free number: highest `01-RESEARCH/NNN-*` plus one, zero-padded to three
digits. Never skip or reuse a number.
2. Create `01-RESEARCH/NNN-descriptive-name/` and a `00-overview.md` in it carrying:
```yaml
---
status: active
initiated: YYYY-MM-DD
touches: [] # areas, subsystems or as-is design docs the effort bears on
---
```
Then a short prose summary: **what** is being investigated, **why**, and **what it
touches**.
3. Do the research in further documents in the same folder — notes, option analyses, evidence,
draft designs. Anything goes. Keep the summary in `00-overview.md` current as the effort changes
shape.
## What makes research worth reading
State evidence, not assertion. *"Zero of 124 modules declare `brain` as a dependency"*
outranks *"the dependency rule is not followed"*. An effort that measured nothing has not
finished.
Research describes real observations but never identifies the mesh it observed. The shape of a
finding survives anonymisation intact — *a node publicly named but behind a household NAT*
carries the whole lesson without naming anything.
## Do not
- Do not put status in prose. It lives in `00-overview.md`'s frontmatter, and the prose must not restate it.
- Do not set `became:` while the effort is open — playbook 02 sets it at closure.
- Do not write into `03-DESIGN` from an open effort.
## Closing
An effort never just stops. It closes through playbook [02](02-graduation.md) as `graduated`
or `abandoned`, always with `became:` pointing at what it turned into. Nothing is deleted —
what was rejected, and why, is the more expensive half to rediscover.
+57
View File
@@ -0,0 +1,57 @@
# Playbook 02 — Graduation and design change
**Trigger.** A research effort concludes, or an existing design must change.
**Who runs it.** Anyone, with the decision recorded before the design moves.
## Graduating research
1. **Check it against GENESIS.** An effort graduates only if its conclusion is traceable to
[`mission.md`](../mission.md), [`context.md`](../context.md) and
[`effect.md`](../effect.md). If it conflicts, either the effort is wrong or GENESIS is —
say which, in writing, before proceeding.
2. **Record the decision.** Write a record in [`02-DECISIONS/`](../../02-DECISIONS/) taking the next free
number. Format and rules are in [`02-DECISIONS/README.md`](../../02-DECISIONS/README.md). State evidence,
not assertion, and record the options that were rejected — that is the half worth keeping.
3. **Write the design.** Create the document under `03-DESIGN/01-to-be/` with frontmatter:
```yaml
---
layer: to-be
status: designed
code: [] # owning code repo(s); set at build handoff, empty before
updated: YYYY-MM-DD
decisions: [02-DECISIONS/NNNN-....md]
---
```
4. **Close the effort.** Set the effort's `00-overview.md` frontmatter to `status: graduated` and
`became:` pointing at the design document and the decision record.
## Amending an existing design
A design changes only through a decision.
1. Write the decision record. If it reverses an earlier one, the earlier record's `status:`
becomes `superseded-by: 02-DECISIONS/NNNN-....md` — **its text is never edited**.
2. Edit the to-be design document and set `updated:` to today.
3. If the amendment came from an issue, set that issue's `amended-design:` to the document
path.
## When something ships
Implementation state is a third axis, independent of both design and decision.
1. Write or update the matching document under `03-DESIGN/00-as-is/` so it describes what now
runs — including anything that shipped differently from the intent. A design that shipped
bent is an as-is fact, not a design amendment.
2. Set the to-be document's `status: implemented` and its `code:` to the owning repositories
from [`repos.md`](../repos.md).
3. `status: implemented` must be defensible from the owning repository's main branch, not from
intent. If it cannot be checked, it is `in-progress`.
## Do not
- Do not move a to-be document into `00-as-is/`. Write the as-is document; both stand.
- Do not edit an as-is document to describe an intention. That is what the to-be layer is for.
- Do not change a decision record's meaning. Supersede it.
+47
View File
@@ -0,0 +1,47 @@
# Playbook 03 — Issues
**Trigger.** Something is wrong at the level of the mesh's design or governance — a rule that
turns out to be unenforced, a stated behaviour that does not happen, a silent failure the
design permits.
**Who runs it.** Anyone may open an issue. No localisation is required to report one.
## What belongs here, and what does not
| Belongs in `04-ISSUES` | Belongs in the knowledge base |
|---|---|
| The design permits a failure to be silent | How to fix one occurrence of it |
| A documented rule is enforced by nothing | A command that works around it |
| A stated invariant is false in practice | A node-specific quirk |
| The owning component is unknown and finding it needs the whole mesh in view | Symptom → fix, once the answer is known |
The knowledge base already holds the operational record and is indexed on symptoms. This
folder is not a second copy of it. An issue here is a question HQ must **answer**, not an
incident someone must **clear**.
## Steps
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`:
```yaml
---
status: open
opened: YYYY-MM-DD
located-in: [] # owning repo(s)/module(s), filled by diagnosis
fixed-by: # PR or commit reference, filled at resolution
amended-design: # design doc path, when the root cause was a design gap
---
```
Then the symptom **as observed**, in plain terms, with the evidence that it happened.
2. Investigate in `01-diagnosis.md` in the same folder — the trail, dated, including what was
ruled out. Move `status:` to `diagnosing`, then `located` once the owner is known.
3. Resolve. Set `status: resolved`, fill `fixed-by:`, and if the root cause was a design gap,
run playbook [02](02-graduation.md) and fill `amended-design:`.
## Rules
- Closed issues are never deleted — they are the mesh's symptom-to-component memory.
- An issue whose answer is a general lesson should also be written to the knowledge base, so
the next person searching a symptom finds it. Both, not either.
- `status: wontfix` is legitimate and requires a sentence saying why.
+30
View File
@@ -0,0 +1,30 @@
# Playbook 04 — Build handoff
**Trigger.** A to-be design is settled and work is about to start in a code repository.
**Who runs it.** Whoever starts the build.
## Steps
1. **Confirm the design is settled.** Its frontmatter reads `status: designed`, and every
claim in it traces to a record in [`02-DECISIONS/`](../../02-DECISIONS/). An open question in the text is a
reason to run playbook [01](01-research.md), not to start building around it.
2. **Name the owner.** Set `code:` in the design's frontmatter to the repositories from
[`repos.md`](../repos.md). If the repository does not exist yet, add it to `repos.md` in
the same change.
3. **Check the as-is.** Read the matching `03-DESIGN/00-as-is/` document. What is being
replaced is stated there; if it is not, write it before changing it. Building against an
undocumented as-is is how a shipped behaviour gets lost.
4. **Flip the status.** `status: in-progress`, `updated:` today.
5. **Build in the code repository.** HQ is not a code repository and never carries
implementation.
6. **On completion**, run the "when something ships" section of playbook
[02](02-graduation.md).
## Rules
- Every merge is a human checkpoint, without exception.
- Never open a pull request unprompted. A permissions list saying it is allowed is not a
request.
- Work in an isolated worktree, never a shared checkout. A failed `cd` in a shared checkout
commits to the wrong branch, and the error scrolls past.
+37
View File
@@ -0,0 +1,37 @@
# Playbook 05 — Constitution sync
**Trigger.** [`how-we-build.md`](../how-we-build.md) changed a rule that the mesh enforces at
runtime.
**Who runs it.** Whoever made the change.
## Why this playbook exists
The mesh injects a constitution into every eligible design meeting; agents check their output
against it and a constitution-check phase can block a meeting. That text is **derived**.
`how-we-build.md` is the source.
Two texts stating the same rules will drift, and the enforced copy winning by default means
the reasoned copy quietly stops being true. This playbook is the mechanism that stops that —
and, per the repository's own rule, it is how the rule "HQ is the source" is checked.
## Steps
1. Edit [`how-we-build.md`](../how-we-build.md). Each rule keeps the reasoning that earned it;
the derived page carries the rule alone.
2. Record the change as a decision — a rule the mesh enforces is architecturally significant.
Playbook [02](02-graduation.md).
3. Publish the derived page to the knowledge base under the constitution slug, replacing its
body. Keep the section numbering stable: the meeting orchestrator and the review fragments
cite sections by number.
4. Verify the derived page reads back with the change present. A publish that reported success
and did nothing is exactly the failure class this repository exists to name.
5. Note the sync in the decision record's Consequences.
## Rules
- **Never edit the derived page directly.** An edit there survives until the next sync and
then vanishes, taking its reasoning with it.
- The derived page may only be **tightened** by per-team override pages, never relaxed.
- If the sync cannot be performed, say so in the record. An unsynced rule is a rule the
mesh does not enforce, whatever `how-we-build.md` says.
+51
View File
@@ -0,0 +1,51 @@
---
status: canonical
updated: 2026-08-23
---
# The Novox repositories
The map of where implementation lives. Humans use it for orientation; agents use it for issue
triage (playbook [`process/03-issues.md`](process/03-issues.md)). The `code:` frontmatter
field in design documents points at entries here.
Repository *names* are recorded; hosts, URLs and owners are not — this repository is public,
and a forge address is an operational detail (see [`README`](../README.md)).
| Repository | Owns |
|---|---|
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0028](../02-DECISIONS/0028-hq-is-company-scoped.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
## What lives where inside the monorepo
Named by role, because the layout is itself part of the as-is design — see
[`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/).
| Area | Holds |
|---|---|
| Module catalogue | One directory per module, each with a manifest. Core modules sit under the mesh's own namespace; everything else at the top level. |
| Node runtime | The daemon and interactive runtime that every node runs. |
| Bootstrap scripts | First-node initialisation, joining an existing mesh, and node rescue. |
| Shared library | The SDK every module builds against. |
| Pipeline test harness | End-to-end coverage of the delivery pipeline. Currently unbuildable — see [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md). |
## Why applications do not live in the monorepo
A standalone application in the monorepo is a convention violation, and reviewers reject it.
The reasoning is recorded in [`02-DECISIONS/0010`](../02-DECISIONS/0010-applications-live-in-their-own-repository.md):
the mesh installs, provisions for, and ships an application through exactly the same machinery
whether or not its source sits beside the mesh's own — so co-location buys nothing and costs
the monorepo's review cadence.
## There is no npm workspace
Each module is a standalone package that consumes its dependencies from the private registry,
not from a sibling directory. The workspace was removed after it caused build-versus-development
divergence — a workspace member importing another resolved to local unbuilt source in the
pipeline and to a published version in development. Recorded in
[`02-DECISIONS/0007`](../02-DECISIONS/0007-no-npm-workspace.md).
Consequence, and it is a real one: a cross-package change is two steps — publish, then consume
— and a repository-wide `npm install` does not exist.
@@ -1,6 +1,12 @@
---
status: active
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/02-modules-and-manifests.md, 03-DESIGN/00-as-is/10-module-catalogue.md, 03-DESIGN/01-to-be/00-work-breakdown.md]
became: [02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md]
---
# 001 — Module domain decomposition
- **Status:** ONGOING — ADR 0001 accepted; graduates when `02-DESIGN` carries the per-context specifications
- **Initiated by:** jochen, 2026-08-22
- **Areas touched:** every `hal/*` and `noxflow/*` module; the pipeline's dependency
graph; the knowledge base; agent identity and credentials.
@@ -44,3 +50,18 @@ boundary fault:
## Open questions
Tracked in [`analysis.md`](analysis.md) under "Open questions".
## Deliberately not decided
Recorded so they are not mistaken for oversights. Each is open, and each comes out of
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md); this effort stays
`active` until they are answered.
| Question | Status |
|---|---|
| `hal/scheduler` — infrastructure, or part of the work context. | Open. |
| Which context owns the executor. | Open. |
| Catalogue destination — one repository or many. | Open. Phase 4. |
| What the shared library keeps after extraction. | Open. Phase 3. |
| Where human agent modality is recorded — which user, on which node, a human agent acts as. | Open. Required by the model; not yet stored. |
| Which domains the modules outside the platform core group into. | Open, from [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md), which settles the principle and deliberately not the list. |
@@ -1,3 +1,8 @@
---
effort: 001-module-domain-decomposition
updated: 2026-08-22
---
# Current state → ideal state
## 1. The count is a symptom
@@ -232,6 +237,6 @@ returns. That mechanism is the subject of a separate ADR.
pipeline resolves dependencies across the registry rather than the filesystem?
4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or
does that belong to `hal/stream`? Everything imports it, which argues both ways.
5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which
5. **Human agent modality.** ADR 0015 requires a fact the mesh does not record: which
user, on which node, a human agent acts as. Where does it live — an attribute of the
agent, or of the agent-node binding?
@@ -1,6 +1,12 @@
---
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/04-delivery.md, 03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
---
# 002 — A mesh that runs locally
- **Status:** GRADUATED — the design is [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md)
- **Initiated by:** jochen, 2026-08-22
- **Areas touched:** `install.d/`, `hal/meshware`, `hal/coordinator`, `hal/brain`,
`hal/developer` (`dev_up`), `hal/sdk` (env generation, feature handlers, artifact
@@ -9,7 +15,7 @@
## Summary
Phase 0 of [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) requires
Phase 0 of [`03-DESIGN/00-work-breakdown.md`](../../03-DESIGN/01-to-be/00-work-breakdown.md) requires
a mesh that comes up in containers, runs its own pipeline, and reproduces known faults on
demand. Nothing else in the decomposition starts until it exists, because every fault the
decomposition addresses was found in production — there was nowhere else to find it.
@@ -18,7 +24,7 @@ This effort establishes what already runs in a container, what is welded to the
what it would take to close the gap. It does **not** choose an approach: the central
question — how a containerised node executes a module service, when a module service is
defined today as a systemd unit shelling to `docker compose` in `/services/` — is not
answered by ADR 0001 and is recorded below rather than decided.
answered by ADR 0015 and is recorded below rather than decided.
## What was established
+9 -4
View File
@@ -1,3 +1,8 @@
---
effort: 002-local-mesh
updated: 2026-08-22
---
# A mesh that runs locally — current state and obstacles
Evidence for Phase 0. Every claim here is either a file location or something measured on
@@ -231,7 +236,7 @@ over either way.
a local-mesh implementation detail — if the answer is that the mesh should own
supervision, Phase 0 should not build a container-only workaround first. Under
investigation; findings will land in [`003-service-supervision`](../003-service-supervision/)
and, if it goes ahead, ADR 0002.
and, if it goes ahead, a decision record of its own.
2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable
root keeps one code path; container-specific targets keep the host paths untouched. This
@@ -239,7 +244,7 @@ over either way.
3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap
scripts, and it will run one OS where the real mesh is heterogeneous by design
(`00-GENESIS/context.md`). Stating the divergence up front is what stops "it works
(`00-META/context.md`). Stating the divergence up front is what stops "it works
locally" from becoming its own class of silent failure. Now sharper, because a
development environment people use daily is trusted far more than a rig — and drifting
from production costs correspondingly more.
@@ -248,9 +253,9 @@ over either way.
## References
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this
- [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — the decision this
phase unblocks
- [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) — Phase 0 tasks
- [`03-DESIGN/00-work-breakdown.md`](../../03-DESIGN/01-to-be/00-work-breakdown.md) — Phase 0 tasks
and checkpoint
- `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open
- `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped
@@ -1,6 +1,12 @@
---
status: active
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: []
---
# 003 — Who supervises a service
- **Status:** ONGOING — evidence gathered, options costed, decision open.
**No longer blocks Phase 0** (see below).
- **Initiated by:** jochen, 2026-08-22, in response to
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
@@ -34,7 +40,7 @@ This effort answers the cost half. It does not choose.
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
production then have the same shape rather than a translation layer between them.
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
- **It cannot be all-or-nothing**, and ADR 0015 already says why: a human agent acts through a
shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes **does not
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
@@ -43,5 +49,5 @@ Detail and costs in [`analysis.md`](analysis.md).
## Decision needed
Which supervision model the mesh adopts, recorded in ADR 0002 before Phase 0 builds
Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds
anything. The options and their costs are in `analysis.md` under "Options".
@@ -1,3 +1,8 @@
---
effort: 003-service-supervision
updated: 2026-08-22
---
# Who supervises a service — the cost of leaving systemd
Measured 2026-08-22. Every claim is a file location or a count.
@@ -123,10 +128,10 @@ What it costs, honestly:
---
## 5. Why it cannot be all-or-nothing — and ADR 0001 already says so
## 5. Why it cannot be all-or-nothing — and ADR 0015 already says so
Some of what runs under systemd today **cannot** be containerised, and the reason is
already in the domain model. ADR 0001:
already in the domain model. ADR 0015:
> a non-human agent acts through a spawned session — a human agent acts through a shell or
> desktop
@@ -191,7 +196,7 @@ hand.
This is worth recording for two reasons. It weakens any argument that systemd-adjacent
self-healing is already wired — it is not. And it is another instance of the pattern this
whole refactor is about: **a documented mechanism that does not exist, believed because it
was written down.** `00-GENESIS/how-we-build.md` calls this out as a rule; here it is again,
was written down.** `00-META/how-we-build.md` calls this out as a rule; here it is again,
found by grep.
---
@@ -203,7 +208,7 @@ found by grep.
**This no longer gates Phase 0.** When this was written, the local mesh was assumed to be
built from application containers, which forced the question — there is no natural way to
run an init system inside one. The decision of 2026-08-22 to build development nodes as
**system containers** (see [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md))
**system containers** (see [`03-DESIGN/01-end-to-end-testing.md`](../../03-DESIGN/01-to-be/01-end-to-end-testing.md))
removes that pressure entirely: a system container runs a real init, so the existing model
works unmodified and the lab needs no answer here to exist.
@@ -218,7 +223,7 @@ fate-sharing reason in §3.
## References
- [`002-local-mesh`](../002-local-mesh/analysis.md) — the effort this came out of
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — agent modality, which
- [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — agent modality, which
decides what cannot leave the host
- `modules/hal/meshware/daemon/src/cerebellum.ts:815-828` — the self-restart workaround
- `modules/hal/meshware/systemd/hal-module@.service` — the per-module Docker lifecycle
@@ -1,6 +1,12 @@
---
status: active
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/01-mesh-and-transport.md, 03-DESIGN/01-to-be/01-end-to-end-testing.md]
became: [02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
---
# 004 — Reproducing the mesh network in a lab
- **Status:** ONGOING — topology established and mapped; not yet stood up
- **Initiated by:** jochen, 2026-08-22 — *"the most difficult part of our VM setup will be
the networking part"*
- **Areas touched:** `modules/wireguard`, `modules/dnsmasq-app`, `modules/traefik`,
+6 -1
View File
@@ -1,3 +1,8 @@
---
effort: 004-lab-network
updated: 2026-08-22
---
# Reproducing the mesh network in a lab
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
@@ -194,5 +199,5 @@ Several manifests declare `scope: public` on firewall rules — `wireguard`, `tr
`modules/unifi/module.yml:52-93` does deliberately.
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is
This is the same shape as the rule in `00-META/how-we-build.md` — *an unenforced rule is
indistinguishable from a wrong one, and costs more, because people believe it.*
@@ -0,0 +1,61 @@
---
status: active
initiated: 2026-08-23
touches:
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
- 03-DESIGN/00-as-is/10-module-catalogue.md
- 03-DESIGN/00-as-is/02-modules-and-manifests.md
became: []
---
# 005 — Which domains the catalogue groups into
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) settles
that modules outside the platform core are grouped by domain rather than by single function,
and deliberately does not settle the list. This effort settles the list — and, first, tests
whether the premise survives measurement.
## What is being investigated
Eighty-nine modules sit outside the platform core. The question is which of them belong
together, and the method is evidence rather than intuition: **which modules actually change
together**, measured across the full history of the code repository.
## Why
The argument in ADR 0017 is that the catalogue's shape records what was installed rather than
what anything is for — that four modules constituting "how a node is reachable" have no
relationship the mesh can see, so a change to connectivity is made four times.
That argument is testable. If those modules genuinely change together, the grouping is
justified by more than tidiness. If they do not, the premise needs revising before a list is
drawn from it.
## What it touches
The catalogue's shape, the manifest, and — for one candidate grouping — the provisioning
reference itself, since a requirement names a provider **module**. Grouping providers would
change what a consumer names.
## Status
Measurement is done and is in [`analysis.md`](analysis.md). It **partly contradicts the
premise**, in a way that narrows the effort usefully:
- Non-platform modules overwhelmingly change **alone** — 10% of commits touch more than one,
and 50 of 89 never co-change with anything.
- Two clusters do exist. One of them, reachability, holds up as a domain.
- The other, the provisioned infrastructure providers, co-changes for a reason that argues
**against** grouping rather than for it.
The remaining work is the list itself, for the modules where grouping is justified, plus the
open questions below.
## Open questions
| Question | Why it is open |
|---|---|
| Whether provider modules group at all, and if so what a consumer's requirement names instead of a module. | The provisioning reference is load-bearing; getting it wrong is expensive. Opinion and evidence in the analysis; not yet decided. |
| Whether applications group into domains now and leave the monorepo later as a unit, or leave first. | Decided in principle — group first, then split — but the migration order has real cost either way. |
| What to do with the ~50 modules that co-change with nothing. | The evidence gives no grouping signal for them at all. That may mean they are correctly sized already. |
| Whether "group or leave" is even the right pair of options. | [Research 006](../006-mesh-from-scratch/code-skeleton.md) finds a third fate — **absorbed into the node host**, ceasing to be a module at all — and argues it is the correct answer for the reachability cluster this effort measured. If so, the cluster this effort found is evidence for absorption rather than for grouping. |
+144
View File
@@ -0,0 +1,144 @@
---
effort: 005-domain-grouping
updated: 2026-08-23
---
# Which modules actually change together
## Method
Every commit in the code repository's main branch that touches the module catalogue, reduced
to the set of modules it touched. Platform-namespace modules are excluded — their
decomposition is settled by
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md). Modules that no
longer exist are excluded, because pre-rename names dominate the raw signal and describe a
catalogue nobody works in.
Commits touching more than eight modules are excluded from pair counting: a sweep across
forty modules says the sweep happened, not that the modules are related.
## Finding 1 — these modules overwhelmingly change alone
| Measure | Value |
|---|---|
| Non-platform modules today | 89 |
| Commits touching at least one | 312 |
| Commits touching two to eight together | 33 — **10%** |
| Modules that never co-change with anything | **50 of 89** |
Nine commits in ten touch exactly one module. More than half the catalogue has never been
edited alongside anything else.
This is the first thing a grouping proposal has to survive, and it is not what the premise
predicts.
## Finding 2 — two clusters exist
Everything above two co-changes falls into one of two groups.
**The provisioned infrastructure providers.** The densest cluster: the relational stores with
each other and with the cache, the broker, the object store and the image registry — a
strongly connected set, six to two co-changes per pair.
**Reachability.** The reverse proxy with the resolver, the firewall and the VPN; the firewall
with the intrusion filter. Three to two co-changes per pair.
Outside those: a torrent client with the VPN, twice. A login manager with a media player,
twice. Nothing else reaches two.
## Finding 3 — the provider cluster is manifest churn, not cohesion
This is the result that matters, and it is only visible by reading the commits rather than
counting them.
Every multi-provider commit is a **cross-cutting change to the machinery, applied N times**:
| Date | What the commit did |
|---|---|
| 2026-04-15 | modules own their tools — decentralised tool serving |
| 2026-04-24 | filesystem-driven feature detection; manifest boilerplate removed |
| 2026-04-27 | remove runtime dependencies that were only build dependencies |
| 2026-05-01 | hooks move into their own directory with their own package |
| 2026-05-05 | hooks renamed to a stage-and-feature convention; legacy install scripts removed |
| 2026-08-06 | verify **by shape in the SDK**, not by copying a script into every module |
| 2026-08-11 | bind declared volume paths that leaked an anonymous volume per deploy |
| 2026-08-23 | scope services to the local network and mesh; respect the profile gate |
**Not one of them is a change to what a database is.** They are all "the manifest contract
changed, therefore every manifest changed".
Two consequences follow.
Merging the providers into one module would not have prevented a single one of these commits.
It would have made the same edit land in one file rather than six — which is a diff-size
improvement, not a boundary.
And the history already shows the correct fix being applied, repeatedly and successfully:
*verify by shape in the SDK, not by copying a script into every module*. When the same edit
must be made in every provider, the answer that worked was **moving the concern into the
machinery**, not merging the modules that carry it.
Co-change here measures coupling to the manifest contract. It does not measure domain
cohesion, and using it as though it did would group the catalogue by which modules are most
boilerplate-heavy.
## Finding 4 — reachability holds up
The same reading applied to the reachability cluster gives a different answer. Most of its
twenty-five multi-module commits are the same machinery churn — but not all, and the
remainder are genuine:
| Date | What the commit did |
|---|---|
| 2026-07-29 | let the mesh's proxy coexist with another listener on the same port |
| 2026-08-03 | derive public split-DNS from node accessors so requests stop hairpinning |
| 2026-08-06 | firewall mesh-only by default, public by declaration |
Each is one intent — *change how a node is reachable* — landing across the proxy, the
resolver, the firewall and the VPN together. That is exactly the shape ADR 0017 describes, and
it is the only place in the catalogue where the measurement finds it.
The 2026-08-23 scoping commit is the sharpest case: it spans the reachability cluster **and**
two providers, because "which network is this exposed on" is a reachability question asked of
a database.
## What this means for ADR 0017
The record's principle stands, and its scope needs narrowing. Grouping by domain is:
- **Justified by evidence** for reachability. One intent, several modules, repeatedly.
- **Argued against by evidence** for the providers. The coupling is to the manifest contract,
and the demonstrated fix is to move the concern into the machinery.
- **Unsupported either way** for the fifty modules that co-change with nothing. Silence is not
evidence of independence — many are simply rarely touched — but there is no measured basis
for grouping them, and a proposal that groups them is drawn from intuition. It should say so.
The catalogue's shape is still wrong in the way the record describes. Measurement says
grouping is the right fix in fewer places than the record implies, and that a second fix —
moving cross-cutting concerns into the machinery — accounts for most of what looks like
grouping pressure.
## Recommendation on the provisioning reference
Asked directly, and stated as an opinion because it is not yet decided.
**Do not group the provider modules.** Three reasons, in order of weight:
1. **The evidence for grouping them is the wrong evidence.** Finding 3.
2. **A requirement names a provider module.** Grouping providers means a consumer names a
resource type and something else chooses the implementation. That is not a folder move; it
is implementation selection, a substantially larger design with its own failure modes, and
nothing currently asks for it.
3. **It would hide which implementation serves a requirement** — the one place the mesh most
needs to be explicit, and precisely the indirection ADR 0017 warns grouping causes.
A provider module is already exactly one purpose: it provisions one resource type. That is a
boundary, not an accident of installation.
## Still to do
- Draw the list for the cases where grouping is justified, and say plainly which entries rest
on measurement and which on judgement.
- Decide the ~50 silent modules: correctly sized, or unmeasured?
- Test the reachability grouping against the migration cost — the modules in it are among the
most-changed in the catalogue, so churn during a move is not hypothetical.
@@ -0,0 +1,94 @@
---
status: active
initiated: 2026-08-23
touches:
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
- 03-DESIGN/00-as-is/00-overview.md
- 03-DESIGN/01-to-be/00-work-breakdown.md
became: []
---
# 006 — The mesh designed from nothing
## What is being investigated
What the mesh would look like if it were laid out today, with the requirements known and none
of the accumulated shape — expressed as a **skeleton**: repositories at the root, modules
inside them, and whatever turns out to be the right leaf unit below that.
The deliverables are [`skeleton.md`](skeleton.md) — tiers, repositories and the four design
moves — and [`code-skeleton.md`](code-skeleton.md) — the tier test, what a module looks like on
disk, and where today's catalogue lands.
## Why
Every structural decision so far has been a **correction**: eight contexts replacing thirty-three
modules ([ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)), domains
replacing single-function modules
([ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)). A
correction inherits the frame of the thing it corrects, and two of the mesh's oldest problems
look unsolvable from inside that frame:
- **The bootstrap circularity.** The mesh needs a database, a bus, a registry and an identity
provider. Those are modules the mesh installs. The mesh cannot install them before it exists.
This has been worked around repeatedly and never designed away.
- **Participation requires privilege.** Everything assumes root on a machine whose packages and
services the mesh owns. A phone cannot participate on those terms, and neither can a machine
someone else administers.
Designing from nothing is a way to find out which parts of the current shape are requirements
and which are residue.
**And there is more residue than expected, from a knowable source.** This began as a
**dotfiles repository** — the first two days of history adopt dotfiles, add per-node dotfile
overrides, and introduce service symlinking with an ignore file. The flat one-directory-per-tool
catalogue, linking rather than copying, adoption of already-configured machines, per-node
overrides, and the desktop modules are all inherited from that, not chosen for a mesh. Recorded
in [`03-DESIGN/00-as-is/10-module-catalogue.md`](../../03-DESIGN/00-as-is/10-module-catalogue.md).
That makes this effort's question sharper than "what would we do differently": much of what
looks like design is a generalisation of *place files on my machines*, never revisited because
it was never stated as an assumption.
## The requirements this is designed against
Stated by the operator, recorded here so the skeleton can be checked against them rather than
against taste:
1. The mesh manages multiple computers — **full control**, through modules installed to nodes.
2. Mesh state lives in a **database**: which modules on which nodes, logs, configuration.
3. Configuration has **several touchpoints** — tool surface, web interface, others — all hosted
by the mesh itself.
4. **Connectivity** is core: every node reachable from every other over a shared overlay, some
nodes publicly exposed, firewalls configured.
5. The mesh **hosts applications** — and requires some of them itself. This is the circularity.
6. **Arch Linux only for now**; ideally any device, including phones, on lighter terms.
7. The end goal is to **operate an IT company** on it — development, design, deployment, full
circle, self-hosted. Personal cloud infrastructure.
8. **Agents make it self-improving and self-healing.**
9. It is **end-to-end testable on one machine**
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
## Status
A first skeleton exists, with four design moves that the current shape does not have. It is
`active` because two of them are unproven and one contradicts a record that is already
accepted.
**Finding worth stating up front:** [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)
names nine bounded contexts and **none of them owns connectivity** — no overlay, no resolution,
no firewall, no ingress. Requirement 4 has no home in the accepted decomposition, while
[research 005](../005-domain-grouping/analysis.md) found reachability to be the *only* part of
the catalogue where modules genuinely change together under one intent. The skeleton adds it.
## Open questions
| Question | Why it is open |
|---|---|
| Does the record — the event log contexts integrate through — belong to the substrate or the control plane? | It is infrastructure by shape and domain by content. Placing it wrong reintroduces a circularity. |
| One repository per tier, or per context? | Already open from ADR 0015 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. |
| Does an unprivileged node earn a place in the inventory, or only a presence? | Decides whether "node" means one thing or two. |
| Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large? | It is the skeleton's biggest unproven claim. A binary whose whole argument is that it has no dependencies now carries six concerns. |
| Four substrate services or five? | The identity provider passes the tier test only if the control plane delegates authentication rather than doing it natively. |
| Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. |
@@ -0,0 +1,272 @@
---
effort: 006-mesh-from-scratch
updated: 2026-08-23
---
# The code skeleton
[`skeleton.md`](skeleton.md) laid out tiers and repositories. This is the level below: what a
module looks like on disk, and — the question that forced a correction — where a given piece of
today's catalogue actually lands.
## The tier test
A decision procedure, so placement is answerable rather than argued. Ask in order; first match
wins:
1. **Does it apply state on a machine?** → tier 0, inside the host.
2. **Can the control plane *start* without it?** If *no* → tier 1, substrate. Note the verb:
*start*, not *function fully*. A capability the control plane loses without something is not
the same as a thing it cannot come up without — see "becoming self-hosting" below.
3. **Does it decide what should be true across nodes?** → tier 2, a control-plane context.
4. **Is it a way to talk to tier 2, holding no logic of its own?** → tier 3, a surface.
5. **Otherwise** → tier 4, a workload.
## Worked example — where postgres ends up
The obvious answer is "twice": once as the mesh's own database in the substrate, once as a
hosted database in the catalogue. That answer is wrong, and seeing why fixes something.
Run the test. *Can the control plane exist without a relational store?* No. **Postgres is
tier 1.** It lives once:
```
mesh-substrate/store/postgres/
```
There is no second copy in the catalogue, because the thing that differs between the mesh's own
database and a project's database is **not the module**. It is how that instance is brought up:
| | The mesh's own instance | A project's database |
|---|---|---|
| Brought up by | the host, from the pinned bundle, with no control plane present | the ordinary delivery and provisioning path |
| Declared in | `mesh-substrate/bundle.yml` | the consuming module's `requires:` |
| Exists because | the control plane cannot start without it | something asked for it |
Same module, two roles. **Tier is a property of the module** — what must exist before what — and
**the bundle is a property of the mesh's own instance.**
This is also legal under the dependency rule, which is worth checking rather than assuming: a
tier-4 workload that requires a database depends on tier 1, which points *downward*. The
inverse — substrate reaching into the catalogue for a module — would not be, and is the shape
the naive "twice" answer would have created.
The same reasoning places the rest of the substrate: the bus, the object store, the image
registry. Each fails step 2, each lives once, each is pinned. That is the whole substrate: four
services, and the deliberate absence of a fifth.
**The identity provider was raised as a boundary case and is settled: it is not substrate.**
The mesh does not require one — tier 2 authenticates its own callers natively, and an identity
provider is a service the mesh hosts like any other. The substrate is four services, not five.
The test earned its keep here by turning a vague unease into one answerable question — *does the
control plane delegate authentication?* — rather than a debate about how important identity
feels.
## Becoming self-hosting — the forge and the registries
Self-improvement means the mesh hosts the things it improves itself with: a forge, an image
registry, a package registry. The obvious worry is that these duplicate — a `mesh-gitea`
for the mesh and a `gitea` for everyone else. **They do not, and the reason is worth stating
carefully, because it is the same reason the bootstrap keeps failing today.**
### They are not substrate
Run the test with the sharpened verb. *Can the control plane start without a forge?* **Yes.** It
comes up, holds inventory, answers questions and manages nodes with the modules it already has.
What it cannot do is **change itself**. That is a capability, not a precondition.
So: forge, image registry and package registry are **tier 4 workloads**. One module each, in the
catalogue, exactly like the media server. There is no mesh-specific copy.
This is not a technicality. It buys a property worth having: **if the forge dies, the mesh keeps
running.** Nodes stay managed, services stay up, only self-modification stops. Putting the forge
in the substrate would make losing it fatal, for no gain.
### But delivery needs them — is that not an upward dependency?
It would be, stated naively, and that would break the one rule the whole skeleton rests on.
It is resolved the way the constitution already says to resolve it — **depend on abstractions,
not on concrete dependencies**. Tier 2's delivery context does not require *the forge module*.
It declares requirements:
| Delivery requires | Satisfied by |
|---|---|
| a source of record for module code | whichever module provides it |
| somewhere to publish images | whichever module provides it |
| somewhere to publish packages | whichever module provides it |
| somewhere to put build artifacts | the substrate's object store |
Tier 2 defines the requirement; tier 4 provides the implementation; the binding is data. The
dependency points **downward from the provider to the interface**, which is legal, and the
control plane never names a concrete module.
The mechanism for this already exists and is the mesh's most valuable one: **provisioning**. A
module declares what it provides; a consumer declares what it requires; the mesh binds them.
The only new idea is that **the control plane is itself a consumer** — it has requirements, and
they are satisfied the same way a workload's are.
That generalisation is significant enough to need its own study, and is not settled here.
### Self-hosting is a state the mesh reaches, not a precondition
This is the part today's mesh gets wrong, and it explains a recurring class of pain.
A first node comes up from **pinned external artifacts** — upstream images, by digest, carried
in the bundle. It has to: the mesh's own registry does not exist yet, and cannot. The mesh at
this point is running and manages nodes, and is not yet self-hosting.
Self-hosting is then **reached**: the forge is installed as an ordinary workload, the mesh's own
source moves into it, the registries come up, and delivery's requirements are re-bound from
external providers to internal ones. From that point the mesh builds and deploys itself.
Stated as a lifecycle:
```
pinned external artifacts ─► mesh runs, manages nodes
│
│ forge + registries installed as workloads
│ delivery's requirements re-bound
▼
mesh builds and deploys itself
```
**Today's mesh assumes the second state from the first moment.** Its source, its packages and
its images are all expected to be self-hosted before there is anything to host them — which is
why raising a first node needs a script that exists solely to paper over the impossibility, and
why that script is the least-exercised path in the system.
Making the transition explicit has a second benefit: it is reversible. A mesh whose forge is
broken can re-bind delivery to external providers and keep improving itself while it repairs
the forge. Today that escape hatch does not exist, because the dependency is not expressed
anywhere it could be changed.
### So, concretely
One `gitea` module. One image-registry module. One package-registry module. Each a tier-4
workload. The mesh's own instances are distinguished from any other instance **by what they are
bound to, not by being different modules** — precisely the same answer as postgres, arrived at
by the same test.
## The five fates of a module
Research 005 asked whether a catalogue module should be **grouped into a domain** or **leave
the repository**. Working the tier test across the catalogue surfaces a third answer that
neither option covers, and it is the most common one.
| Fate | Means | Examples from today |
|---|---|---|
| **Absorbed into the host** | It is not a module at all. It is part of what "managing a machine" means, and belongs in tier 0. | overlay membership, packet filtering, package management, service supervision, container runtime, filesystem management |
| **Substrate** | The control plane cannot exist without it. Pinned, host-applied. | relational store, bus, object store, image registry |
| **Control-plane context** | It decides something across nodes. | connectivity policy, inventory, delivery, provisioning, observability |
| **Workload module** | The mesh hosts it. Grouped per [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md). | media library, desktop session, collaboration tooling |
| **Leaves the repository** | A standalone application, per [ADR 0010](../../02-DECISIONS/0010-applications-live-in-their-own-repository.md). | the applications identified in research 005 |
**The first fate is the finding.** Research 005 measured the reachability cluster — proxy,
resolver, firewall, overlay — as the only place in the catalogue where modules genuinely change
together under one intent. The skeleton explains *why*: they are not four modules that ought to
be one domain module. They are four facets of one thing the host should own, currently
expressed as modules because a module was the only unit available.
Under this skeleton the overlay module and the firewall module **stop existing**. The host holds
membership and applies filtering; tier 2 decides the policy; the swappable backends stay
modules. That is a different and better answer than grouping them, and it was not visible from
inside the current frame.
It also partly answers research 005's other open question — the fifty modules that co-change
with nothing. Several are host concerns rather than domains: package management, container
runtime, filesystem tooling. Silence was the right signal after all; the wrong conclusion was
that grouping was the only available fix.
## What a module looks like on disk
Per [`skeleton.md`](skeleton.md) Move 4, a module declares **parts** — independently selectable
pieces of desired state — and produces **artifacts** — things built once per version.
```
<module>/
module.yml identity, what it provides, what it requires,
which host profiles it can land on
parts/
service/ desired state: container, volumes, exposure
provisioner/ how it grants its resource to consumers
migrations/ its own persistent state
tools/ capabilities it contributes
artifacts/
<name>/ source for something built and published
```
A module with no artifacts consumes an upstream image and builds nothing. A module with no
parts is not a module.
Worked through for the substrate's relational store:
```
mesh-substrate/store/postgres/
module.yml provides: database · profiles: [managed]
parts/
service/ the container, its volume, its network exposure
provisioner/ grants a database and role to a consumer
migrations/ none — it holds no state of its own
artifacts/ none — upstream image, pinned by digest in bundle.yml
```
## The tree, at file level
```
mesh-host/ TIER 0
cmd/host/
internal/
apply/ reconcile declared state
inventory/ what this machine is and can do
link/ outbound connection to the control plane
overlay/ membership: address, keys, tunnel
filter/ packet filtering from tier-2 policy
packages/ package management
services/ supervision
containers/ container runtime
store/ embedded local state
profile/ managed · user · edge
substrate.lock pinned tier-1 descriptor
mesh-substrate/ TIER 1
bundle.yml the pinned set, by digest
store/postgres/
bus/<broker>/
objects/<object-store>/
images/<registry>/
mesh-control/ TIER 2
record/ the event log contexts integrate through
inventory/ nodes · modules · assignments · versions
config/ settings · secrets · derivation
connectivity/ addresses · resolution · exposure · filtering · certificates
provisioning/ grants between modules
delivery/ source → artifact → node
observability/ health · logs · metrics
identity/ agents · humans · services · authorisation
work/ tasks · workflows · runs
knowledge/ memory · documents · retrieval
api/ the one interface surfaces speak to
mesh-surfaces/ TIER 3
tools/ web/ cli/
mesh-catalog/ TIER 4
<domain>/<module>/ layout as above
mesh-lab/ mesh-sdk/ mesh-hq/
```
## What this does not settle
- **The identity boundary case.** Four substrate services or five.
- **Where the record lives.** Still the open question from
[`00-overview.md`](00-overview.md), and the tier test does not resolve it: the record is
needed by tier 2 and is *of* tier 2, which is exactly the shape that produces a circularity.
- **Whether absorbing into the host makes the host too large.** Six internal concerns is already
a lot for a binary whose whole argument is that it has no dependencies. The counter-argument
is that each is small and none can be optional — but this is the skeleton's biggest unproven
claim, and it should be tested by writing the host's interface before anything else.
- **The migration.** Nothing here says how today becomes this.
@@ -0,0 +1,280 @@
---
effort: 006-mesh-from-scratch
updated: 2026-08-23
---
# The skeleton
Repositories at the root, modules inside them, parts at the leaf. Four tiers, and a dependency
rule that only points downward.
## What a tier is
A tier answers one question: **what has to exist before this can exist?** It is not importance,
and it is not a layer in the networking sense. It is bootstrap order, made explicit.
Walk a bare machine to a running mesh and the tiers fall out of the story:
```
bare machine
│ one command lands ONE binary. Nothing else exists. TIER 0 host
│
│ it reads a pinned file it already carries and raises a
│ database, a bus, an object store and an image
│ registry — locally, alone. TIER 1 substrate
│
│ on those, the mesh's brain starts: which nodes exist,
│ what runs where, what is reachable. TIER 2 control plane
│
│ ways to talk to that brain. TIER 3 surfaces
│
└ everything the mesh then carries. TIER 4 workloads
```
**Dependencies point only downward.** The substrate never references the control plane. That
one constraint is the entire bootstrap answer, because it guarantees there is always a place to
start.
Why it earns its keep here: **the current mesh violates this, and that is the circularity that
keeps recurring.** The mesh database is a module; modules are installed by the delivery
pipeline; the pipeline needs the database. No order works, so a first-node script exists to
paper over it, and every later substrate change has to pretend the problem is not there.
A tier is not a repository and not a bounded context. Those are different cuts: a context says
*who owns this concept*, a tier says *what must already be running*.
## The tree
```
mesh-host/ TIER 0 — the only thing ever installed by hand
apply/ reconcile declared state on this machine
inventory/ what this node is, has, and is capable of
link/ the single outbound connection to the control plane
store/ embedded local state — authoritative while disconnected
profile/ capability detection: managed · user · edge
substrate.lock pinned tier-1 descriptor, appliable with no mesh present
mesh-substrate/ TIER 1 — declarations only, no logic of its own
store/ relational state
bus/ commands and events
objects/ blobs and build artifacts
images/ container images
bundle.yml the pinned set tier 0 can raise alone
mesh-control/ TIER 2 — the control plane
record/ the event log every context integrates through
inventory/ nodes · modules · assignments · versions
config/ settings · secrets · derivation onto nodes
connectivity/ overlay · resolution · exposure · filtering · certificates
provisioning/ resource grants between modules
delivery/ source → artifact → node
observability/ health · logs · metrics · alerts
identity/ agents · humans · services · authorisation
work/ tasks · workflows · runs
knowledge/ memory · documents · retrieval
api/ the one interface every surface speaks to
mesh-surfaces/ TIER 3 — thin; no logic lives here
tools/ the agent-facing tool surface
web/ the operator-facing interface
cli/ the shell-facing interface
mesh-catalog/ TIER 4 — what the mesh hosts
<domain>/ grouped per ADR 0017, list per research 005
mesh-lab/ the whole mesh, disposable, on one machine
mesh-sdk/ contracts shared across tiers — types, not behaviour
mesh-hq/ this repository
```
## The dependency rule
**A tier may depend only on tiers below it.** Substrate never references the control plane.
The control plane never reaches into a node except through the host. A surface holds no logic
a second surface would have to reimplement.
This is the whole of the bootstrap answer, and per this repository's own rule it must say how
it is checked: a dependency-direction lint in the build, failing on an upward import. A tier
rule enforced by intention is the same as no tier rule — that is
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to architecture.
## Move 1 — the substrate is applied, not delivered
**The problem.** The mesh needs a database, a bus, an object store and an image registry.
Today those are modules, and modules are installed by the delivery pipeline, which
needs the database and the bus. The first node is therefore raised by a special script that
exists only because of the circularity, and every later change to the substrate has to pretend
the circularity is not there.
**The move.** The host can apply a declaration without anyone telling it to. The substrate is
a **pinned bundle** the host carries: a fixed, versioned, self-contained descriptor of the
four services and nothing else. Raising a first node is `host apply substrate.lock` — not a
special path, just the ordinary one with no control plane on the other end.
The circularity disappears rather than being worked around: **the substrate is applied by tier
0, the mesh is delivered by tier 2, and they are different mechanisms on purpose.**
The price is real and should be named: the substrate is upgraded by bumping a pin and
re-applying, not by the pipeline. It gets less machinery than everything else — no per-node
selection, no provisioning, no fan-out — and that is the point. Five services justify a
simpler mechanism than a hundred.
## Move 2 — the host is one binary with capability profiles
**The problem.** Everything assumes root on a machine whose packages, services and network the
mesh owns. A phone cannot offer that, and neither can a work laptop. The current answer would
be a lightweight fork, which means two implementations and one of them rotting.
**The move.** One host binary, and a **profile** it detects rather than is told:
| Profile | Can | Typical |
|---|---|---|
| `managed` | packages, services, network, filesystem — the full surface | a machine the mesh owns |
| `user` | user-level services and tools; no package or network management | a shared or administered machine |
| `edge` | report presence, relay, expose a tool surface; hold nothing | a phone |
A module declares which profiles it can land on. Assignment to an incapable node fails at
declaration time, not at deploy time — a phone is not a machine that fails to install a
firewall; it is a node the firewall module cannot be assigned to.
This makes the phone case a **capability question rather than a platform question**, which is
what keeps it from becoming a second implementation. It also removes the current unstated
assumption that every node is equivalent — already false, and today handled by remembering.
## Move 3 — connectivity becomes a context
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) names nine contexts
and none of them owns the overlay, the resolver, the firewall or the ingress. `config` owns
PKI, which is the closest thing, and it is not close.
Meanwhile [research 005](../005-domain-grouping/analysis.md) measured the whole catalogue and
found that reachability is the **only** place where modules genuinely change together under one
intent — the proxy with the resolver, the firewall with the overlay, repeatedly, because *how a
node is reachable* is one question asked in four places.
So the evidence and the gap point the same way. `connectivity` owns:
- the overlay every node joins, and the addresses on it
- name resolution, internal and public
- exposure — which services answer from outside, on which names
- filtering — what may reach a node at all
- certificates for both name spaces
This is an addition to an accepted record, so it is a decision, not a drafting choice. It
belongs in a new record that extends ADR 0015 the way
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) does —
not written here.
### Where the networking actually lives
A context is an authority, not a running thing, so naming one does not say who brings the
overlay up. That splits three ways, and the split is the design:
| Concern | Tier | Why there |
|---|---|---|
| **Overlay membership** — this node joins, holds an address, keeps the tunnel up | **0, in the host** | Everything cross-node needs it *before* the substrate is reachable from elsewhere. A module cannot provide it, because installing a module is itself a cross-node operation. |
| **Policy** — who holds which address, what resolves, what is exposed, what is filtered | **2, `connectivity`** | Bookkeeping and authority. It decides; it runs nothing. |
| **Machinery** — resolver, reverse proxy, firewall backend, certificate issuance | **4, modules** | Swappable, and not every node needs them. A node without a reverse proxy is still a node. |
There is a second circularity hiding here, and it has to be closed explicitly: **the host's link
to the control plane does not run over the overlay.** If it did, the overlay would have to be up
before the host could be told how to join it. The link is ordinary outbound internet to a public
endpoint; the overlay carries node-to-node traffic only.
Joining is therefore: host lands → links out with a join token → control plane returns an
address and keys → host raises membership → the substrate on other nodes becomes reachable.
This also makes the `edge` profile honest rather than special-cased. A phone can hold overlay
membership in userspace without privilege, and cannot run the machinery. That is the profile
distinction doing its job.
## Move 4 — `feature` splits in two
The invitation was to check whether the concept survives. It does not, in one piece.
Today a **feature** means both *a thing built once* and *a thing selected per node*, and the
delivery pipeline is hard to reason about precisely because those have different cardinality
and one word ([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)
is the pipeline half of the same confusion).
Split it:
| Concept | Is | Cardinality |
|---|---|---|
| **artifact** | something built and published — an image, a bundle, a package | once per module version |
| **part** | an independently selectable piece of a module's desired state | chosen per node, per assignment |
A module declares desired state in parts, and produces artifacts. An assignment names the node
and the parts. Delivery builds artifacts once and applies parts per node — and the two words
now carry the two cardinalities that the pipeline already has.
This keeps what features are genuinely for — per-node opt-in of *some* of a module, which the
work breakdown already calls for — and drops the conflation that makes the current model
confusing.
## What the mesh is, versus what it hosts
Tiers 0–3 are the mesh. Tier 4 is everything it carries, and the boundary is stated by
requirement rather than by taste: **a module is part of the mesh if removing it stops the mesh
managing nodes.** A media server does not. A relational store does — which is why it sits in the
substrate and not the catalogue, despite being, in every other respect, an application like any
other. An identity provider, notably, does **not**: the control plane authenticates its own
callers, so identity is a hosted service like the media server.
That test also settles the IT-company goal without a special category. Development, design and
deployment tooling are **workloads** — tier 4, hosted, provisioned, delivered like anything
else. The mesh does not grow a "company" feature; it hosts the tools a company runs on, and
the fact that it runs its own development on them is dogfooding, not architecture.
## What agents are, structurally
Self-improvement and self-healing are not a tier. Agents are participants
([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)) that hold identity in
tier 2, act through tier 3 like any other caller, and run as workloads in tier 4.
This matters for one reason: **an agent must not have a privileged path**. Anything an agent
can do to the mesh, a person can do through the same surface, and anything it cannot express
through the tool surface is a gap in the surface rather than a reason for a back door. Self-
healing built on a private channel is unreviewable, and would be the one part of the mesh with
no human checkpoint.
## How this is tested
The lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) raises the
tree above on one machine: virtual machines as nodes, a real overlay between them, the real
substrate bundle, the real control plane, the real delivery path.
The tier rule is what makes that affordable. A scenario needing only tiers 0 and 1 is one
virtual machine and a pinned bundle — which is also, exactly, the bootstrap path. **The
hardest thing to test becomes the cheapest scenario to run**, and the first-node path stops
being the one thing nobody exercises until it breaks.
## What this skeleton does not answer
- Where the record lives. It is infrastructure by shape and domain by content, and putting it
in the substrate risks recreating a circularity in the one place the design just removed one.
- Whether tier 2's contexts are one repository or several. Open from ADR 0015 already.
- Whether an `edge` node is in the inventory or merely present — which decides whether "node"
is one concept or two.
- The migration. Nothing here says how today's mesh becomes this, and the skeleton is worth
little until that is costed.
## A naming near-miss, recorded
The tier-0 binary was first called `mesh-agent`, because "node agent" is the reflex everywhere
else in the industry. That is wrong here, and wrong in the specific way
[`how-we-build.md`](../../00-META/how-we-build.md) §4 exists to catch: **Agent** is a
first-class concept in this mesh — a participant, some of whom are human, holding identity and
memory ([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)). One document
carried both meanings.
It is the same failure as the anatomy naming in the current runtime: an evocative domain word
pointing at infrastructure.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) supplies the fix in
its own title — *the mesh brokers capabilities; nodes host; agents think.* Three verbs, three
components: the control plane **brokers** (`mesh-control`), the tier-0 binary **hosts**
(`mesh-host`), the participant **thinks** (`agents`, untouched).
`mesh-node` was the alternative and was rejected: *Node* is the aggregate in the inventory — the
record of a machine — while the binary is what runs on it and does the hosting.
@@ -0,0 +1,51 @@
---
status: active
initiated: 2026-08-23
touches:
- 03-DESIGN/00-as-is/03-provisioning.md
- 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
became: []
---
# 007 — Provisioning as the mesh's core mechanism
## What is being investigated
Provisioning is the mechanism the whole mesh rests on: a module declares what it needs, and the
mesh makes it exist, generates the credential, records the grant, and puts the values where the
module will read them. [ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)
calls it the mesh's core concern rather than its plumbing.
[Research 006](../006-mesh-from-scratch/code-skeleton.md) then asks it to carry **more**: the
control plane becomes a consumer with its own requirements — a source of record, an image
registry, a package registry — satisfied by the same mechanism. That generalisation is only
safe if the mechanism is sound, and the as-is record says it is not, in named ways.
## Why now
Four weaknesses are already documented in
[`03-DESIGN/00-as-is/03-provisioning.md`](../../03-DESIGN/00-as-is/03-provisioning.md), each
observed rather than theorised:
- **Rotation has no fan-out.** A shared credential can be rotated without telling the peers
holding the old one. This has locked the mesh out of its own broker.
- **A grant is not a check.** The record says a resource was provisioned. Nothing verifies it
still exists, still has that credential, or is reachable from where the consumer runs.
- **A frozen password outlives its generation.** A generated secret written once diverges from a
persistent data directory initialised earlier, and presents as an authentication error.
- **No requirements is indistinguishable from provisioning that did not run.** A module that
declares nothing skips the stage, which is correct, and looks identical to failure.
Generalising a mechanism with these properties to the control plane's own dependencies would
make each of them fatal rather than annoying.
## The questions
| Question | Why it matters |
|---|---|
| What does a **grant** mean, exactly — a record that a resource was created, or a claim about the world that is continuously reconciled? | The difference between the current model and one where "provisioned" is checkable. Almost every weakness above is a symptom of the first answer. |
| How is a credential **rotated** with fan-out to every holder? | The mechanism grants easily and regrants not at all. This is the most damaging gap and it has taken the mesh down. |
| Can the **control plane** hold requirements, and what satisfies them before anything is installed? | The generalisation research 006 needs. Ties directly to the self-hosting transition. |
| What happens when a requirement **cannot** be satisfied — no provider, provider on an unreachable node, provider not yet installed? | Today this is silent or a stall. It should be a stated, visible state. |
| Does a requirement belong to a **module** or to one of its **parts**? | Research 006 splits `feature` into artifact and part. A part-scoped requirement means a database is not provisioned where the part that needs it is not installed. |
@@ -0,0 +1,53 @@
---
status: active
initiated: 2026-08-23
touches:
- 03-DESIGN/00-as-is/04-delivery.md
- 02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md
- 02-DECISIONS/0013-an-artifact-is-build-output.md
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
became: []
---
# 008 — The coordinator: a change checked in becomes a deployed state
## What is being investigated
The mesh's own continuous delivery: a change is committed, and the mesh ends up in the state
that change describes — across every node the change touches, with a verdict that says whether
it worked.
The coordinator is what orchestrates that, and it is the mesh's most consequential machinery:
everything reaches every node through it.
## Why now
The as-is record ([`03-DESIGN/00-as-is/04-delivery.md`](../../03-DESIGN/00-as-is/04-delivery.md))
names problems that are structural rather than incidental:
- **A green pipeline proves transport, not effect.** The stages report that a message was
dispatched and accepted, which is not the same as the thing running, correct, or present.
This is the mesh's single most consistent failure shape.
- **Detection is the most fragile input.** A merge that creates no pipeline, with nothing saying
so, is the characteristic bad outcome — and it has happened for reasons unrelated to the
change.
- **The fan-out point is asymmetric.** The build node has already passed two silos when work
fans out, and code that knew only about the first parked it forever while every other node
deployed cleanly.
- **There is no end-to-end coverage.** The harness has not built since 2026-06-04
([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)).
Research 006 adds a requirement the current design does not have: the coordinator must work
**before the mesh is self-hosting**, when source and artifacts come from outside, and keep
working across the transition to self-hosted providers.
## The questions
| Question | Why it matters |
|---|---|
| What is a **deployed state**, and how does the mesh know it is in one? | Everything follows from this. If a stage reports transport, "deployed" is a claim nobody checked. A desired-state model with reconciliation gives a different answer from a job-completion model. |
| Does the coordinator dispatch **stages**, or converge nodes on a **declaration**? | The current model is a state machine over stages. The alternative is that a node is told what should be true and reports what is. The second makes drift visible; the first cannot see it. |
| How does a change **become** a pipeline, reliably? | Detection has failed for reasons unrelated to the change, silently. |
| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) and to a module carrying its own assertions. |
| How does delivery work **before self-hosting**, and across the transition? | From research 006: source and artifacts start external and are re-bound to internal providers. The coordinator has to be indifferent to which. |
| Does the **three-silo** split survive the artifact/part split? | [ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md) is cardinality-driven, and research 006 renames the thing the cardinality is about. |
+92
View File
@@ -0,0 +1,92 @@
---
status: active
initiated: 2026-08-23
touches:
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
- 03-DESIGN/01-to-be/01-end-to-end-testing.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 03-DESIGN/00-as-is/00-overview.md
became: []
---
# 009 — Getting from the mesh that exists to the mesh that is designed
## What is being investigated
How a running mesh becomes the one in
[research 006](../006-mesh-from-scratch/code-skeleton.md), without losing what it currently
carries.
The proposed shape: **build tiers 0, 1 and 2, then replace the current setup in one move.**
## Why big-bang is the right instinct here
Recorded because incremental is the reflex answer and it is wrong in this case.
- **The two models are structurally incompatible.** The tier rule, the host absorbing what are
now modules, the artifact/part split, provisioning generalised to the control plane's own
requirements — none of these can half-apply. Running both models at once means the old one's
assumptions keep constraining the new one, which is how a migration becomes permanent.
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
same risk as one performed for the first time on the real mesh.
- **Incremental would carry the rot forward.** The as-is layer documents silent failure paths,
a dead test harness and unenforced rules. A gradual migration preserves them by definition.
## The distinction that lowers the risk
**Replace the control plane; do not move the workloads.**
The things that would hurt to lose — mail, media, source, databases, their data directories —
are not the mesh. They are what the mesh manages. They sit in container volumes on nodes, and
they do not need to move for the control plane above them to be replaced.
So the big-bang is: the old control plane stops managing these nodes, and the new one starts —
with the workload data untouched, in place, and re-declared rather than migrated.
That reframing turns "replace the mesh" into "replace the part with no persistent state of its
own", which is a materially smaller act than it first sounds.
## The tension this exposes
Taking over already-running workloads is **adoption**, and adoption was ruled out of scope —
recorded as a legacy path in
[`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../../03-DESIGN/01-to-be/01-end-to-end-testing.md).
The migration appears to need exactly the capability the design declared it would not have.
The way out, to be tested: the new mesh does not adopt anything. It **declares** the workloads
from scratch and points them at data directories that already exist. Nothing inspects a running
machine to learn what is there; the declarations are written from the as-is layer, which is what
that layer is for. Data survives because it was never touched, not because it was adopted.
If that holds, adoption stays out of scope and the migration is ordinary declaration. If it does
not, adoption needs a one-time, explicitly unsupported tool, and that should be a decision rather
than a discovery.
## Sequencing
| Phase | What | Done when |
|---|---|---|
| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
| E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
| F | Cut over. | The real nodes are managed by the new control plane. |
| G | Reach self-hosting — re-bind delivery from external providers to the mesh's own forge and registries. | The mesh builds and deploys itself. |
Phase G is deliberately last. Per research 006, self-hosting is a state the mesh **reaches**;
attempting the cutover and the self-hosting transition in the same move recreates exactly the
circularity the skeleton removes — and would mean a failed cutover could take away the means to
fix it.
## The questions
| Question | Why it matters |
|---|---|
| What state must **survive** the cutover, versus be re-created? | Workload data must. Provisioned credentials could be re-issued. The mesh's own inventory could be re-declared. Each answer changes the risk. |
| What is the **way back**? | A cutover with no rollback is not a plan. If workload data is untouched, reverting may be as small as re-pointing the old control plane at it — to be verified, not assumed. |
| How is the cutover **rehearsed** against something resembling the real mesh, without copying the real mesh into a repository? | The lab must be able to model the real topology's shape without carrying its identity. |
| Does anything have to keep running **during** the cutover? | Mail and source are the obvious candidates. If yes, "big-bang" is really "big-bang with exceptions", and the exceptions should be named now. |
| Is the forge inside or outside the cutover? | If the mesh's own forge goes down with the old control plane, the means of deploying a fix goes with it. This is the self-hosting circularity appearing as a migration risk. |
+37 -7
View File
@@ -4,13 +4,43 @@ Investigations that have not yet hardened into design.
## Structure
Each effort lives in `NNN-descriptive-name/` and **must** contain `status.md` with:
Each effort lives in `NNN-descriptive-name/` and **must** contain `00-overview.md`, carrying its
state in YAML frontmatter and a prose summary below it:
- a short summary of the effort
- who initiated it
- the areas it touches
- current status: `ONGOING`, `GRADUATED`, or `ABANDONED`
```yaml
---
status: active | graduated | abandoned
initiated: YYYY-MM-DD
touches: [] # design docs, subsystems or areas the effort bears on
became: [] # required when status is terminal — what it turned into
---
```
An effort graduates by producing an ADR and a `02-DESIGN` entry. It is abandoned in
place — never deleted. What was rejected, and why, is the more expensive half to
The prose says what is being investigated, why, and what it touches. It does not restate the
status — status lives in one place, and two places is one too many.
Further documents in the same folder hold the work itself: notes, evidence, option analyses,
draft designs.
## Lifecycle
| status | Meaning |
|---|---|
| `active` | Investigation in progress. |
| `graduated` | Checked against `00-META`, decided in `02-DECISIONS/`, and specified in `03-DESIGN` — see `became:`. |
| `abandoned` | Stopped or superseded. Nothing is deleted. |
An effort graduates by producing a decision record **and** a `03-DESIGN` entry. It is abandoned
in place — never deleted. What was rejected, and why, is the more expensive half to
rediscover.
Starting and closing efforts is playbook territory:
[`00-META/process/01-research.md`](../00-META/process/01-research.md) and
[`02-graduation.md`](../00-META/process/02-graduation.md).
## Rules
- Markdown only. Do not skip or reuse a sequence number.
- **Evidence, not assertion.** An effort that measured nothing has not finished.
- Research describes real observations but **never identifies the mesh it observed**. The shape
of a finding survives anonymisation intact.
@@ -0,0 +1,68 @@
---
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 1. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
> deliberation that happened.
## Context
The mesh is a set of machines that must call each other's capabilities. On the day the
repository was founded there was no inter-node transport at all — each node was configured
independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
## Considered options
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
participant. It also makes node liveness a caller's problem: a request to a sleeping node
is an error rather than a wait.
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
independent of demand, and request/reply has to be built on top of it by hand.
3. **A central message broker with per-node exchanges.** Chosen.
## Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange
named for itself and a request queue; a shared mesh exchange carries commands and events that
are not addressed to one node.
Three message shapes, and only three:
- **RPC** — request/reply, for calling a capability that lives on another node.
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
- **Events** — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
## Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker
connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually
right and occasionally the wrong thing entirely — a queued command for a node that never
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
where the capability is.
## References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.
@@ -0,0 +1,73 @@
---
status: accepted
date: 2026-03-14
deciders: jochen
reconstructed: true
---
# 2. Everything is a module, and one manifest describes all of them
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh carries several kinds of thing: containerised services with data and ports, pure
capability providers with no service at all, and bare markers whose only content is that a
node has them. Before this decision these were separate concepts with separate handling —
the earlier vocabulary was *capabilities*, and services were installed by a different path
than tools.
Every distinct kind of thing needs its own install path, its own change detection, its own
place in the delivery pipeline, and its own documentation. Three kinds means three of each,
and every new feature has to be built three times or, more commonly, once — leaving two kinds
quietly unsupported.
## Considered options
1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag
list. Rejected: it is what existed, and the cost was paid in every cross-cutting change.
2. **One manifest, kind inferred from directory contents.** Chosen.
3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service
still declares itself — but the general rule became inference, because a declared list and
the directory it describes drift, and the directory is the one that is true.
## Decision
Everything the mesh installs is a **module**: a directory with a manifest. The manifest
declares identity, environment variables, what the module provides, what it requires, and how
it is exposed. What kind of module it is follows from what the directory contains:
| Contains | Is |
|---|---|
| a compose definition | a service |
| a tools directory | a capability provider |
| a daemon or unit directory | a long-running process |
| a configs directory | a source of managed files |
| nothing but a manifest | a flag — presence is the whole content |
A module may be several of these at once. Each is a **feature**, and the delivery pipeline
addresses features, not modules.
The mesh's own components are modules on exactly these terms. They get no privileged install
path, no separate registry, and no exemption from the pipeline.
## Consequences
- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches
everything the mesh carries.
- Dogfooding stops being a discipline and becomes structural: if the mesh's own components
need an exception, the machinery is unfinished, and that is visible immediately.
- Feature detection from directory contents means a directory rename silently changes what a
module *is*. This has bitten repeatedly — a hook named for a feature the module does not
have is skipped without complaint.
- The manifest becomes load-bearing and grows. It is now the largest single point of
coupling in the mesh.
## References
- `Rename capabilities → modules across the entire codebase`, 2026-03-14.
- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive
by conversion rather than by creation.
- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`.
- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`,
`troubleshooting/health-check-tools-index-false-positive`.
@@ -0,0 +1,65 @@
---
status: accepted
date: 2026-04-02
deciders: jochen
reconstructed: true
---
# 3. The mesh database is the source of truth; the repository is node-agnostic
> Reconstructed after the fact from the evidence cited below.
## Context
Two things must be known to run the mesh: **what exists** — which modules there are, what each
declares, how each is built — and **what runs where** — which node hosts which module, with
which settings, at which version.
The repository is the natural home of the first. It was initially also the home of the second:
per-node directories held that node's configuration, and adopting a machine meant committing
its files. That has three costs. A node cannot be changed without a commit, so runtime state
and source share a review cadence they do not share a rhythm with. Two nodes cannot be
reconciled, because nothing holds both. And the repository becomes an inventory of the
installation, which is exactly the content that cannot be made public.
## Considered options
1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding
change is a commit and a deploy, and the repository accumulates an inventory of one
particular mesh.
2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no
authority: two nodes disagreeing have no arbiter, and drift is invisible until something
breaks.
3. **A mesh database as the single authority, cached locally for resilience.** Chosen.
## Decision
A single database holds every binding: which node hosts which module, at which selection, with
which environment overrides, plus mesh-level settings that all nodes read. The runtime loads
its configuration from that database at startup and falls back to a local cache when the
database is unreachable.
**The repository defines what exists. The database defines what runs where.** No node-to-module
mapping is ever committed.
A node is therefore not described anywhere in source. Bringing one into the mesh is a database
operation.
## Consequences
- The repository becomes node-agnostic, and can be published without disclosing an
installation. This repository's public stance rests on that property.
- A binding changes without a commit, a build, or a deploy.
- The local cache means a node survives losing the database, but a node running from cache is
running from a snapshot with no indication of its age. Divergence is silent by construction.
- The database is the hardest dependency in the mesh. It is also a module, provisioned like
any other, which makes its bootstrap circular — resolved by the first-node initialisation
script, and the reason such a script exists.
- Nothing on a node is authoritative. That is what makes the next decision necessary.
## References
- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration
tables that landed with it.
- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments."
- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`.
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-04-03
deciders: jochen
reconstructed: true
---
# 4. Managed files are generated onto nodes and never edited there
> Reconstructed after the fact from the evidence cited below.
## Context
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh
database. But the things that consume those bindings — environment files, service
definitions, daemon configuration, firewall rules — are files on a node's disk, because that
is what the software reading them requires.
So the same value exists twice: authoritatively in the database, and materialised in a file.
Any edit to the file is a change to a copy. Before this decision, environment values could be
pushed from a node back into the database, which made the direction ambiguous in both
directions at once.
## Considered options
1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed.
Two writers and no arbiter: whichever synced last wins, and neither is authority.
2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts
ADR 0003 and returns to state that cannot be reconciled across nodes.
3. **Strictly one-directional: the database is written, files are generated.** Chosen.
## Decision
Every managed file is **derived**. A synchroniser regenerates it from the mesh database
whenever the underlying values change. The write path is the mesh tool that owns the value;
the file is an output.
This applies to generated environment files, service definitions, managed configuration, and
anything else a synchroniser lists as its own.
An edit to a managed file survives until the next synchronisation and is then overwritten,
without a warning, taking whatever it was fixing with it.
## Consequences
- **A file edited on a node is a bug with a delay on it.** This is now one of the mesh's core
values, and it is a consequence of this decision rather than a stance taken independently.
- To change a value you must know which tool owns it. That is a real cost, paid every time,
and the reason the mesh provides a way to ask whether a given file is managed.
- Debugging by editing a file no longer works, and fails in the most confusing way available:
it works, and then stops working later for no locally visible reason.
- Recovery is cheap. A node's entire managed surface can be regenerated from the database.
- Values resolve by precedence — database override, then existing file value, then generated,
then manifest default — which means an unset override does not clobber a generated
password. The subtlety is real and has caused its own confusion.
## References
- `Extract hal/env-sync module, remove syncEnvToDb`, 2026-04-03 — the commit that removed the
node-to-database direction.
- Knowledge base: `conventions/no-direct-file-mutation`, `provisioning` (resolution priority).
- Regeneration gaps: `troubleshooting/changed-manifest-default-not-rerendered`,
`troubleshooting/config-removed-from-manifest-not-pruned`.
@@ -0,0 +1,67 @@
---
status: accepted
date: 2026-04-06
deciders: jochen
reconstructed: true
---
# 5. Capabilities are provisioned on declaration, not configured by hand
> Reconstructed after the fact from the evidence cited below.
## Context
Most modules need something another module holds — a database, a cache, a bucket, a message
vhost, an identity client. Wiring that by hand means creating the resource, creating a user,
generating a credential, putting it in the consumer's configuration, and repeating all of it
on every node the consumer runs on.
Every step is a place to make a mistake that surfaces much later, and the credential ends up
written somewhere it can be read.
## Considered options
1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a
description of the mistakes people will make.
2. **A shared credential per resource type**, distributed to all consumers. Rejected: no
isolation, and rotation becomes a mesh-wide outage.
3. **Declared requirements, satisfied by the provider module.** Chosen.
## Decision
A module declares what it **provides** and what it **requires**. A requirement names the
provider, the resource type, optionally a name and a target node, and a mapping from the
resource's connection fields to the consumer's environment variables.
The mesh satisfies it: a provisioner belonging to the provider creates the resource and its
credential, records the grant, and writes the mapped values as database overrides. The
synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then
materialises them. Neither the credential nor the topology is ever written by hand.
A requirement may name a provider on another node. The grant records consumer and provider
nodes separately, so cross-node wiring is the same declaration.
## Consequences
- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a
capability; where it lives is the mesh's problem. This is the property
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model
around.
- Credentials are never authored, so they are never authored badly, and they are never in the
repository.
- Each consumer gets its own credential, so revocation is per-consumer.
- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the
peers holding the old one, and this has taken the mesh down. The declaration model makes
granting easy and says nothing about fan-out.
- A module with no requirements skips the stage entirely, which is correct and also means the
absence of provisioning is indistinguishable from provisioning that did not run.
## References
- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the
provisioner daemon becomes its own component.
- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then-
environment-then-start sequence becomes the coordinator's.
- Knowledge base: `provisioning`, `provisioning/requires`.
- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`,
`troubleshooting/provision-adoption-rotates-live-credential`.
@@ -0,0 +1,68 @@
---
status: accepted
date: 2026-05-14
deciders: jochen
reconstructed: true
---
# 6. Schema and state changes are numbered migrations, in the same language as the code
> Reconstructed after the fact from the evidence cited below.
## Context
Modules own persistent state. That state has to change as they change, across nodes that are
at different versions, some of which have data that predates the change.
Two things were being done that do not survive contact with a second node. Schema was created
at startup, so what a table looked like depended on which version last started. And migrations
were shell scripts, so they could not use the types, connection handling or helpers the module
already had, and were not compiled or checked with it.
## Considered options
1. **Startup SQL / create-if-missing.** Rejected. It converges only for a node that started
with the newest version. A node that never restarts never migrates; a node that restarts on
an old version can undo a change.
2. **Shell migrations.** Rejected. Unchecked, untyped, and a separate dialect from the module
they belong to. Also, as later discovered, packaged differently — and therefore
occasionally not packaged at all.
3. **Numbered migrations in the module's own language, compiled with it.** Chosen.
## Decision
Every schema or state change is a numbered migration file, written in the same language as the
module and compiled with it. There is no startup schema creation and no ad-hoc statement.
Rules that come with it:
- The initial migration is **frozen** once it has run anywhere. It is never modified; a change
is a new number.
- Every statement is **idempotent** — guarded so that re-running is safe.
- A change needs **both** a baseline for a fresh installation and an incremental migration for
installations that already exist. Code referencing a column requires that the migration
creating it exists.
- Migration numbers are unique. A duplicate prefix is a defect, not a style issue.
## Consequences
- A node at any version converges to the current schema by running the migrations it has not
run.
- Migrations are checked by the same compiler as the code, and a migration that does not
compile fails the build rather than the deployment.
- The rules are enforced unevenly. Duplicate prefixes have shipped repeatedly and been fixed by
renumbering afterwards; the mesh now validates for them, which is the check this rule needed
in order to be real.
- A migration directory is a feature like any other, which means it is packaged like any other
— and when packaging is wrong, migrations silently do not ship. This has happened.
- Freezing the initial migration means a fresh installation replays the entire history. That
cost grows and nothing currently bounds it.
## References
- `fix(agents): convert workflow migrations to TypeScript` (#50), 2026-05-14.
- `feat(dev_validate): guard against duplicate migration numeric prefixes` (#379), 2026-06-26 —
the rule acquiring a check.
- Knowledge base: `migrations/schema-drift`, `noxflow/troubleshooting/migration-number-collision`.
- Packaging failures: `troubleshooting/shell-migrations-never-packaged`,
`troubleshooting/provision-migration-never-applied`.
+65
View File
@@ -0,0 +1,65 @@
---
status: accepted
date: 2026-06-04
deciders: jochen
reconstructed: true
---
# 7. No workspace — each module is a standalone package consuming published dependencies
> Reconstructed after the fact from the evidence cited below.
## Context
Modules depend on each other, above all on the shared library every module builds against.
A workspace was the obvious way to express that: sibling packages, resolved locally, one
install at the root.
It produced a divergence that is worth stating precisely, because it is not obvious. In
development, a workspace member importing a sibling resolves to that sibling's **local source**.
In the pipeline, each module is built alone, from a clone, without its siblings present — so
the same import resolves to the **published version**. The two environments were therefore
building different code from identical source, and the failure appeared only in the pipeline,
in a module that had not been touched.
## Considered options
1. **Keep the workspace and make the pipeline replicate it** — clone every module, build the
graph. Rejected: it makes every build a whole-repository build, which is the cost the
per-module pipeline exists to avoid, and it does not extend to modules in their own
repositories.
2. **Keep the workspace and pin siblings to published versions.** Rejected as the worst of
both: the workspace's local resolution silently overrides the pin, so the divergence
remains while looking solved.
3. **No workspace. Every module is standalone and consumes published dependencies.** Chosen.
## Decision
There is no workspace. Each module is an independent package that declares its dependencies
and consumes them from the private registry, including the mesh's own shared library.
A cross-package change is therefore two steps: publish the producer, then consume it. The
pipeline does the first on push and resolves the levels so that a module always builds against
its dependencies' freshly published versions.
## Consequences
- Development and the pipeline resolve imports identically. The divergence is gone by
construction rather than by discipline.
- A module in its own repository is not a special case. It builds exactly as a module in the
monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md)
cheap.
- A cross-package change costs a publish-and-consume round trip. This is the real price, paid
on every shared-library change.
- There is no repository-wide install and no repository-wide build. Anything that assumed one
broke, and one thing that assumed one has stayed broken: the end-to-end pipeline harness has
not built since this decision landed. See
[`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).
## References
- `fix(noxflow): kill npm workspace, restore encryption inside PgAdminRepo` (#240),
2026-06-04. The reason is recorded in the root package manifest, which still carries the
note explaining why no workspace exists.
- The divergence it fixed is named there: workspace members importing each other resolved to
local unbuilt source in the pipeline.
@@ -0,0 +1,71 @@
---
status: accepted
date: 2026-06-05
deciders: jochen
reconstructed: true
---
# 8. A step that fails must fail the job
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh's expensive faults are not crashes. They are the operations that reported success and
did nothing: an artifact that partially downloaded and was extracted anyway, a package that
404ed from every mirror while the job went green, a hook that never ran because it was named
for a feature the module does not declare, a deploy that reported the transport succeeded
rather than that the effect happened.
Each of these was found long after it happened, by someone investigating an unrelated symptom.
The cost is not the failure; it is the interval between the failure and anyone learning of it,
during which decisions are made on the assumption that the thing worked.
## Considered options
1. **Continue on error and report at the end.** Rejected — it is largely what existed. A
summary nobody reads is not a report, and later steps run against the state the failed step
should have produced.
2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a
precise, located failure into a vague one discovered elsewhere, and requires a health check
for every possible partial state.
3. **Fail the step, fail the job, say which step.** Chosen.
## Decision
A step that fails stops the sequence it is part of, and the failure is surfaced where the work
was requested — not only in a log.
Concretely, and these are the forms it takes:
- A scripted sequence gates each step on the previous one. A directory change that fails must
stop the commands that assumed it.
- An artifact that does not fully download is not extracted.
- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must
mean the thing is running, not that a command returned.
- A template that cannot resolve a variable is not written half-rendered.
**Prefer failing to lying.** A green result that is not true costs more than a red one.
## Consequences
- Failures are noisier and land earlier, on the person who caused them.
- Some jobs that used to complete now stop. In every case examined so far, that job was
producing a partial result that something downstream trusted.
- This is a rule the mesh has adopted repeatedly rather than once, because each instance is
written in a different place — a shell hook, a download path, a deploy stage. It is not
enforced by a mechanism, and cannot currently be checked in general. New instances are still
being found; the package-install case remains open as
[`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md).
## References
- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05.
- `A flavor template with an unresolved variable is written to disk instead of failing`
(#710), 2026-08-08.
- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`,
`troubleshooting/service-started-is-not-ready`,
`troubleshooting/green-pipeline-means-transport-not-effect`,
`troubleshooting/silent-failures-and-stale-state`.
- The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must
be loud."
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 9. The mesh is governed by a constitution, injected where work is decided
> Reconstructed after the fact from the evidence cited below.
## Context
By mid-2026 the mesh was doing a large share of its own design and implementation work through
agents. The rules those agents were expected to follow existed — in operating instructions, in
convention documents, in the knowledge base — but they were **retrieved**: an agent had to know
a rule existed in order to look it up.
Rules that must be looked up are followed by whoever already knows them, which is precisely the
population that does not need them. The rules being violated were the ones nobody thought to
search for.
## Considered options
1. **Documentation plus review.** Rejected — it is what existed. Review catches a violation
after the work is done, and only if the reviewer knows the rule.
2. **Lint and automated checks only.** Rejected as insufficient, not wrong. A check catches
what can be expressed mechanically; most of these rules are about judgement — what belongs
in a repository, when a criterion counts as verified.
3. **A canonical rule set, injected into context wherever work is decided, with a check phase
before output is accepted.** Chosen.
## Decision
A single canonical document states the mesh's non-negotiable rules. It is **injected
proactively** into every eligible design and analysis session — agents do not fetch it, it
arrives — and a check phase verifies the session's output against it before the work proceeds.
It is a governed document, not a page. Changing it requires a proposal, sign-off by reviewers
who are not the proposer, and a recorded decision. Drive-by edits are reverted.
Scoped override pages may **tighten** it for a team or product. They may never relax it.
## Consequences
- A rule reaches the work whether or not anyone remembered it existed.
- The check phase makes a violation a blocking outcome rather than a review comment.
- Two copies of the same rules now exist: this document, and the reasoning in HQ that earned
them. The enforced copy wins by default, so the reasoned copy quietly stops being true —
which is why [`00-META/how-we-build.md`](../00-META/how-we-build.md) is now the source
and the governed page is derived from it, via playbook
[`05-constitution-sync.md`](../00-META/process/05-constitution-sync.md).
- Injection costs context on every eligible turn, and grows with the document. Nothing
currently bounds that.
- The amendment process requires two reviewers, which a mesh with one human operator satisfies
only by counting agents. That tension is real and unresolved.
## References
- The governed page was authored 2026-07-10 and carries its own amendment process.
- `feat(noxflow): HAL architectural conformance gate for reviewer + architect` (#296),
2026-06-10 — the check phase, predating the document it checks against.
- Knowledge base: `platform/constitution`.
@@ -0,0 +1,64 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 10. Applications live in their own repository; the monorepo is for the mesh
> Reconstructed after the fact from the evidence cited below.
## Context
The module system makes adding anything to the monorepo trivial — a directory and a manifest.
That ease is the problem. Standalone applications, sites and side-projects accumulated beside
the mesh's own components, and once there they inherited the monorepo's review cadence, its
pipeline detection, and its history.
The mesh's own code and an application that merely runs on the mesh have nothing in common
except the manifest format. They change for different reasons, are reviewed by different
criteria, and have no reason to share a branch.
## Considered options
1. **Everything in the monorepo.** Rejected — it is what existed. The monorepo becomes an
inventory of one installation's applications, and every application change queues behind
mesh review.
2. **A second monorepo for applications.** Rejected: the same coupling with an extra name.
Applications have no more in common with each other than with the mesh.
3. **One repository per application, registered with the mesh as a build source.** Chosen.
## Decision
Every standalone application, site or side-project lives in its own repository, with a manifest
at the root. It registers with the mesh as a build source and is then built, provisioned,
deployed and verified by exactly the same pipeline as anything in the monorepo.
The monorepo holds the mesh: the runtime, the core modules, the delivery machinery, and the
shared infrastructure the mesh itself provisions against.
Creating an application directory in the monorepo is a convention violation, and reviewers
reject it.
## Consequences
- An application's cadence is its own. It is not reviewed as mesh code and does not queue
behind mesh work.
- The separation is safe **only because** the pipeline and provisioning are identical either
side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without
standalone packages this decision would fork the build.
- The monorepo stops being an inventory of the installation, which is a precondition for
publishing anything about it.
- Discovery gets harder: there is no single listing of everything the mesh runs, and the
registry of build sources becomes the closest thing to one.
- A module source that is not registered is silently skipped by the pipeline. The cost of
being outside the monorepo is that being forgotten is possible.
## References
- The rule is stated in the governed constitution page authored 2026-07-10, §3, as a
convention violation reviewers must reject.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
applications to the modules already in the monorepo.
- Knowledge base: `troubleshooting/unregistered-module-source`.
@@ -0,0 +1,59 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 11. The installer owns linking; nothing else creates a symlink
> Reconstructed after the fact from the evidence cited below. The incident that earned the rule
> predates the record, and its date is not established here.
## Context
A service's definition lives in the module catalogue; its runtime directory and persistent data
live outside it. The mesh connects the two by linking the definition into the runtime location
— deliberately, so that runtime state and source stay separate while the running service reads
a current definition.
A link is also the easiest thing in the world to create by hand while fixing something, and a
container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere
it should not have, and **production data was lost**.
## Considered options
1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss
for a service running a definition nobody can find.
2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of
the mistake — the link looks right and the resolution happens inside the container engine.
3. **One component owns linking; everyone else is forbidden.** Chosen.
## Decision
The installer creates and repairs every link the mesh needs. It reconciles them: a missing
link is created, a stale one is repointed, and a real file found where a link belongs is
adopted into the node's override location and replaced.
**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person
debugging. The prohibition is absolute because the judgement required to make a safe exception
is exactly the judgement that was not available at the moment it mattered.
## Consequences
- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who
has not seen the incident. That is why it is recorded here rather than only asserted.
- Links become reconcilable state rather than incidental filesystem facts.
- The rule is stated for humans and agents and is enforced by convention, not mechanism. A
check does not exist.
- The rule as written governs the mechanism rather than removing it. A link made by the
installer resolves the same way as one made by hand, so the hazard is narrowed and not
closed. [ADR 0018](0018-the-mesh-creates-no-symlinks.md) proposes widening this to "nothing
links, the installer included"; until that is accepted, this record governs.
## References
- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or
service directories have caused production data loss via Docker volume path resolution. The
installer handles all linking. Never create symlinks manually."*
- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files.
@@ -0,0 +1,74 @@
---
status: accepted
date: 2026-07-12
deciders: jochen
reconstructed: true
---
# 12. An agent is a persistent employee, not an instance of a pool
> Reconstructed after the fact from the evidence cited below.
## Context
Agents were originally a **pool**: a named kind of worker, scaled to some number of
interchangeable instances. Work went to whichever instance was free.
That model has no place to put the things that turn out to matter. An agent that accumulates
knowledge of a domain cannot keep it, because the next task lands on a different instance. An
agent cannot own a workspace, because there are several of it. It cannot be held to a policy —
warned for a violation, then dismissed — because there is no continuing subject to warn.
Scaling was also solving a problem the mesh does not have. Instances were being multiplied to
get concurrency, when concurrency is a property of how much work one agent may hold at once.
## Considered options
1. **Keep the pool, attach memory to the pool.** Rejected: shared memory across
interchangeable workers is a knowledge base, not an agent's experience, and the mesh
already has one.
2. **Keep the pool, make instances sticky.** Rejected as a pool pretending to be identities —
identity by scheduling accident, lost on any restart.
3. **One agent is one persistent identity, with concurrency as a property of it.** Chosen.
## Decision
An agent is a **singular, named, persistent identity**: a home node, a workspace on that node,
accumulating memory, and a lifecycle — hired, active, draining, retired. Not a pool member.
Concurrency is a property of the agent, not a count of copies: an agent has a cap on how many
sessions it may hold at once.
Lifecycle is explicit and has verbs. An agent is hired onto a node; it may be reassigned while
idle; it is retired by draining first, and forced only deliberately. Retired agents are not
deleted.
Surge capacity is expressed within the model rather than against it: a template agent is a
blueprint, cloned into a real agent with a lifetime when a queue grows, drained and retired
when it expires. A temporary employee is still an employee.
Some agents are **human**. What differs is modality — how the agent acts — not category. A node
itself is an agent of a kind exempt from the hiring lifecycle.
## Consequences
- Memory, workspace and reputation have a subject to belong to. Policy becomes possible: an
agent that violates a rule can be warned, and warned agents can be dismissed.
- The mesh gained a hiring model, and with it the question of who may hire.
- Scaling by adding instances is gone. If one agent is saturated, either its session cap rises
or another agent is hired — both deliberate acts.
- The transition was not free. Lifecycle columns had to reach every query that selects an
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises:
one kind of participant, differing only in modality.
## References
- `docs(adr): agents as persistent employees + MINERVA librarian` (#495), 2026-07-12 — the
original record, in the code repository.
- `feat(B4): one persistent employee, N sessions — rename max_instances → max_sessions` (#547)
and `feat(noxflow): B3 — workspace provisioner for agent employee model` (#549), 2026-07-20.
- `feat(noxflow): warn-then-fire agents who merge to main without review` (#209), 2026-06-01 —
policy that presumes a continuing subject, predating the model that provides one.
- Knowledge base: `agents/employee-lifecycle`, `agents/temp-surge`, `agents/workspace-layout`.
- The migration cost: `troubleshooting`/`noxflow-agent-enriched-select-missing-lifecycle-columns` (#546).
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 13. An artifact is build output, never a source tree
> Reconstructed after the fact from the evidence cited below.
## Context
A module is built once and deployed to every node assigned to it. What travels between those
two events is the artifact.
For a long time the artifact was a filtered copy of the module's source directory. Deploying it
therefore meant resolving and installing its dependencies **on the target node** — which
requires the target to reach a package registry, at deploy time, for every node, every deploy.
A node with no route to the registry could not deploy code that had already been built
successfully.
## Considered options
1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy
becomes a network operation with a failure mode per node, and the code that runs is
assembled independently on each one.
2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the
dependency resolution's platform assumptions along with it.
3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen.
## Decision
The artifact is the module's **build output directory** — compiled and bundled, with its
dependency graph inlined. Deploy is extract-and-run and touches no network.
A build that cannot produce a self-contained output **fails**. It does not fall back to
shipping a dependency tree, because a fallback that works is a fallback that is never fixed —
an application of [ADR 0008](0008-a-failed-step-fails-the-job.md).
## Consequences
- A node can deploy without reaching a registry. What was built is what runs, identically, on
every node.
- Deploys are faster and their failure modes are local.
- **Everything not in the build output does not ship.** This is the decision's whole cost, and
it was paid several times before it was understood: migrations that read the source layout,
provisioning scripts that read the source layout, selection files never packaged at all. Each
worked in development, where the source is present, and silently did nothing after deploy.
- Any file a module needs at runtime must be deliberately placed into the build output. The
rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code,
and that generalisation was the expensive part.
- Bundling has its own failure modes that a compiler will not catch — a bundler can exit
successfully and produce output that cannot load.
## References
- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04.
- The consequences, in order: `Provision migrations and seeds read the source layout, not the
artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07.
- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`,
`troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`,
`troubleshooting/esbuild-silent-tla-breakage`.
@@ -0,0 +1,78 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 14. Build, publish and deploy are three silos with different cardinality
> Reconstructed after the fact from the evidence cited below.
## Context
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
run a different number of times.
- Compiling happens **once per module feature**, on the build node.
- Packaging and uploading happens **once per module feature**, on the build node.
- Installing, configuring, starting and verifying happens **once per module feature per node**.
Conflating them is what made earlier versions slow and hard to reason about. Work that should
happen once was being repeated per node, and the fan-out point was implicit rather than a
boundary anything could observe.
The split had been declared before it was real. Packaging still happened inside the build,
which meant the boundary existed in the documentation and not in the code.
## Considered options
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
Cardinality is then a property of each stage's implementation, and nothing can reason about
the pipeline as a whole.
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
so build must know every module, every feature, and how each composes its artifact —
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
a stale package instead of re-packaging.
3. **Three silos, with an explicit handover between each.** Chosen.
## Decision
Delivery is three silos, and the boundaries are real:
| Silo | Runs | Where |
|---|---|---|
| **build** | once per module feature | the build node |
| **publish** | once per module feature | the build node |
| **deploy** | once per module feature **per node** | every assigned node |
Commands and events are addressed **per feature**, not per module.
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
publishing, so a module whose artifact is a package publishes in the publish silo, not the
build one.
Modules are resolved into dependency **levels**, and a level completes before the next begins,
so a module always builds against its dependencies' freshly published versions.
## Consequences
- Work that should happen once happens once. The fan-out point is explicit and observable.
- A failed upload retries by re-packaging, because packaging belongs to the stage that
uploads.
- The handover is a staged tree in a known location rather than the build's working directory,
which is reference-counted and cannot be assumed to still exist when a later stage runs.
- The build node is now the only node that has already passed through two silos when the
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
stages; code that knew only about the first parked the build node forever while every other
node deployed cleanly.
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
reports success over a stall it cannot see.
## References
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
documents claimed otherwise and were stale until 2026-08-06.
- The build-node stage-tracking failure was observed on pipeline #5557.
@@ -1,8 +1,11 @@
# 1. The mesh brokers capabilities; nodes host; agents think
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
- **Status:** Accepted
- **Date:** 2026-08-22
- **Deciders:** jochen
# 15. The mesh brokers capabilities; nodes host; agents think
## Context
@@ -175,7 +178,7 @@ existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
- [`01-RESEARCH/001-module-domain-decomposition`](../01-RESEARCH/001-module-domain-decomposition/analysis.md)
— current-state evidence, table counts, open questions
- [`00-GENESIS/how-we-build.md`](../00-GENESIS/how-we-build.md) — naming and integration rules
- [`00-META/how-we-build.md`](../00-META/how-we-build.md) — naming and integration rules
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
array that makes a feature a singleton per module
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
@@ -1,12 +1,15 @@
# 2. A lab node is a virtual machine running the real install
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
- **Status:** Accepted
- **Date:** 2026-08-22
- **Deciders:** jochen
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
@@ -128,7 +131,7 @@ than staging means every certificate experiment on a real node consumes issuance
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`02-DESIGN/01-end-to-end-testing.md`](../02-DESIGN/01-end-to-end-testing.md) — what the lab
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
@@ -0,0 +1,97 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0015-mesh-brokers-nodes-host-agents-think.md
---
# 17. Modules outside the platform core are grouped by domain, not by single function
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules
into bounded contexts named after their aggregates, and sends the rest out of the monorepo on
the grounds that they run *on* the mesh rather than being *of* it.
That leaves the larger half unaddressed. Around three quarters of the catalogue are modules
that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an
SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of
media-library services. Today each is its own module, because one module is the unit of *one
piece of software*, and no other grouping exists.
The result is that the catalogue's shape records what was installed, not what anything is for.
Four modules that together constitute "how a node is reachable" have no relationship the mesh
can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a
change to how the mesh handles connectivity has to be made four times.
This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident
rather than by domain* — appearing outside it.
## Considered options
1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else
keeps accident boundaries, so the catalogue becomes harder to read after the refactor than
before it.
2. **One module per piece of software, with a tag or category field.** Rejected. A label is not
a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it
drifts from the thing it labels.
3. **Group them into domain modules, each owning the software that serves one purpose.**
Proposed here.
4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for
the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it
into that model repeats the metaphor-naming mistake ADR 0015 exists to correct.
## Decision
*Proposed — the principle is settled; the domain list is not. See "Open" below.*
Modules that are not part of the platform core are grouped into **domain modules**. A domain
is named for the concern it serves, and owns the software that serves it. The unit stops being
one piece of software and becomes one purpose.
This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own
domain stand unchanged. This decision covers what ADR 0015 leaves outside them.
Naming follows the same rule as the core: **name the domain for what it does, not for what it
is made of**. Connectivity, not a VPN implementation.
## Consequences
- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes
reach each other is a change to one module.
- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from
its module list.
- Swapping an implementation stops being a module replacement, with the data-volume and
provisioning consequences that carries, and becomes a change inside a domain.
- The count drops sharply, which is a symptom of the improvement rather than the point of it.
- **Grouping conceals.** A domain module hides which implementation is in use, and every
operational question — which port, which unit, which credential — gains an indirection.
- The migration is not free and has no obvious increments: a domain is only useful once
everything belonging to it has moved.
- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its
own sake would be the same error in the other direction.
## Open
**The domain list is not settled and this record does not invent one.** What is decided is the
principle; what is not decided is the set. Candidate groupings are visible in the catalogue —
connectivity and reachability, node presentation and desktop, media libraries, observation and
metrics, storage and data services — but naming them here would be reconstructing a decision
that has not been taken.
Settling the list is a research effort, not an act of this record. Until it concludes, this
ADR stays `proposed`. That effort is
[`01-RESEARCH/005-domain-grouping`](../01-RESEARCH/005-domain-grouping/00-overview.md), and its
first measurement already narrows this record's scope: co-change analysis supports grouping for
reachability, argues against it for the provisioned infrastructure providers, and finds no
signal either way for the fifty modules that never change alongside anything.
## References
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this
extends, and its rule about naming a context after its aggregate.
- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are
already out of scope here; they are not domains and do not group.
- [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md)
— the catalogue's current shape, which is the evidence for the problem.
@@ -0,0 +1,101 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0011-the-installer-owns-linking.md
---
# 18. The mesh creates no symlinks — a derived file is a copy
## Context
[ADR 0011](0011-the-installer-owns-linking.md) responded to production data loss — a hand-made
link, resolved through a container engine's volume handling, pointing a mount somewhere it
should not have — by centralising linking in the installer and forbidding it everywhere else.
That narrowed the incident class. It did not close it. The hazard is not *who* made the link;
it is that a path can resolve somewhere other than where it appears to. A link made by the
installer resolves exactly the same way as a link made by hand. The rule made the mechanism
rarer and better-governed while leaving the mechanism in place.
Two things have changed since, and together they remove the argument that kept it.
**The original case for linking was staleness.** A copy of a service definition goes stale
silently while the catalogue moves on, so a link was the cheap way to guarantee the running
node reads a current definition. That argument assumes the node's copy is unmanaged.
**It is not.** [ADR 0004](0004-managed-files-are-generated-never-edited.md) established that
everything on a node's disk is derived from the mesh and regenerated when its inputs change,
and the installer already **reconciles** links rather than assuming them — repointing stale
ones, adopting real files it finds where a link belongs. Reconciling content is the same
operation as reconciling a pointer, plus a comparison.
So the mesh already has the machinery that makes a copy safe, and is using a link to solve a
problem that machinery solves better. Worse, a link is conceptually the wrong shape: it makes
the node's runtime state a *pointer into source*, which is the one thing
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) and ADR 0004 exist to prevent.
State is derived onto nodes; it does not reach back.
## Considered options
1. **Keep ADR 0011 as the final position** — centralised linking, forbidden elsewhere.
Rejected as the status quo. It governs the mechanism rather than removing it, and the
failure it was written for remains reachable by any code path the installer trusts.
2. **Keep links but harden them** — canonicalise before mounting, refuse a link that escapes
an expected root. Rejected: it is a check bolted onto a hazard, and it has to be correct in
every consumer, including container engines the mesh does not control.
3. **Copy, reconciled by the installer, with staleness detected rather than assumed away.**
Proposed here.
## Decision
*Proposed — the position is settled; the migration is not designed. See "Open" below.*
**The mesh creates no symlinks.** A file a node needs is placed on that node as a real file,
derived from the mesh and reconciled by the installer like every other managed file
([ADR 0004](0004-managed-files-are-generated-never-edited.md)).
The prohibition in ADR 0011 stands and widens: it ceases to be "only the installer may link"
and becomes "nothing links, the installer included".
When this is accepted, ADR 0011 becomes superseded rather than edited — its reasoning is why
the rule exists at all, and the incident behind it is the reason anyone believes either record.
## Consequences
- The path-resolution hazard is removed rather than governed. There is no link for a container
engine to resolve, so the class of failure that cost production data is closed by
construction.
- A node's runtime state stops pointing into source. What a node holds is derived output, which
is what the mesh's model already says it is everywhere else.
- **Staleness becomes a real problem that must be answered, not assumed away.** This is the
cost, and it is the whole cost: today a link cannot be stale, and a copy can. The answer has
to be detection — the installer comparing what is on disk against what the mesh says should
be — and it must be loud, because a silently stale definition is exactly the failure shape
this mesh keeps producing ([ADR 0008](0008-a-failed-step-fails-the-job.md)).
- Reconciliation gets more expensive: comparing content rather than checking a pointer's
target, on every module, on every node.
- Disk usage rises, trivially, and is not a consideration.
- Existing links must be converted. A node mid-migration holds both forms, so reconciliation
has to handle finding a link where a file now belongs — the mirror image of the adoption it
already does.
## Open
- **How staleness is detected.** Content hash, version marker, or regeneration on every
reconcile. This is the decision that makes or breaks the change and it is not taken here.
- **Whether anything must keep a link** for reasons outside the mesh's control. If something
does, that is a finding worth recording rather than an exception worth granting quietly.
- **Migration order.** Converting a node's links is a change to how its services resolve their
own definitions, which is not a change to make everywhere at once.
Until those are answered this record stays `proposed`, and ADR 0011 remains the governing rule.
## References
- [ADR 0011](0011-the-installer-owns-linking.md) — the incident, and the rule this widens.
- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — the machinery that makes a
copy safe.
- [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
— what the installer does today, including reconciliation and adoption.
@@ -0,0 +1,69 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
supersedes: none
---
# 19. HQ is its own repository, and it is public
## Context
The mesh's reasoning — mission, research, design, decisions — began inside the code
repository, under a folder there. The objection to moving it out was specific and good: the
mesh already has an operational memory and a structured archive, and adding a third store
repeats the mistake that consolidation was meant to fix.
## Considered options
1. **Keep it in the code repository.** Rejected, but the objection it rests on is correct and
is answered rather than dismissed — see Consequences.
2. **Put it in the structured archive**, alongside the governed documents. Rejected: the
archive is not reviewable as a diff, and a design argument is exactly the thing that needs
line-by-line review and a branch.
3. **Its own repository.** Chosen.
## Decision
HQ is its own repository, and it is **public** — written for a reader who is not its author
and has no access to the mesh it describes.
Three reasons it is separate:
- **The cadence differs.** A decision changes when thinking changes, not when code changes.
Tying documents to a code branch merges them on the code's schedule.
- **The reviewers differ.** A design argument is not reviewed the way an implementation is,
and should not queue behind a build.
- **The scope is wider than one repository.**
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the
monorepo; documentation governing several repositories cannot live inside one of them.
Being public is not incidental. It is enforceable only because
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) made the code repository
node-agnostic: there is no per-node content to leak. Nothing here may carry routable
addresses, real domain names, hosting providers, node names, absolute paths, usernames,
credentials, or operational detail useful only to an attacker.
The test is whether a paragraph would still teach a stranger running an entirely different
mesh.
## Consequences
- A document and the code it describes can no longer land in one commit. Keeping them honest
is a discipline rather than a mechanism — which is why decisions are recorded as they are
taken, and why a document stating a rule must say how the rule is checked.
- Research must state evidence without identifying the mesh it observed. The shape of a
finding survives anonymisation; the instance does not travel.
- **The objection is answered by indexing, not by location** — the claim being that these
documents remain searchable beside everything else, one source with many surfaces.
**That indexing does not exist.** Checked 2026-08-23, it returns nothing. Until it does, the
objection stands unanswered and this repository is the third knowledge store it was argued
not to be. Recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
## References
- Supersedes the earlier position that documentation lives inside the code repository under a
folder there. That position was never recorded separately and has no record of its own.
- [`README.md`](../README.md) — the public-repository rule in full.
@@ -0,0 +1,67 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 20. Design is written in two layers: what is, and what is intended
## Context
HQ held only the intended mesh. Every reader had to already know the running system in order
to understand what the decisions were about, and a statement about current behaviour had
nowhere to live except inside a document describing an intention.
The consequence was invisible until looked for: an as-is claim inside a to-be document is
indistinguishable from the intention around it, so the document silently stops being true as
the system moves — and nobody can tell which half went stale.
The mesh also has a large body of shipped behaviour that nobody would choose again. It is not
design in the sense of "what we decided"; it is design in the sense of "what is there", and it
is exactly the part a person changing the system most needs.
## Considered options
1. **One layer, describing the target.** Rejected — the status quo. The running system goes
undocumented and the target document accumulates unmarked claims about it.
2. **One layer, describing what runs, with intentions only in decision records.** Rejected:
a decision record is an argument, not a specification, and a multi-part intention has
nowhere coherent to live.
3. **Two layers, declared per document, never mixed.** Chosen.
## Decision
`03-DESIGN` holds two layers, and every document declares which it is:
| Layer | Describes | Written from |
|---|---|---|
| `00-as-is/` | The mesh that exists | The implementation and the operational record |
| `01-to-be/` | The mesh being built toward | Decision records |
An as-is document **records what is, not what should be** — including behaviour nobody would
choose again. A layer that keeps only the good decisions is a brochure.
When a to-be design ships it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand: one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning.
Where implementation and intention disagree, the as-is document records the implementation and
says they disagree.
## Consequences
- A reader can tell, from the folder and from one frontmatter field, whether they are reading
a description or a plan. That distinction was previously unavailable at any price.
- Correcting an as-is document requires evidence from the implementation, not agreement — and
needs no decision record, because nothing was decided.
- Two documents must be kept current per subsystem instead of one. This is the cost, and it is
paid on every ship.
- Something that shipped differently from its design becomes a visible divergence rather than
a silently wrong document, and may deserve an issue.
## References
- [`03-DESIGN/README.md`](../03-DESIGN/README.md) — the layer contract and frontmatter schema.
- [`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/) — the first eleven as-is documents, written
2026-08-23 from the monorepo and the operational memory.
@@ -0,0 +1,57 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 21. Every workflow is a playbook, and agents operate through them
## Context
HQ stated a knowledge flow — research becomes design — and nowhere stated how anything moves
along it. What graduation required, who wrote the decision, what closed an effort, what
happened when something shipped: all of it was convention held in one person's head.
A large share of the work here is done by agents. An unwritten convention is not available to
an agent at all, so each one either invents a procedure or asks. Both produce a repository
whose shape depends on who last touched it.
## Considered options
1. **Convention, learned by reading existing documents.** Rejected — the status quo. It
transmits shape but not rules, and it transmits the mistakes along with the pattern.
2. **One long contributing document.** Rejected: it is read once, and the step someone needs
is never the step they are reading.
3. **A playbook per workflow, each with trigger, steps and outputs, wrapped by a thin skill.**
Chosen.
## Decision
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs
it, steps, outputs. Engineers and agents follow the same playbooks, and **agents must not act
outside them**.
Each playbook is wrapped by a thin skill that defers to it as authoritative and adds only
mechanical scaffolding — next free number, frontmatter block, where the file goes. The
playbook holds the reasoning; the skill holds the steps. When they disagree, the playbook
wins.
## Consequences
- An agent arriving with no context can act correctly, because the procedure is retrievable
rather than remembered.
- The playbooks are themselves reviewable. A bad rule can be found and changed, which is not
true of a convention.
- Duplication between playbook and skill is real, and is managed by making the skill thin and
naming the playbook as authoritative in the skill's first lines. Nothing prevents them
drifting; the constraint is that only one carries reasoning.
- A workflow with no playbook is a workflow agents will get wrong. Adding one is part of
adding the workflow.
## References
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the five playbooks
and the flow they implement.
- Modelled on the process layer in the sibling HQ repository for the PAPA platform, which
arrived at the same shape and the same thin-skill split.
@@ -0,0 +1,54 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 22. Status lives in frontmatter; cross-cutting views are generated
## Context
Status was carried in prose — a bold line near the top of a document saying what state it was
in — and indexes were maintained by hand. The decision-record index had already drifted from
the folder it described **after a single addition**, which is about as short a demonstration
as the failure mode offers.
A hand-maintained index is a copy of something the filesystem already knows. It is correct
only for as long as everyone remembers it exists, and its being wrong is silent.
## Considered options
1. **Prose status plus hand-maintained indexes.** Rejected — the status quo, already
demonstrably broken.
2. **A central status file.** Rejected. It centralises the drift rather than removing it: the
file and the documents disagree, and the file is the one people read.
3. **Machine-readable frontmatter per document; every cross-cutting view generated on
demand.** Chosen.
## Decision
Every document carries its state in YAML frontmatter — research overviews, design documents,
decision records, issue reports — with a schema stated in the section README.
**There are no central status files.** Every cross-cutting view — a status matrix, the
decision-record index, the open-issue list — is generated from frontmatter when asked for, and
never written to disk.
Prose does not restate status. One place, and two is one too many.
## Consequences
- A view cannot drift from what it describes, because it does not persist.
- Status becomes queryable. Inconsistencies — a closed effort with nothing in `became:`, an
`implemented` design with no owning repository — are findable mechanically, and the
generator reports them as flags rather than silently rendering around them.
- Frontmatter must be valid and paths in it must resolve, which is now something to check.
- A reader browsing the repository on a forge sees no index. That is the trade: the index is
correct and absent rather than present and wrong.
## References
- [`.claude/skills/hal-status/SKILL.md`](../.claude/skills/hal-status/SKILL.md) — the
generator, including the inconsistencies it flags.
- [`02-DECISIONS/README.md`](README.md) — the hand-written index that drifted, and its removal.
@@ -0,0 +1,60 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 23. Issues have a front door, separate from the operational memory
## Context
Findings that were nobody's task accumulated in a table inside the decision ledger — a package
install reporting success while installing nothing, a manifest key read by no code, an
end-to-end harness dead for months. They were measured, true, and unowned: a table row cannot
be assigned, diagnosed or closed.
The mesh already has an operational memory holding roughly a hundred and thirty entries,
indexed on symptoms. The obvious move — put these there — is wrong, and the reason is the
distinction worth recording.
## Considered options
1. **Leave them in the ledger.** Rejected: a ledger records decisions taken, and these are
the opposite — questions nobody has answered.
2. **Put them in the operational memory.** Rejected. That store answers *how do I fix this
occurrence*; these are *why does the design allow this at all*. Filing them there makes
them findable by symptom and unfindable as open questions, and nothing there has a state
that can be closed.
3. **A numbered issue folder in HQ, deliberately narrow.** Chosen.
## Decision
`04-ISSUES` is the front door for something wrong at the level of **design or governance**:
a rule enforced by nothing, a stated invariant that is false, a failure the design permits to
be silent, or a symptom whose owner cannot be found without the whole mesh in view.
One numbered folder per issue: the report with the symptom as observed and the evidence, and
a diagnosis document carrying the trail, dated, including what was ruled out.
**This is not a second copy of the operational memory.** An issue here is a question HQ must
*answer*; an entry there is an incident someone must *clear*. An issue whose answer is a
general lesson belongs in both — and the operational memory is searched first, because if the
answer is already there this was never an issue.
## Consequences
- A finding gets a number, a state and an owner, and closing it is a visible act.
- The symptom-to-component trail accumulates in a place where the whole mesh is visible, which
is where cross-component diagnosis has to happen.
- The boundary needs judgement on every report, and will sometimes be got wrong. Filing too
narrowly loses a finding; filing too widely rebuilds the operational memory here, which is
the outcome HQ's separation was argued against
([ADR 0019](0019-hq-is-its-own-repository.md)).
- Six issues opened on creation, all previously unowned observations.
## References
- [`04-ISSUES/README.md`](../04-ISSUES/README.md) — the boundary table and the frontmatter
schema.
- [`00-META/process/03-issues.md`](../00-META/process/03-issues.md) — the playbook.
@@ -0,0 +1,76 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 24. The folder numbering is the flow, and decision records run oldest first
## Context
Two orderings were wrong in ways that only show up when someone new reads the repository.
**The folders.** Decisions lived in an unnumbered folder that sorted after the numbered ones,
so the repository's most load-bearing content read as an annex.
The sibling HQ repository for the PAPA platform had already solved this and solved it
crookedly: its design folder existed from its initial commit, and when its decision folder was
finally promoted it took the **next free number** rather than its place in the sequence. That
repository now reads `01 research → 03 decision → 02 design`. A decision precedes the design
it authorises and is numbered after it. By the time this was visible, the design folder was too
settled to renumber.
**The records.** Fourteen decisions had been taken in implementation and never written down —
the broker, the module abstraction, the mesh database, the artifact, the three silos and the
rest. Meanwhile two records existed, holding numbers 0001 and 0002, for decisions taken last.
## Considered options
1. **Match the sibling repository exactly**, inheriting its ordering. Rejected: structural
parity is worth something, but not the cost of copying a scar the other repository would
not choose again.
2. **Leave the folder unnumbered.** Rejected — the annex problem, and it leaves an unexplained
gap for anyone arriving from the sibling repository.
3. **Number by position in the flow, and renumber the records chronologically.** Chosen, on
the grounds that this repository was four commits old and nothing outside it cited a
number. That is the only window in which either renumbering is free.
## Decision
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
So `01-RESEARCH`, `02-DECISIONS`, `03-DESIGN`, `04-ISSUES`. Following the folder numbers walks
the process in the order it happens.
**Decision records are a chronological ledger.** They run oldest first. The fourteen decisions
already taken in implementation were back-filled as records 0001–0014, each dated from the
history, each carrying `reconstructed: true` and saying so in its first lines, and each citing
the commit, pull request or knowledge-base entry it was recovered from. The two existing
records moved to 0015 and 0016.
A reconstructed record is not a transcript. Where the deliberation is not recoverable it states
what the alternatives were and why the chosen one won on the evidence available — not a
discussion that did not happen. Where a date is not establishable it says so.
The foundational folder is `00-META`, matching the sibling repository.
## Consequences
- The repository reads in process order, and the gap at `03` that a reader coming from the
sibling repository would notice is explained by this record.
- The design documents can cite reasoning instead of asserting rules, because the reasoning now
exists.
- Structural divergence from the sibling repository, deliberately, in exactly one place. It is
recorded here so that the difference reads as a choice rather than an accident.
- **Record numbers are now stable and renumbering is over.** This decision spends the one
window that existed; a future record takes the next free number regardless of its date.
- Reconstructed records carry a standing risk: they are the most confident-sounding documents
in the repository and the least directly witnessed. The `reconstructed` flag exists so that
is never invisible.
## References
- The sibling repository's restructure of 2026-07-13 moved its decision folder in a single
commit of twelve renames with no content change, alongside the same status-into-frontmatter
and playbook changes made here.
- [`02-DECISIONS/README.md`](README.md) — the format, and the note on reconstructed records.
@@ -0,0 +1,67 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 25. HQ is the source of the mesh constitution
## Context
[ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set,
injected into every eligible design session and checked before output is accepted. It lives in
the knowledge base, where the orchestrator reads it.
HQ separately carried a document stating overlapping rules with the reasoning that earned each
one. Two texts, one enforced and one not.
That arrangement has a predictable outcome and it is not a tie. The enforced copy wins by
default, because it is the one that blocks work. The reasoned copy quietly stops being true,
and the rules survive without the incidents that justify them — at which point a rule reads as
arbitrary, and an arbitrary rule is the kind people route around.
## Considered options
1. **The knowledge-base page is the source; HQ points at it.** Rejected, though it is the
honest description of what was already happening. It leaves the reasoning downstream of the
rule, and the reasoning is the part that makes a rule survive a challenge.
2. **Accept the overlap and let both stand.** Rejected: two authorities is no authority, and
the drift is silent.
3. **HQ is the source; the governed page is derived and published from it.** Chosen.
## Decision
[`00-META/how-we-build.md`](../00-META/how-we-build.md) is the source. The governed page the
mesh injects is **derived** from it — the rules without the reasoning — and is never edited
directly.
Publishing is a playbook step, not a manual act, and it ends with **reading the page back and
verifying the change is present**. A publish that reported success and did nothing is exactly
the failure class this mesh keeps producing
([ADR 0008](0008-a-failed-step-fails-the-job.md)).
Section numbering is stable, because the orchestrator and the review fragments cite sections by
number.
## Consequences
- One source, many surfaces — the same argument HQ's separation already rests on
([ADR 0019](0019-hq-is-its-own-repository.md)), applied to the rules themselves.
- Each rule keeps the incident that earned it, in a place that is reviewed as a diff.
- An edit to the derived page survives until the next sync and then vanishes. The playbook says
so, and nothing mechanically prevents it.
- **The sync is manual and is the weak point.** An unsynced rule is a rule the mesh does not
enforce, whatever the source says — so the playbook requires the failure to be stated rather
than passed over. This is the same class of gap as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md),
and it is worth watching for the same reason.
- The document grew from four rules to seven sections, because it now has to carry everything
the mesh enforces rather than only what someone thought to write down.
## References
- [`00-META/process/05-constitution-sync.md`](../00-META/process/05-constitution-sync.md) —
the sync, including the read-back.
- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it
exists.
@@ -0,0 +1,89 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 26. Every decision is a record; there is no ledger
## Context
HQ carried a decision ledger at its root: a chronological table of forty-one numbered
decisions, each with who decided and a pointer to where the reasoning lived. It was created
deliberately, to make decisions findable and to give a home to decisions too small to warrant
a document.
By the time the decision records were back-filled
([ADR 0024](0024-the-numbering-is-the-flow.md)) the ledger had become three things at once,
and only one of them was still needed.
Classified, its forty-one entries were: ten restating a record, eleven restating design
documents, fifteen describing how this repository works — with the reasoning in a README rather
than anywhere citable — three small rules with no home at all, and two superseded stubs.
So the ledger was mostly a copy. Worse, it was a **hand-maintained index**, which
[ADR 0022](0022-status-lives-in-frontmatter.md) had just finished rejecting for the
decision-record index on the grounds that it had drifted after a single addition. Keeping one
copy of that pattern while removing another is not a position.
It had also produced a naming collision that a directory listing makes plain: `DECISIONS.md`
beside `02-DECISIONS/`, holding different things.
## Considered options
1. **Keep the ledger.** Rejected. It duplicates the records, restates status, and is the exact
hand-maintained index this repository decided against elsewhere.
2. **Keep it, renamed, for small decisions only.** Rejected, and this is the option worth
arguing with — it is genuinely useful to record a decision without writing a document. But a
decision small enough to be one table row is almost always a **rule** rather than a
decision, and a rule belongs in [`how-we-build.md`](../00-META/how-we-build.md) where it is
enforced and where its reasoning is kept. That is where the three orphans went.
3. **Every decision is a record; nothing else.** Chosen. This is how the sibling HQ repository
for the PAPA platform works, and it has no ledger of any kind.
## Decision
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is
not recorded.**
`02-DECISIONS` holds every decision. There is no ledger, no index file, and no central status
of any kind. The chronological view — decisions in the order they were taken — is *generated*
from record frontmatter, which is what the ledger was actually for.
Content that was only in the ledger was rehomed rather than dropped:
| Was | Went to |
|---|---|
| Decisions about how this repository works | Records [0019](0019-hq-is-its-own-repository.md)–[0025](0025-hq-is-the-source-of-the-constitution.md) |
| Small rules with no record | [`how-we-build.md`](../00-META/how-we-build.md) — the package rule, and two already there |
| Lab decisions not stated in the design | [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) |
| "Deliberately not decided" | The research effort and design document each question belongs to |
| Unowned observations | [`04-ISSUES`](../04-ISSUES/) ([ADR 0023](0023-issues-have-a-front-door.md)) |
## Consequences
- One place to look, and nothing to keep in sync. The collision between the ledger and the
record folder is gone.
- Structural parity with the sibling repository on decisions, which
[ADR 0024](0024-the-numbering-is-the-flow.md) deliberately broke on folder numbering. The
divergence is now exactly one thing, and it is the one thing that was argued for.
- **Writing a record is now the only way to record a decision, and a record is more work than
a table row.** The real risk is that a small decision goes unrecorded because nobody wanted
to write a document. The mitigation is that a small decision is usually a rule, and
`how-we-build.md` takes rules cheaply — but this is a cost, not a solved problem, and it is
the thing to watch.
- The chronological view now depends on the generator existing and being run. It did not
before.
- Two superseded ledger stubs had no record of their own. The position that documentation
lives inside the code repository is now recorded only as superseded context in
[ADR 0019](0019-hq-is-its-own-repository.md); the system-container position is explained in
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md). Neither is lost.
## References
- The sibling PAPA HQ repository: root holds only agent instructions and a README; every
decision is a numbered record, and its graduation playbook has no path for an unrecorded
decision.
- [ADR 0022](0022-status-lives-in-frontmatter.md) — the hand-maintained-index argument this
applies consistently.
@@ -0,0 +1,112 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 27. The product is Novox Mesh; Nox is an identity, not a second system
## Context
The name `HAL` was never chosen. This began as a dotfiles repository, the first commits in
February 2026 adopt dotfiles and per-node overrides, and the name arrived with the code — as
recorded in
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md),
most of the current shape is inherited from that origin rather than designed for a mesh. The
name is part of the inheritance.
Three things make it worth changing rather than living with.
**It is borrowed, and borrowed badly.** HAL is the canonical *untrustworthy* machine
intelligence. For infrastructure whose entire proposition is that it manages your machines,
heals itself, and is trusted with credentials, that is an unhelpful flag to fly, and it is not
a name anyone owns.
**There is a name available that is owned.** The company is Novox. A product of Novox should
carry that lineage rather than a film reference.
**A platform and a persona are different things, and one name was doing both.** `HAL` named the
mesh *and*, implicitly, the thing an operator talks to. Those are separate concerns — the
platform is what runs; the persona is who answers.
## Considered options
1. **Keep `HAL`.** Rejected. Every reason to keep it is sunk cost, and the sunk cost is at its
smallest today.
2. **Rename everything to a single new name covering platform and persona.** Rejected: it
repeats the conflation that made `HAL` ambiguous.
3. **Separate the two: a product name and an identity.** Chosen.
## Decision
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths.
**`Nox` is an identity of Novox**, and specifically an **agent identity within the mesh's own
model** — a named participant, exactly as
[ADR 0012](0012-agents-are-persistent-employees.md) defines one. Not a separate product, not a
separate runtime, not a privileged path.
**Nox is the agent of the mesh, not of a node.** This is the part that carries weight:
- **Every node keeps its own identity.** That already exists and stays — a node is a named
participant with its own character, and addressing one directly remains possible and normal.
- **Nox is scoped to the whole mesh.** It is what the mesh is called when the mesh itself
speaks, rather than one machine within it.
- **Nox addresses node identities.** Asking Nox for something that lives on one node is Nox
talking to that node, not a human choosing a machine.
- **A human mostly talks to Nox.** It is the front door.
That last point makes Nox the concrete form of the vision in
[`00-META/mission.md`](../00-META/mission.md): *an agent states an intent and the mesh carries
it out — no console to open, no runbook to follow, no remembering which node holds which
thing.* Nox is who that intent is stated to. The mission described the behaviour; this names
the thing that has it.
Nox holds no private channel. Whatever it can do, it does through the same surfaces every other
agent uses — which is not a naming detail: a persona with its own path would be the one part of
the mesh with no human checkpoint, and the skeleton already rules that out.
`HAL` is retired.
**Timing is the substance of this decision, not an aside.** The skeleton in
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) is not built. Renaming
before it exists costs a search and replace across research documents. Renaming after costs the
same class of migration as everything else this repository is trying to avoid, and would
therefore not happen.
## Consequences
- **The as-is layer keeps `HAL`.** It describes what runs, and what runs is called HAL. The
to-be layer uses `mesh`. The rename is part of the migration, and the two-layer split
([ADR 0020](0020-design-is-written-in-two-layers.md)) is what makes holding both names
coherent rather than confusing.
- **Records 0001–0026 keep `HAL`.** They are immutable and they say what was decided when it
was decided. No record is edited for a name.
- Tier 2 cannot be `mesh-mesh`. The control plane is **`mesh-control`**; `mesh-broker` was
rejected because the substrate already contains a message broker.
- The namespace, environment variable prefix, service names and on-disk paths all change. In
the existing system that is a migration and is not attempted here.
- **`mesh` is a generic word**, and it already means something specific in infrastructure — a
service mesh is a different thing. Recorded as a known trade rather than an oversight: the
full name `Novox Mesh` is distinctive, and the short form is internal.
- The persona has a name before it has behaviour. That is the right order — it is an identity in
a system that already has a model of identities, so it needs no new machinery to exist.
- **Except in one respect, and it is a real gap.**
[ADR 0012](0012-agents-are-persistent-employees.md) binds every agent to a home node, one to
one, with a workspace on that machine. A mesh-scoped agent has no home node by definition, so
the model does not currently have a shape for Nox. Extending it — an agent whose scope is the
mesh rather than a machine — is a decision of its own and is not taken here.
- Two levels of identity now exist where there was one: the node, and the mesh. The distinction
has to stay visible in every surface, or "ask Nox" and "ask a node" collapse into each other
and it stops being clear who is answering.
## References
- [ADR 0012](0012-agents-are-persistent-employees.md) — what an identity is in this system, and
why `Nox` needs no separate mechanism.
- [ADR 0020](0020-design-is-written-in-two-layers.md) — why the as-is and to-be layers can
legitimately use different names for the same system.
- The dotfiles origin, and the naming inheritance it explains:
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md).
+89
View File
@@ -0,0 +1,89 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 28. HQ is company-scoped; the mesh is its first product
## Context
This repository was `hal-hq` — one product's headquarters, named for the product. Then the
product was renamed ([ADR 0027](0027-the-product-is-novox-mesh.md)), which forced the question
of what the repository is actually the headquarters *of*.
Two facts settled it, and both were checked rather than assumed.
**Novox already delivers other things.** The company's forge organisation holds live projects
beside the mesh, and they are registered as build sources — meaning the mesh already builds and
deploys them. They are not hypothetical future products; they exist and ship today.
**They are tenants, not peers.** They run *on* the mesh. Every one of them is developed,
delivered and hosted by it. So the mesh is not one product among several — it is the ground the
others stand on.
That distinction decides the scope. If the mesh were a product beside others, a per-product HQ
would be right. Because it is the substrate the company operates on, a decision about the mesh
is a decision about how the company works.
## Considered options
1. **`mesh-hq` — one HQ per product.** The safe choice, and the reversible one: a second
product creates its own HQ and shared practice graduates upward later. Rejected, knowingly,
because it models the mesh as a peer of things that are actually its tenants.
2. **A company HQ *and* a product HQ, from the start.** Rejected as ceremony — two repositories
for one operator, and the constitution's own YAGNI rule says not to.
3. **One company-scoped HQ, `novox/hq`, with the mesh as its first product.** Chosen.
## Decision
The repository is **`novox/hq`** — Novox's headquarters, not the mesh's.
It holds the reasoning behind what Novox builds. Today almost all of that is the mesh, because
the mesh is what Novox is building. That is a fact about the present, not a definition of the
repository.
**The scope of each document is fixed now, so the eventual split is mechanical rather than
archaeological:**
| Scope | Documents | Moves if products separate? |
|---|---|---|
| **Company** | [`how-we-build.md`](../00-META/how-we-build.md), [`process/`](../00-META/process/), [`repos.md`](../00-META/repos.md), this record and [0019](0019-hq-is-its-own-repository.md)–[0027](0027-the-product-is-novox-mesh.md) | No — they stay at the top |
| **Product (mesh)** | [`mission.md`](../00-META/mission.md), [`context.md`](../00-META/context.md), [`effect.md`](../00-META/effect.md), `01-RESEARCH`, `03-DESIGN`, `04-ISSUES`, records 0001–0018 | Yes — into a product section |
The folders are **not** restructured now. One product's content under a company name is
correct while there is one product's worth of it, and nesting before there is anything to nest
is the ceremony option 2 was rejected for.
## Consequences
- Engineering practice has a home that does not belong to the mesh. `how-we-build.md` — never
write to production directly, migrations for schema changes, runtime evidence for behavioural
criteria — is true of any Novox project, and its being in a mesh repository was always a
slight mislabelling.
- The constitution derived from it ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md))
can legitimately govern work outside the mesh. Under a product HQ it could not have, without
either duplicating or reaching across repositories.
- **The bet is not entirely forward-looking, and that is worth being honest about.** Novox
already has work that is *not* a mesh tenant — client engagements and at least one product
that is developed outside it. So the company genuinely has a scope wider than the mesh
**today**, which strengthens the case for a company HQ and simultaneously means the split in
the table above is closer than "some day". The table is not a precaution; it is a plan whose
trigger already half-exists.
- What has *not* happened yet is any of that work needing the constitution. That is the actual
trigger ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md)): the moment something
outside the mesh must be governed by the same rules, product-level content moves down a level
and this repository becomes what its name already claims.
- A new repository was created rather than the old one transferred — the forge predates the
transfer API. `jschoubben/hq` is left in place untouched; it is not the source of truth and
nothing points at it.
- The mesh's own documents now live one conceptual level below the repository they are in. A
reader arriving at `01-RESEARCH` should understand it as the mesh's research, not Novox's.
Nothing in the folder names says so, and that is the cost of not restructuring.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name that forced the question.
- [ADR 0019](0019-hq-is-its-own-repository.md) — why HQ is a repository at all. Unchanged; only
its scope moves.
+67
View File
@@ -0,0 +1,67 @@
# 02-DECISIONS
Architecture decision records — the "why" trail behind the rules in
[`00-META`](../00-META/) and the specifications in [`03-DESIGN`](../03-DESIGN/).
**Numbered `02` because a decision precedes the design it authorises.** Research concludes,
the decision is recorded here, and only then is the design written. Following the folder
numbers walks the process in the order it happens.
One file per decision, numbered, never deleted. A superseded record has its `status:` changed
and gains a pointer to what replaced it — **its text is never edited**. The reasoning that was
rejected is the expensive half to rediscover.
The records run in the order the decisions were taken, oldest first.
**Every decision is a record.** There is no ledger and no index file — if a decision is worth
recording it is worth a record, and if it is not worth a record it is not recorded
([ADR 0026](0026-every-decision-is-a-record.md)). A "decision" small enough to be one line is
almost always a **rule**, and a rule belongs in
[`00-META/how-we-build.md`](../00-META/how-we-build.md), where it is enforced and keeps the
incident that earned it.
## Frontmatter
```yaml
---
status: proposed | accepted | superseded
date: YYYY-MM-DD # when the decision was taken, not when it was written down
deciders: name
reconstructed: true|false # true when the record was written after the fact from evidence
superseded-by: # 02-DECISIONS/NNNN-....md, when status is superseded
extends: # 02-DECISIONS/NNNN-....md, when this record widens an earlier one
---
```
## Body
```
# N. Title in plain language
## Context what was true, with evidence
## Considered Options numbered, each with why it was rejected
## Decision what was decided
## Consequences what follows, including what got harder
## References commits, pull requests, knowledge-base entries, prior art
```
State evidence, not assertion. *"Zero of 124 modules declare `brain` as a dependency"*
outranks *"the dependency rule is not followed"*.
## Reconstructed records
Records 0001–0014 were written on 2026-08-23, after the decisions they describe. Records 0015 onward were taken as records. Those
decisions were taken in implementation rather than in a document; the records state what was
decided and the evidence it was decided from, and each carries `reconstructed: true` and says
so in its first lines.
A reconstructed record is not a transcript. Where the deliberation is not recoverable, the
options section states what the alternatives were and why the chosen one won on the evidence
available — not a discussion that did not happen. Where a date is not establishable it says so
rather than guessing.
## Index
The index is **generated, not maintained** — run the `hal-status` skill, which reads the
frontmatter of every record. A hand-written index drifts from the folder it describes, and
this one had already done so after a single addition.
-10
View File
@@ -1,10 +0,0 @@
# 02-DESIGN
The authoritative specification. Implementation is built against what is written here.
A document enters this folder only after the decision behind it is recorded in
[`adr/`](../adr/) and the research that produced it is marked `GRADUATED`.
Empty for now: the decomposition in ADR 0001 is decided but not yet specified. The first
entries will be the per-context designs — `hal/mesh` brokering, the `hal/stream` record,
and the feature model that `hal/delivery` owns.
+98
View File
@@ -0,0 +1,98 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh as it stands
A set of machines, each running the same runtime, each loading only the parts of the catalogue
it has been assigned. They hold no shared filesystem and make no direct connections to one
another. What makes them a mesh is a database that knows what should run where, and a message
broker that carries everything between them.
## Three nouns
**A node** is a machine that runs the runtime. Nodes differ in what they are assigned and in
what they can reach — some carry a public name, some sit behind a household connection with no
inbound route at all — and the mesh is designed so that difference stays a property rather than
becoming a special case. A node holds no authoritative state: everything it needs is derived
onto it and can be regenerated.
**A module** is a directory with a manifest, and it is the only unit the mesh installs. A
containerised service is a module. A set of capabilities with no service behind them is a
module. A bare marker whose whole content is that a node has it is a module. The mesh's own
components are modules on exactly the same terms as everything else it carries
([ADR 0002](../../02-DECISIONS/0002-everything-is-a-module.md)).
**An agent** is a participant. Some agents are human. What differs is modality — how the agent
acts — and not category: both hold identity, both act, both accumulate memory
([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)).
## Where truth lives
The repository defines **what exists**: the modules, what each declares, how each is built.
The mesh database defines **what runs where**: which node is assigned which module, at which
selection, with which overrides, plus the settings every node reads. No node-to-module mapping
is ever committed ([ADR 0003](../../02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md)).
Everything on a node's disk is **derived** from those two, and is regenerated rather than
edited ([ADR 0004](../../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)). A node that
loses its database keeps running from a local cache, which is deliberate and has the obvious
cost: the cache carries no indication of its own age.
## How anything moves
Nothing dials a node. Every node dials the broker outbound, owns an exchange named for itself,
and consumes from its own request queue
([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)). Three message shapes carry
everything: requests expecting a reply, commands instructing that a stage of work be done, and
events stating that something happened.
A capability that lives on another node is reached the same way a local one is. At startup a
node asks its peers what they host and creates a local stand-in for each remote capability, so
the caller does not know or care where the work happens. Credentials never travel: the call
goes to where the capability is.
## How change reaches a node
A push to the forge is the only trigger. What follows is three silos with deliberately
different cardinality: compile once, package and upload once, then install-configure-start-
verify **on every assigned node**
([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)). What travels between
build and node is a self-contained build output, so a deploy is extract-and-run and touches no
network ([ADR 0013](../../02-DECISIONS/0013-an-artifact-is-build-output.md)).
Modules are resolved into dependency levels and a level completes before the next begins, so a
module always builds against its dependencies as they were just published.
## What the mesh does for a module
A module declares what it **provides** and what it **requires**. The mesh satisfies the
requirement: it creates the resource, generates the credential, records the grant, and writes
the values where the module will read them. The module never learns which node its database
lives on, and nobody ever writes a credential by hand
([ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)).
This is the property the mesh's whole shape rests on, and it is why provisioning is treated as
a core concern rather than as plumbing.
## The shape of its failures
Worth stating in an overview, because it is the most consistent thing about the system: the
mesh's expensive faults are almost never crashes. They are operations that reported success
and did nothing — a download that half-completed, a hook that was never called because it was
named for a feature the module does not declare, a stage that reported it had dispatched a
message rather than that the effect happened, a package that 404ed from every mirror while the
job went green.
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) is the response, and it is applied
instance by instance rather than enforced by a mechanism. New instances are still being found.
That is an as-is fact, not a criticism: it is the single most useful thing to know about this
system before changing it.
+113
View File
@@ -0,0 +1,113 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0003-the-mesh-database-is-the-source-of-truth.md
---
# The mesh and its transport
Two things make a set of machines into a mesh: a database that holds every binding, and a
broker that carries every message. Neither is optional and neither is replaceable at present.
## The mesh database
One relational database holds the bindings. Its content divides cleanly:
| Holds | Describes |
|---|---|
| Node records | Which nodes exist, and each node's own properties — its name in the mesh, whether it carries a public name, its identity text |
| Assignments | Which node hosts which module, at which selection, and whether it starts automatically |
| Overrides | Per-node, per-module values that take precedence over anything the manifest generates |
| Mesh settings | Values every node reads — where the broker is, where the forge is, where the registry is |
| Grants | Which consumer holds which resource from which provider, with the credential |
The runtime loads this at startup. If the database cannot be reached it falls back to a local
cache and continues.
**The repository contains none of this.** It defines what exists; the database defines what
runs where. This is what makes the repository node-agnostic, and it is the property that lets
anything about the mesh be published at all.
### What the cache costs
Running from cache is the difference between a node that survives a database outage and one
that stops. It is the right trade and it has a cost that is worth naming: a node running from
cache looks identical to a node running from the database. There is no age on the cache and
nothing reports divergence, so a node can be running yesterday's assignment set indefinitely
without any signal that it is.
### Where the settings are
A mesh-level setting is a value every node needs and no node owns — the broker's location, the
forge's, the registry's. These live in the database rather than in any node's configuration,
so a node learns where the broker is from the mesh rather than from a file, and moving the
broker is a database change rather than a fleet-wide edit.
The circularity is real: a node must reach the database to learn where the broker is, and the
database is itself a module the mesh provisions. It is resolved by the initialisation script
that stands up the first node, which is the reason such a script exists separately from
everything else.
## The broker
Every node connects **outbound** to a single broker. Nothing ever connects to a node.
Each node declares a topic exchange named for itself and consumes from its own request queue.
A shared mesh exchange carries traffic addressed to no particular node — pipeline commands and
the events they emit.
Three message shapes, and only three:
- **Requests** expect a reply. This is how a capability on another node is called.
- **Commands** instruct that a stage of work be done. They are addressed by what is to be done
and consumed by whichever node is meant to do it.
- **Events** state that something happened, addressed to nobody. Anything interested subscribes.
### Consequences the design accepts
A node behind a household connection with no inbound route participates exactly as a publicly
named one does. This is the property the transport was chosen for.
A call to a node that is down **waits** rather than failing. Usually this is what is wanted. It
is also how the mesh's most confusing stalls happen: a command queued for a node that never
returns is a stall with no error anywhere, and the pipeline has produced this shape more than
once — an empty pipeline that never completes blocks every pipeline queued behind it.
Two consumers accidentally sharing one queue silently split the traffic between them, each
receiving half of what it expects. This has happened between a module's daemon and its
capability server.
The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation, and doing it wrong has taken the broker
down.
## Reaching a capability on another node
A node hosts some capabilities and can reach the rest.
At startup, a node asks every peer what it hosts. For anything hosted elsewhere it creates a
local stand-in that forwards over the broker. For anything hosted both locally and elsewhere
it wraps the local one so a caller can name a target.
The effect is that a caller does not know where a capability runs. The important half is what
does **not** move: the work happens where the capability is, so its credentials never leave
that host. A remote call transports a request and a reply, never a secret.
Discovery happens at startup and is bounded by a short timeout. A peer that is slow or absent
at that moment is simply not discovered, and the node runs without that capability until it
restarts. Nothing re-discovers on a schedule.
## Names and reachability
Nodes address each other by names that resolve on the mesh's own overlay, not on whatever the
underlying network provides. A node's mesh name is its overlay address; its public name, if it
has one, is a separate fact used by things outside the mesh.
Two lessons are embedded in that separation, both learned the expensive way. A name resolved
by local multicast discovery introduces a delay and a failure mode that appears on one node and
not others, so mesh names are not multicast names. And a node must not pin its own public name
locally: the duplicate record breaks resolution for everything else that needs it.
@@ -0,0 +1,113 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0006-schema-changes-are-numbered-migrations.md
- 02-DECISIONS/0007-no-npm-workspace.md
---
# Modules, manifests and features
Everything the mesh installs is a module: a directory with a manifest. There is no second
mechanism.
## What a manifest declares
| Declares | Meaning |
|---|---|
| Identity | Name and version. **Version is owned by the builder** — a hand-edited version is a defect, and reviewers revert it. |
| Environment | Every variable the module reads, with how each is produced: a static default, a generated secret, a value pulled from the node's own record, or a template composed from the others. A variable not declared here is invisible to the mesh and will not be generated, injected or audited. |
| What it provides | The resource type this module can provision for others, and the network on which it is reachable. |
| What it requires | Resources it needs from other modules, and the mapping from each resource's connection fields onto its own environment variables. |
| Service shape | The primary container, and the data directories that must exist with the right ownership before it starts. |
| Exposure | The public names this module's interfaces answer on, declared portably so the reverse proxy configuration can be generated rather than written. |
| Images | Container images this module builds, so the pipeline builds and publishes them before publishing the module. |
## Features are the unit of work
A module is not the unit the pipeline addresses. A **feature** is.
A feature is a kind of content a module can carry: a service, a set of capabilities, a
long-running process, managed configuration files, migrations, firewall rules, an installable
application. One module can carry several.
Features are **detected from directory contents**, not declared. A module with a capabilities
directory has that feature; a module with a daemon directory has that one. An explicit
declaration was supported and is now discouraged, because a declared list and the directory it
describes drift, and the directory is the one that is true.
Every pipeline command and event names a feature. There is no per-module build.
### What detection costs
Detection makes the manifest shorter and the truth singular, and it makes the directory
structure load-bearing in a way that is not obvious from reading a manifest. Renaming a
directory changes what a module *is*, silently. The recurring failure is a hook named for a
feature the module does not carry: it is skipped without complaint, and the change it was
supposed to make simply never happens.
An unknown key in a manifest is likewise accepted in silence — which is how a firewall rule
can appear to restrict a port and restrict nothing (see
[`04-ISSUES/003`](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)).
## Kinds of module
The kinds are not a type system — they are what the detected features add up to.
- **A service module** carries a container definition. It gets a runtime directory, generated
environment, data directories, and is started under supervision.
- **A capability module** carries capabilities and no service. It contributes what a node can
do, locally and to its peers.
- **A flag module** carries nothing but a manifest. Its presence in a node's assignment is the
entire content: it gates behaviour elsewhere.
- **Combinations** are ordinary. A database module is a service *and* a capability provider
*and* a provisioner.
## Selections
A module can ship variants of the same feature and a node takes the one that fits it — a build
for one accelerator or another, a configuration for a public node or a private one. The
artifact stays selection-blind; the choice is a property of the assignment, held in the mesh
database.
This is one of the areas where behaviour has repeatedly diverged from intent, in both
directions: selection files that were never packaged into the artifact at all, and a stale
staged override on a node that silently won over the newly selected one. Both classes are
recorded in the knowledge base; both presented as "the change did not apply" with no error.
## Dependencies between modules
Modules depend on each other, above all on the shared library they all build against. There is
**no workspace** ([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)): each module is a standalone
package consuming published dependencies, including the mesh's own.
The pipeline resolves modules into dependency **levels** and completes a level before starting
the next, so a module always builds against its dependencies as just published.
The cost is a publish-and-consume round trip for every cross-package change, and the absence of
any repository-wide build. One thing that assumed a repository-wide build has stayed broken
since (see
[`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)).
## Persistent state
A module that owns state owns its migrations: numbered, written in the module's own language,
compiled with it, frozen once they have run anywhere, and idempotent so that re-running is safe
([ADR 0006](../../02-DECISIONS/0006-schema-changes-are-numbered-migrations.md)).
Two kinds exist and the distinction matters: migrations against the module's **own** local
state, and migrations against a **provisioned** resource, which run on the node that consumes
the resource rather than on the node that built the module.
## A documented rule with no enforcement
Every module exposing capabilities is documented as required to declare the mesh's core runtime
as a dependency. **Zero of the catalogue's modules do.**
This is recorded here rather than quietly corrected, because it is the clearest instance of the
rule this repository states about itself: a rule whose enforcement does not exist is
indistinguishable from a wrong one, and costs more, because people believe it. Whether the rule
or the catalogue is wrong has not been decided.
+88
View File
@@ -0,0 +1,88 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md
- 02-DECISIONS/0004-managed-files-are-generated-never-edited.md
---
# Provisioning
A module states what it needs. The mesh makes it exist, generates the credential, records the
grant, and puts the values where the module will read them. Nobody writes a credential and
nobody writes a topology.
This is the mesh's core concern rather than its plumbing — the property everything else is
built on.
## The declaration
A **provider** declares the resource type it can create and the network on which that resource
is reachable from a container.
A **consumer** declares, per requirement: the provider module, the resource type, optionally a
name and a target node, and a mapping from the resource's connection fields onto its own
environment variables.
The consumer must also declare those variables as empty in its environment section. A mapped
field with no declared variable is dropped — the value is produced and then discarded, because
the generator only emits variables the manifest knows about.
## What happens
Inside the pipeline the sequence is orchestrated by the coordinator and is not optional:
1. The provisioner creates the resource and its credential, records the grant, and writes the
mapped values as database overrides.
2. The synchroniser regenerates the module's environment from those overrides.
3. Only then is the service started.
Each step waits for the previous one's event. A module that declares no requirements skips the
first step entirely — which is correct, and means the absence of provisioning is
indistinguishable from provisioning that did not run.
Outside the pipeline the same two steps exist as direct operations, for debugging. They are
not the normal path.
## What can be provisioned
Providers exist for relational databases of two kinds, an object store, a cache, message-broker
virtual hosts and access, an identity provider's clients, and a download category. Each returns
the connection fields a consumer maps from — host, port, user, password, and whatever else the
resource type implies.
A provider with no registered provisioner still records a grant, with an empty connection. This
is deliberate and easy to misread: the grant exists, so the requirement looks satisfied, and
nothing was created.
## Cross-node grants
A requirement may name the node whose provider should satisfy it. The grant records consumer
node and provider node separately, so a module on one node holding a database on another is
the ordinary case rather than a special one.
For a containerised consumer, the provisioner returns the provider's routable name rather than
a loopback address — the value has to be correct from **inside** a container on another
machine.
## Where this hurts
**Rotation has no fan-out.** A resource whose credential is shared by several consumers can be
rotated by provisioning a new one, and the peers holding the old credential are not told. This
has locked the mesh out of its own broker, and has caused a node's adoption to rotate a live
shared password without informing anything that held it. Granting is easy; regranting is not
modelled.
**A frozen password outlives its generation.** A generated password is written once. If the
resource's persistent data directory already exists from an earlier initialisation, the stored
credential and the generated one diverge, and the symptom is an authentication failure that
looks like a configuration error.
**Migrations against a provisioned resource run on the consumer's node**, not the build node,
and read the deployed artifact rather than the source tree. Both facts were wrong in the
implementation for a period during which new provisioning migrations silently never ran.
**A grant is not a check.** The record says a resource was provisioned. Nothing verifies it
still exists, still has the recorded credential, or is reachable from where the consumer runs.
+100
View File
@@ -0,0 +1,100 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md
- 02-DECISIONS/0013-an-artifact-is-build-output.md
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
---
# Delivery — from a push to a running node
One trigger, three silos, and a fan-out point that is the most consequential boundary in the
mesh.
## The trigger
A push to the forge. The forge calls a webhook; the receiving module verifies its signature
and emits an event. The coordinator resolves which modules the pushed commits affect, orders
them into dependency levels, and creates a pipeline per level.
There is no second path. A manual trigger exists for a module the push detection missed, and
using it routinely is a sign the detection is wrong rather than a workflow.
Detection is the pipeline's most fragile input. It has failed for reasons that have nothing to
do with the change: a webhook truncating its commit list on a large merge, a forge address
whose port broke the module-path matching. The characteristic outcome is the bad one — **a
merge that created no pipeline, and nothing said so**.
## Three silos
Cardinality is the whole point, and the three differ
([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)):
| Silo | Runs | Where | Does |
|---|---|---|---|
| **build** | once per module feature | the build node | compile and bundle into a self-contained output |
| **publish** | once per module feature | the build node | package that output and upload it; a package-registry feature publishes here |
| **deploy** | once per module feature **per node** | every assigned node | install, configure, start, verify |
Commands and events are addressed per feature, not per module.
Build hands over a **staged tree**, not a package. Packaging belongs to publish, so a failed
upload retries by re-packaging rather than by re-sending something stale, and build never needs
to know how each module composes its artifact. The handover is a tree in a known location,
because the build's own working directory is reference-counted and may be gone by the time a
later stage runs.
## The artifact
The artifact is **build output** — compiled and bundled with its dependency graph inlined —
never a filtered copy of source ([ADR 0013](../../02-DECISIONS/0013-an-artifact-is-build-output.md)).
A deploy is extract-and-run and touches no network.
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after deploy — migrations reading a source layout, provisioning
scripts reading a source layout, selection files never packaged at all.
## The fan-out, and the build node
After publish, work fans out to every assigned node. The build node is the only node that has
already passed through two silos when this happens, and that asymmetry has its own defect
class: anything advancing a node's stage must account for **both** pre-fan-out stages. Code
that knew only about the first parked the build node forever while every other node deployed
cleanly — and the recovery sweeper, which knew the same subset, reported nothing to recover.
A recovery mechanism that knows less than the thing it guards is worse than none, because it
reports success over a stall it cannot see.
## Levels
A level completes before the next begins, so a module builds against its dependencies as they
were just published. The shared library is at level zero, which means anything that breaks it
breaks the first module of every cascade.
## What green proves
**A green pipeline proves transport, not effect.** The stages report that a message was
dispatched and accepted, which is not the same as the thing being running, correct, or present.
This is the mesh's most consistent failure shape and it is not incidental to the design — it is
what the stage reporting currently measures. Documented instances include a service reported
started when the container command merely returned, an image pull failure that did not fail the
deploy, a package install that 404ed from every mirror while the job went green (see
[`04-ISSUES/001`](../../04-ISSUES/001-failed-package-install-reports-success/00-report.md)),
and a node left on old code after a failed artifact download with a version marker that had
already advanced.
A verify stage exists to close this gap. It has been built and, for a period, was never
scheduled, because the coordinator's stage list did not include it.
## What is not covered
The end-to-end harness for this pipeline has not built since the workspace was removed, and
nothing reports that nothing runs it
([`04-ISSUES/005`](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). The
pipeline's coverage is currently assumed rather than checked — which is the same class of claim
this repository exists to make people stop making.
@@ -0,0 +1,107 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0011-the-installer-owns-linking.md
---
# The node runtime, and how a node comes into being
Every node runs the same runtime. What differs is its assignment.
## Two modes, coexisting
The runtime runs in two modes at once on any node that needs both.
**Daemon** mode is a headless consumer: it connects to the broker, consumes the node's request
queue, routes each request to a local capability, and emits the node's lifecycle events. This
is what makes a node a participant — it is reachable whether or not anyone is logged in.
**Interactive** mode exposes the node's capabilities to a session on that machine over a local
protocol. Its capability surface is larger, because it includes stand-ins for every capability
discovered on peers.
The two are the same code with the same catalogue. A capability is written once and is
available to both.
## Anatomy naming, and what it obscures
The runtime's components are named after brain anatomy: an entry point that bootstraps, a
headless listener, an interactive surface, an installer daemon, and a provisioner daemon.
This is the mesh's most-cited naming problem and it belongs in the as-is layer because it is
what a reader will actually encounter. The names are evocative and describe nothing: the most
suggestive word in the system names the node runtime, and the component whose manifest says
"mesh messaging" is documented elsewhere as the interactive runtime. Anatomy makes attractive
names and poor boundaries.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) replaces this with names
taken from what each part owns. Until then, this is the vocabulary in the code.
## Starting a module
For each assigned module the runtime, at startup:
1. Reads the manifest, if there is one — a flag module has nothing to read.
2. Skips a module's capabilities if a variable they require is unset. This is quiet by design
and hard to distinguish from a module that has no capabilities.
3. Loads the capabilities the module carries.
4. For a service module, ensures it is installed and running.
Installation is idempotent and does the unglamorous work: ensure the runtime directory,
reconcile the link from the catalogue's definition into it, generate the environment, run any
outstanding local migrations, create data directories with the right ownership, then start the
service under supervision.
**The installer is the only thing that creates a link** ([ADR
0011](../../02-DECISIONS/0011-the-installer-owns-linking.md)). It reconciles rather than assumes: a
missing link is created, a stale one repointed, and a real file found where a link belongs is
adopted into the node's override area and replaced. Nothing else — not a hook, not a fix, not a
person debugging — creates one.
That is the as-is. The intent is to remove linking altogether and derive a real file instead,
which the reconciliation machinery already makes possible
([ADR 0018](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), proposed). What is described above
is what runs today.
## Supervision
Services run under the host's init system via a templated unit, one instance per module. It is
a thin layer: the unit starts and stops a container group.
Whether the mesh keeps this, drops the per-module layer, containerises the daemons, or writes
its own supervisor is **open** — costed in research effort 003 and deliberately undecided. It
no longer gates anything.
Two failures worth knowing about, both in the shape of "the change did not apply". A
per-instance copy of the unit template shadows the template, so edits to the template do
nothing. And a session-scoped one-shot job loses the environment it was given, because the
import is one-time and not persisted.
## How a node comes into being
Three bootstrap scripts, and which one runs depends on the situation:
- **First node.** Nothing exists yet, so the script stands up the database the rest of the
mesh reads from, publishes the catalogue, and starts the mesh. This resolves the
circularity of a mesh whose source of truth is itself a provisioned module.
- **Joining.** The node registers, takes its assignment from the database, and syncs.
- **Rescue.** A node that cannot reach the mesh is brought back far enough to.
Bringing a node into being is therefore a **database operation with a script attached**, not a
checkout. There is no per-node content in the repository to copy.
Adoption of a pre-existing machine's configuration was the original path and is now a legacy
one, explicitly out of scope for the lab
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)).
## Node identity
Each node carries an identity text in its own record, which the daemon writes onto the node at
startup so that a session on that machine knows which node it is on and how that node
presents itself. It carries identity only; shared rules live separately.
Like everything else derived onto a node, it is generated and not edited there.
@@ -0,0 +1,86 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0004-managed-files-are-generated-never-edited.md
- 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md
---
# Configuration and secrets
Every value a module reads at runtime comes from a file on a node's disk. Every one of those
files is **generated**.
## The rule
A managed file is derived from the mesh database. A synchroniser rewrites it when the values
behind it change. The write path is the mesh operation that owns the value; the file is an
output ([ADR 0004](../../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)).
An edit to a managed file survives until the next synchronisation and is then overwritten
silently, taking whatever it was fixing with it — bringing back the bug the edit had removed,
with a delay, and with no error to connect the two events.
There is a way to ask whether a given file is managed. That question has to be asked, because
the answer is not visible from the file.
## What is managed
Generated environment files for each module, service definitions in each module's runtime
directory, managed configuration files a module declares, and node-level settings. The list is
not a category — it is whatever a synchroniser claims, which is why the question is asked of
the tooling rather than answered from a rule.
## How a value is decided
Values resolve by precedence, highest first:
1. **A database override** — set deliberately, or written by a provisioner.
2. **The value already in the generated file** — preserved for anything without an override.
3. **A generated value** — a random secret, a template composed from other variables, or a
value pulled from the node's own record.
4. **The manifest default.**
Two consequences follow, and both are subtle enough to have caused confusion.
The second rule is what keeps a generated password **stable** across regenerations. It is not
an oversight; without it every regeneration would issue a new secret and break whatever holds
the old one.
The same rule means that **changing a manifest's default does not change anything already
using it**. The existing file's value wins. The new default reaches only installations that
never had one.
Removing a declaration is worse than changing it: the old override row and the file it produced
are both left behind. Configuration is additive in practice, whatever the manifest says.
## Secrets
Generated secrets are produced by the mesh, never authored. Provisioned credentials arrive as
database overrides written by the provisioner and are marked as such, so they can be
distinguished from a deliberate override and cleaned up when the grant is removed
([ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)).
Nothing in the repository contains a credential. The repository has no per-node content at all,
which is what makes that guarantee structural rather than a matter of care.
Two known weaknesses, both recorded rather than resolved:
**Generated environment files were world-readable.** The mode passed at write time only applies
when the file is created, so regeneration left the previous mode in place. The class of bug is
worth remembering beyond the instance: a permission set at creation is not a permission
maintained.
**Rotation is not a mesh operation.** Secrets can be generated and granted; there is no
mechanism that rotates one and informs everything holding it. Where a rotation has been done,
it has been done by hand, and doing it wrong has taken services down.
## Node-level and mesh-level values
A node-level value applies to everything on one node. A mesh-level value applies everywhere and
is read by every node — the broker's location is the canonical example, and a wrong one is how
a mesh fails to form.
Neither is a file that anyone edits. Both are database records that produce files.
+78
View File
@@ -0,0 +1,78 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions: []
---
# Knowledge
The mesh keeps two knowledge stores. They are not redundant, and knowing which is which is the
difference between finding an answer in one search and rediscovering it over several hours.
## The operational memory
A store of operational notes, written and read by whoever — human or agent — is working. Each
note is a slug and a body: how something works, what went wrong, what the fix was, what
assumption turned out to be false.
It is indexed on **symptoms**. The entry someone needs is usually titled after the error they
are staring at, which is why the standing instruction is to search the literal error text
before forming a hypothesis rather than after one fails.
Its content is overwhelmingly the record of previous debugging: a large body of
troubleshooting entries, module conventions, and standing notes about work that is open. It is
the mesh's institutional memory of *what has already gone wrong*.
The cost of skipping it is documented in the mesh's own record: entries have been rediscovered
from scratch, over hours, in sessions where the search was skipped because the trail felt
confident. It fires hardest on familiar ground, not unfamiliar ground.
## The structured archive
A second store, structured rather than flat: spaces, pages, revisions, tiers, and full-text
search. Where the operational memory is a note, this is a document with an owner and a
lifecycle.
Content is promoted through tiers — private, then team, then platform — with a librarian agent
owning approval and promotion at the boundary. Proposals to edit are reviewed rather than
applied.
This is where the mesh's **governed** documents live, including the constitution injected into
design sessions ([ADR 0009](../../02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md)).
## Why both
The distinction is by lifecycle, not by subject.
| Operational memory | Structured archive |
|---|---|
| Written the moment something is learned | Written deliberately, reviewed |
| Flat, symptom-indexed | Structured, tiered, owned |
| Anyone writes; nothing approves | Promotion is approved |
| Truth is "this happened" | Truth is "this is agreed" |
Collapsing them would cost one of the two properties: either every hard-won note waits for
review, or governed documents can be changed by anyone mid-incident.
## Where this repository sits
This repository is a third thing, and the objection was raised when it was created: a fourth
knowledge system repeats the mistake the split was made to fix.
The answer given was **indexing, not location** — that these documents are indexed into the
knowledge base so that a symptom search returns them alongside everything else. One source,
many surfaces.
**That indexing does not currently exist.** A search for this repository's content returns
nothing. The claim is load-bearing for the decision to separate the repository at all, and
until it is true, this repository is exactly the fourth knowledge system the objection
described. Recorded here because it is a statement about how the mesh's knowledge actually
works today, and as [`04-ISSUES/006`](../../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
## The librarian
A single agent owns the archive's approvals and promotions. Its approval capabilities have at
times not been reachable as tools, which does not affect the operational memory but does mean
promotion stops silently — the store keeps accepting proposals that nothing can approve.
+88
View File
@@ -0,0 +1,88 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0012-agents-are-persistent-employees.md
- 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md
---
# Agents and work
The mesh does a large share of its own design and implementation. Agents are how, and the
model they run under is the employee model, not a worker pool.
## An agent is an employee
An agent is a singular named identity with a home node, a workspace on that node, accumulating
memory, and an explicit lifecycle
([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)).
| Property | Meaning |
|---|---|
| Lifecycle state | Active, draining, or retired. Retired agents are kept. |
| Home node | Where its workspace lives. One node per agent. |
| Session cap | How much work it may hold at once. **Concurrency is a property of the agent, not a count of copies.** |
| Kind | Whether it is hirable, or is a node's own agent and exempt from hiring |
The verbs are explicit: an agent is **hired** onto a node, **reassigned** only while idle, and
**retired** by draining first — forcing it is a deliberate act that aborts work in flight.
Surge capacity lives inside the model rather than against it. A template agent is a blueprint
with no life of its own; when a queue grows past a threshold it is cloned into a real agent
with a lifetime, which drains and retires when that expires. A temporary employee is still an
employee.
Because there is a continuing subject, **policy becomes possible**: an agent that violates a
rule can be warned, and a warned agent can be dismissed. A pool cannot be warned.
## Some agents are human
There is one kind of participant. What differs is **modality** — a non-human agent acts through
a spawned session and the record; a human agent acts through a shell, a desktop, or a message.
Both hold identity, both act, both accumulate memory.
The mesh does not currently record modality completely. Which user, on which node, a human
agent acts as is **required by the model and not stored** — an open question carried over from
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md).
## Work
Work is expressed as tasks moving through workflows. A workflow names the states a kind of work
passes through and what must be true to leave each one; a task carries its acceptance criteria
and its trail.
Several workflow shapes exist for different sizes of work — a single implementation, a larger
container of related work, and shapes that add analysis or design stages ahead of
implementation.
The area's characteristic defects are **transition** defects rather than logic defects: a task
bouncing between review and implementation because a guard was evaluated on stale state, a
result that cannot be recorded in the same act as the transition it justifies. The workflow
engine's correctness is about atomicity, and that is where it has been wrong.
## Meetings
Some work is decided in a **meeting**: several agents in turns, with distinct roles, over a
template that names the phases.
This is where governance meets execution. The constitution is injected into every eligible
meeting turn — agents do not fetch it, it arrives — and a check phase verifies the meeting's
output against it before the meeting may proceed
([ADR 0009](../../02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md)). A named violation
blocks progress.
Meeting turns run on the orchestrator's node regardless of where the participating agents are
pinned. That is a known divergence between the model and its execution, not a design intent.
## What this rests on that is not built
The work domain shares one large schema with several other domains. That is the concrete
instance of a rule stated in [`how-we-build.md`](../../00-META/how-we-build.md) — *contexts
integrate through the record, never through a shared schema* — being violated by the mesh's
own largest component, and it is the reason work that belongs to one domain keeps having to be
implemented in another.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) dissolves that arrangement.
Until it does, this is the shape.
@@ -0,0 +1,78 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0001-nodes-communicate-over-a-broker.md
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
---
# Interfaces and observability
How the mesh is reached, and how anyone can tell what it is doing.
## Capabilities are the primary interface
The mesh's primary interface is not a web console. It is a set of **capabilities**, exposed to
a session and callable in language.
A capability is contributed by a module and is available on any node, wherever it actually
runs: local ones directly, remote ones through a stand-in created at startup that forwards over
the broker. The caller does not know the difference, and the credentials never move.
This is the mesh's stated vision made concrete — an agent states an intent and the mesh works
out which node holds the thing. It is also why a capability's **schema** is load-bearing in a
way that is easy to underestimate: a parameter name that collides with the transport's own
reserved names breaks the call, and a validation-library version mismatch has silently dropped
every argument while the call still appeared to succeed.
## The board
A web interface presents the mesh — nodes, modules, pipelines, agents, work. It is a **view**.
Its own guidance is that shared logic belongs in the mesh's library rather than inline in the
board, precisely so the board does not quietly become a second implementation of the mesh's
rules.
## Public exposure
Nodes carrying a public name run a reverse proxy as the sole entry point. A module declares the
names its interfaces answer on, portably, and the proxy's configuration is **generated** from
those declarations rather than written — generated files are marked as such and anything
hand-written beside them is left alone.
Nodes without a public role use a local equivalent with locally-trusted certificates. The
generation step degrades quietly on a node with no proxy, which is intended and is one more
place where "nothing happened" is the correct outcome and looks identical to a failure.
Certificate issuance currently always targets the public authority's production endpoint,
which consumes real quota for every experiment
([`04-ISSUES/004`](../../04-ISSUES/004-certificate-issuance-targets-production/00-report.md)).
## Health
Nodes run checks and report. The mesh's health surface answers whether things are up.
What it does **not** answer is whether they are correct, and that gap is the recurring theme of
this whole system: the deploy path reports transport rather than effect, so absence reads as
success. A check that confirms a service is running does not confirm the service is running the
code that was just deployed, and a node has been left on old code with a version marker that
had already advanced.
## Thoughts
Every node's daemon runs a periodic loop that surfaces observations from that node's own
context. They are stored in the mesh and can inform a session or trigger action.
It is the one part of the mesh that is not request-driven — the mesh noticing things rather
than being asked.
## The honest summary
Observability tells you the mesh is **up**. Establishing that it is **right** currently means
reading the operational record and checking by hand.
That is the gap the lab is designed to close
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)): a place where a
change can be run end to end and a verdict produced, cheaply enough that producing one is
routine.
+126
View File
@@ -0,0 +1,126 @@
---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0010-applications-live-in-their-own-repository.md
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
---
# The catalogue, and what its shape says
The catalogue holds **124 modules**. Thirty-three belong to the mesh's own domain; the other
ninety-one run *on* the mesh rather than being *of* it
([ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)).
The count is not the finding. The **shape** is.
## How it is organised today
By namespace, and the namespace records origin rather than purpose:
- **The mesh's own namespace** holds the platform: the runtime and its daemons, the shared
library, delivery, provisioning, configuration synchronisation, knowledge, identity, the
board, developer tooling, and node presentation.
- **A second namespace** holds the work domain — tasks, workflows, agents, meetings — split
across a handful of packages that share one schema.
- **Everything else sits flat at the top level**, one directory per piece of software.
## What the flat level actually contains
Grouped by what they are *for* — a grouping the catalogue itself does not express:
| Purpose | Roughly |
|---|---|
| Data and storage services the mesh provisions against | Relational and document databases, a cache, an object store, a package registry, a time-series store |
| Messaging and identity | A message broker, an identity provider |
| Reachability | A VPN, a firewall, an intrusion filter, an SSH daemon, a resolver, a certificate authority, a reverse proxy, network equipment control |
| Forge and container plumbing | Forge integrations, an image registry, container lifecycle and retention |
| Media libraries | Acquisition, organisation, playback, transcoding, streaming |
| Workstation and desktop | Browser, file manager, monitors, session management, audio, package management, runtime managers |
| Hardware-specific support | Power and firmware control for particular hardware, filesystem management |
| Collaboration and productivity | File sync, office tooling, boards, automation, chat and messaging bridges, mail, analytics, dashboards, home automation, issue trackers and wikis |
| Third-party organisation integrations | Systems belonging to organisations outside the mesh |
Every row is several modules, and **no row is a thing the mesh can see**. Four modules
together constitute "how a node is reachable", and they have no relationship the mesh can
assign, version, reason about or replace as one unit. A change to how the mesh handles
connectivity is made four times.
## What the shape records
**The catalogue's shape records what was installed, not what anything is for.** One module is
the unit of one piece of software, because that is the only granularity the module system
offers.
This is the same failure [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)
names for the platform core — *boundaries drawn by deployment accident rather than by domain* —
appearing outside it, at four times the scale. The core is being recomposed; the flat level is
addressed in principle by
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md), which
deliberately does not yet settle the domain list.
## Where the shape came from
The catalogue's shape is not arbitrary and it is not a series of mistakes. **This repository
began as a dotfiles repository**, and most of what looks inexplicable is inherited from that
directly.
The evidence is in the first two days of history: *adopt desktop dotfiles and scripts from the
original dotfiles repository*; *per-node dotfile overrides*; *service symlinking, node
`.dfignore`, headless server support* — all on 2026-02-24 and 2026-02-25, before anything
resembling a mesh existed. Forty-five commits mention dotfiles.
Read that way, several things stop being puzzles:
| Feature of today's shape | Dotfiles ancestor |
|---|---|
| One flat directory per piece of software | Exactly how a dotfiles repository is organised |
| **Linking rather than copying** as a stated design principle | How every dotfiles manager works — the whole category is built on it |
| **Adoption** — taking over a machine that already exists, with its own configuration | The core dotfiles verb, and the reason a machine could be brought in at all |
| **Per-node overrides** | Per-host dotfile overrides, generalised into per-node module settings |
| Desktop and workstation modules — browser, file manager, editors, session, audio, personal scripts | Dotfiles content that became "modules" when modules became the only unit |
| An ignore file controlling what is placed on a machine | A dotfiles manager's ignore file |
The generalisation from *place files on my machines* to *manage a mesh of nodes* was the right
move and it worked. What it did not do is revisit the assumptions underneath, because they were
never stated as assumptions — they were just how the thing already worked.
**This is the most useful single fact for anyone changing the catalogue**, and it is why the
linking principle in particular reads as a deliberate architectural choice when it is an
inheritance. See [ADR 0018](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), whose case
this strengthens: the argument for links was never made *for a mesh*.
It also explains the measurement in
[research 005](../../01-RESEARCH/005-domain-grouping/analysis.md). Fifty modules that never
change alongside anything are not fifty missing domains — many are dotfiles-era entries for one
piece of software, which never shared a domain because they never had one.
## Two properties worth keeping
Whatever replaces the shape, two things about it are right.
**Uniformity.** A media server and the mesh's own coordinator are installed, provisioned,
delivered and verified by identical machinery. The mesh's own components hold no privilege —
which is what makes dogfooding structural rather than a discipline, and what makes moving a
module out of the repository safe.
**Placement is already decided.** A standalone application belongs in its own repository
([ADR 0010](../../02-DECISIONS/0010-applications-live-in-their-own-repository.md)), and reviewers reject
it in the monorepo. The catalogue's flat level is not a dumping ground by policy; it is one by
history.
## Known inconsistencies in the catalogue itself
Recorded because a reader will meet them:
- A documented requirement that every capability-exposing module declare the core runtime as a
dependency is met by **zero** modules.
- A firewall-scoping key is declared by five manifests and read by none
([`04-ISSUES/003`](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)).
- A connections block in the manifest is metadata: it describes a module's reachability and
wires nothing.
- At least one module deliberately runs outside the standard per-module supervision, for
reasons recorded in the operational memory. The standard path is not universal.
+30
View File
@@ -0,0 +1,30 @@
# 03-DESIGN / 00-as-is
The mesh as it stands. These documents describe what runs, including the parts nobody would
choose again — an as-is layer that only records the good decisions is a brochure.
They are written from the implementation and from the operational record, not from intent.
Where the two disagree, the implementation wins and the disagreement is stated.
| Document | Covers |
|---|---|
| [`00-overview.md`](00-overview.md) | The whole in one pass — what a node is, what a module is, how work reaches it |
| [`01-mesh-and-transport.md`](01-mesh-and-transport.md) | The mesh database, the broker, discovery, and how a call reaches another node |
| [`02-modules-and-manifests.md`](02-modules-and-manifests.md) | The module, the manifest, and features as the unit of work |
| [`03-provisioning.md`](03-provisioning.md) | Declared requirements, provisioners, credentials, and cross-node grants |
| [`04-delivery.md`](04-delivery.md) | Push to running: the three silos, levels, and what a green pipeline proves |
| [`05-runtime-and-installation.md`](05-runtime-and-installation.md) | The node runtime, its modes, and how a node comes into being |
| [`06-configuration-and-secrets.md`](06-configuration-and-secrets.md) | Managed files, value resolution, and where secrets live |
| [`07-knowledge.md`](07-knowledge.md) | The two knowledge stores, and what each is for |
| [`08-agents-and-work.md`](08-agents-and-work.md) | Agents as employees, tasks, workflows, and the meeting model |
| [`09-interfaces-and-observability.md`](09-interfaces-and-observability.md) | How the mesh is reached and watched — tools, board, proxy, health, thoughts |
| [`10-module-catalogue.md`](10-module-catalogue.md) | The catalogue's shape, and what its shape says |
## What these documents are not
They are not a runbook. Operational procedure — how to fix one occurrence of something — lives
in the knowledge base, which is indexed on symptoms and is the right place to search when
something is broken.
They are not exhaustive. A subsystem is described to the depth at which its **design** is
visible; below that is code.
@@ -1,6 +1,14 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md]
---
# Work breakdown — the decomposition
How ADR 0001 gets built, in what order, and where a human must look.
How [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) gets built, in what order, and where a human must look.
Ordering is not preference. Each phase removes a constraint the next one needs gone.
@@ -27,7 +35,7 @@ An agent may, without asking:
- **merging anything** — every merge is a human checkpoint, without exception
- **a decision the ADRs do not already answer** — record the question in the relevant
research effort rather than picking and moving on
- **any change to `hq/00-GENESIS`** — it is stable by nature
- **any change to `hq/00-META`** — it is stable by nature
### Definition of done for every task
@@ -74,7 +82,7 @@ The decomposition is impossible while a feature is a singleton per module.
| # | task | done when |
|---|---|---|
| 1.1 | ADR 0002 — named features, per-node opt-in | accepted |
| 1.1 | Decision record — named features, per-node opt-in (next free number) | accepted |
| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build |
| 1.3 | Selection: `always` / flavor-selected / `optional` | a node installs a subset; artifacts stay flavor-blind |
| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed |
@@ -1,3 +1,11 @@
---
layer: to-be
status: designed
code: [hal]
updated: 2026-08-23
decisions: [02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
---
# End-to-end testing
**What is under test is a module.** The mesh is the harness.
@@ -10,6 +18,18 @@ This is the Phase 0 prerequisite from [`00-work-breakdown.md`](00-work-breakdown
---
## Two boundaries this design sets
**Adoption is out of scope.** Bringing a node into being is the lab runner's job. Adopting a
machine that already exists, with its own configuration, is a legacy path and the lab does not
reproduce it — a scenario starts from nothing every time, which is what makes it a fixture
rather than a snapshot.
**The real topology is one test among many, not the baseline.** A scenario declares the mesh
its question needs. If something can only be tested against the shape the mesh happens to have
today, that is a gap in the vocabulary rather than a reason to privilege that shape.
## Where this sits in the way work happens
Work reaches the mesh along one path today:
@@ -33,7 +53,7 @@ Work reaches the mesh along one path today:
**The gate is a human reading a diff, and the test is production.** That is workable at a
change a day and it is the constraint at ten. For autonomous work it is worse than a
constraint: an agent's output arrives as a diff that *looks* right, carrying no evidence
that it runs, and the only reviewer is the condition `00-GENESIS/context.md` calls mandatory
that it runs, and the only reviewer is the condition `00-META/context.md` calls mandatory
— *human agents are few, often one, and usually asleep.*
The missing step goes between the pull request and the merge:
@@ -146,7 +166,7 @@ that only the developer path can do is a divergence, and it will drift.
the pipeline result is the verdict
```
This is the same property `00-GENESIS/mission.md` asks for: *the mesh's own components ship
This is the same property `00-META/mission.md` asks for: *the mesh's own components ship
through the same machinery as anything else it carries — if they need an exception, the
machinery is not finished.* A test that needed its own delivery path would be that exception.
@@ -350,6 +370,12 @@ reads credentials from host paths, because there is nowhere else to put a mesh.
is, a workstation stops being collateral.
**The supervision question stops gating anything.**
[`01-RESEARCH/003-service-supervision`](../01-RESEARCH/003-service-supervision/analysis.md)
[`01-RESEARCH/003-service-supervision`](../../01-RESEARCH/003-service-supervision/analysis.md)
remains open on its own merits — and once this exists, its options are cheap to try rather
than expensive to argue about.
## Deliberately not decided
**Whether the lab verdict is a workflow guard or an advisory check on the pull request.** Open,
and deliberately trivial — a policy detail, changeable in an afternoon, not an architectural
choice. Recorded so it is not mistaken for an oversight.
+22
View File
@@ -0,0 +1,22 @@
# 03-DESIGN / 01-to-be
The mesh being built toward. Every statement here traces to a record in
[`02-DECISIONS/`](../../02-DECISIONS/); nothing arrives by drafting.
A document here describes an intention. What currently runs is in
[`00-as-is/`](../00-as-is/), and the two are never merged — when something ships, the as-is
document is written and this one's status becomes `implemented`.
| Document | Covers | Rests on |
|---|---|---|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) |
## Not yet written
- **The eight bounded contexts.** [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)
decides the decomposition; the per-context specifications do not exist yet. The work
breakdown says in what order they are needed.
- **Domain grouping outside the core.** [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)
settles the principle and explicitly does not settle the domain list. That is a research
effort, not a design document, until it concludes.
+51
View File
@@ -0,0 +1,51 @@
# 03-DESIGN
The authoritative specification. Implementation is built against what is written here.
## Two layers
| Folder | What it is |
|---|---|
| [`00-as-is/`](00-as-is/) | **The mesh that exists today.** Shipped behaviour, described as it is — including behaviour nobody would choose again. |
| [`01-to-be/`](01-to-be/) | **The mesh being built toward.** Every statement traceable to a record in [`02-DECISIONS/`](../02-DECISIONS/). |
They are never mixed. A statement about the future does not belong in an as-is document, and
an as-is document is never edited to describe an intention.
When a to-be design ships, it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand — one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning, which is
the expensive half.
## Frontmatter
Every design document (not the READMEs) carries:
```yaml
---
layer: as-is | to-be
status: designed | in-progress | implemented | abandoned
code: [] # owning code repo(s), from 00-META/repos.md
updated: YYYY-MM-DD # date of the last status change, not of text edits
decisions: [] # 02-DECISIONS/ records this document rests on
---
```
For an as-is document, `status: implemented` is the normal state — it describes something that
runs — and `code:` names where that implementation lives.
Status changes when **implementation state** changes, never because design text was edited. An
`implemented` claim must be defensible from the owning repository's main branch, not from
intent. If it cannot be checked, it is `in-progress`.
Cross-cutting views are generated from this frontmatter by the `hal-status` skill and never
written to disk.
## What belongs here
Functional analysis, architectural description, and specification — **prose and diagrams
only, no code**. A manifest field may be named; a manifest may not be pasted. A document
enters the to-be layer only after the decision behind it is recorded in [`02-DECISIONS/`](../02-DECISIONS/)
and the research that produced it is closed.
Subfolders are encouraged where a layer grows enough to need them.
@@ -0,0 +1,49 @@
---
status: open
opened: 2026-08-22
located-in: []
fixed-by:
amended-design:
---
# 001 — A failed package install does not fail the job
## Symptom
A module declared a package. The install produced, from every mirror:
```
error: failed retrieving file … 404
```
followed by:
```
-> error installing repo packages
```
The prepare job then reported **success**. The package is absent; the pipeline is green.
## Why this matters more than one missing package
The first thing the lab work asked the mesh to install demonstrated the exact fault the lab
exists to catch — a step that failed, reported success, and left the next step to run against
state that was never produced.
It is also a direct violation of a decision already taken and recorded:
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) says a step that fails must fail the
job. That record notes the rule is applied instance by instance and enforced by no mechanism.
This is an instance where it was never applied.
## Evidence
- Observed 2026-08-22 while declaring the virtualisation package required by
[ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md).
- A fix is written and open as a pull request, unmerged since 2026-08-20.
## Open questions
- Why is the failure swallowed — is the exit status discarded, or never checked?
- Is this specific to package installation, or does the surrounding stage swallow every
non-zero exit?
- The fix has been open for two days. What is the review path for a change of this class?
@@ -0,0 +1,40 @@
---
status: open
opened: 2026-08-22
located-in: []
fixed-by:
amended-design:
---
# 002 — A declared package can fail purely because the node's index is stale
## Symptom
A package install ran without first synchronising the node's package index. It therefore
requested a version the mirrors had already superseded, and received a 404 from every one of
them.
The package exists. The declaration is correct. The node's view of what exists is old.
## Why this matters
The failure has nothing to do with the module, the manifest or the mirror. It is a property of
when the node last synchronised, which nothing in the mesh manages or reports. Two nodes given
the same declaration on the same day can produce different outcomes, and neither says why.
Combined with [issue 001](../001-failed-package-install-reports-success/00-report.md), the
failure is not only environmental but silent: today the node ends up without the package and
the job is green.
## Evidence
- Observed 2026-08-22, same declaration as issue 001.
- Documented as a recurring shape in the knowledge base under package installation failures,
where it is recorded as appearing in two disguises.
## Open questions
- Should the mesh own package-index freshness as a node property, the way it owns module
versions — or is an index sync part of the install step?
- A partial sync is unsafe on the platform in use; a full upgrade is the only sanctioned fix.
Does that make index freshness a scheduled node concern rather than a pipeline one?
@@ -0,0 +1,40 @@
---
status: open
opened: 2026-08-22
located-in: []
fixed-by:
amended-design:
---
# 003 — A firewall rule's `scope:` is read by no code
## Symptom
Five module manifests declare a `scope:` key on firewall rules. The key is not part of the
firewall rule type and nothing reads it. Real scoping is expressed by a different field.
A manifest can therefore appear to restrict a port and restrict nothing.
## Why this matters
This is the failure mode [`how-we-build.md`](../../00-META/how-we-build.md) names directly:
*an unenforced rule is indistinguishable from a wrong one, and costs more, because people
believe it.* Here it is worse than unenforced — the declaration reads as a restriction, so a
reviewer checking whether a port is scoped will find that it is, and be wrong.
It also says something about the manifest as a whole: an unknown key is accepted silently. Any
misspelled or invented key behaves this way, and this one was found by reading rather than by
any check.
## Evidence
- Five manifests carry the key. Zero code paths consume it.
- Recorded as an observation on 2026-08-22.
## Open questions
- Should the manifest reject unknown keys outright? That is the general fix; this is one
instance of it.
- Were the five declarations intended to restrict something that is currently open? Each needs
checking against what the node actually exposes — the declaration cannot be trusted either
way.
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-08-22
located-in: []
fixed-by:
amended-design:
---
# 004 — Certificate issuance always targets the authority's production endpoint
## Symptom
The reverse proxy sets no staging endpoint for its certificate resolver. Issuance therefore
goes to the public authority's production endpoint in every case, including experiments.
## Why this matters
Production issuance is rate-limited per domain and per account. Every certificate experiment on
a real node consumes quota that is not replenished quickly, and exhausting it is not
recoverable by retrying — it removes the ability to issue a certificate anyone actually needs.
The consequence lands hardest on exactly the work most likely to iterate: standing up a new
node, changing how names resolve, or testing the lab's certificate authority split
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
## Evidence
- The resolver configuration declares no staging endpoint.
- Observed 2026-08-22.
## Open questions
- Should the endpoint be a node property — production for nodes serving real traffic, staging
everywhere else — rather than a fixed proxy setting?
- The lab issues its own certificates and so does not consume public quota at all. Does that
make this a problem only for experiments run outside the lab, and therefore an argument for
running them inside it?
@@ -0,0 +1,47 @@
---
status: open
opened: 2026-08-22
located-in: [hal]
fixed-by:
amended-design:
---
# 005 — The end-to-end pipeline harness has not built since the workspace was removed
## Symptom
The repository's end-to-end pipeline test harness depends on a workspace that no longer
exists. It has not been buildable since 2026-06-04. Nothing runs it, and nothing reports that
nothing runs it.
## Why this matters
The delivery pipeline is the mesh's most consequential machinery — every module reaches every
node through it — and its only end-to-end coverage has been silently dead for over two and a
half months.
That interval is not incidental. Several of the pipeline's most expensive defects landed
inside it: packaging that was never actually split from build, migrations and provisioning
scripts reading a source layout that no longer ships, selection files never packaged at all.
Whether this harness would have caught any of them is unknown — which is itself the point. The
coverage was assumed, not checked.
## Evidence
- The workspace was removed by pull request #240 on 2026-06-04
([ADR 0007](../../02-DECISIONS/0007-no-npm-workspace.md)).
- The harness has not built since that date.
- Recorded in the knowledge base as a standing entry, not as a fixed incident.
## Relationship to the lab
[`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../../03-DESIGN/01-to-be/01-end-to-end-testing.md)
designs end-to-end testing on a lab mesh, which would replace this harness rather than repair
it. That is a reason to decide its fate deliberately, not a reason to leave it broken and
unmentioned: until the lab exists, this is the coverage the pipeline is presumed to have.
## Open questions
- Repair, or retire in favour of the lab? Leaving it in the repository unbuilt is the one
option that keeps the false impression of coverage.
- Was anything relying on it, or had it already stopped running before the workspace removal?
@@ -0,0 +1,57 @@
---
status: open
opened: 2026-08-23
located-in: [hal, hq]
fixed-by:
amended-design:
---
# 006 — This repository is not indexed into the knowledge base, and the claim that it is holds up a decision
## Symptom
[`README.md`](../../README.md) states, as the answer to the objection against creating this
repository:
> These documents are still indexed into the knowledge base, so `recall_search` returns them
> beside everything else. One source, many surfaces — which was always the actual requirement.
Searching the knowledge base for this repository's content returns nothing.
## Evidence
Verified 2026-08-23, two searches against the mesh's operational memory:
| Query | Result |
|---|---|
| The full title of a decision record in this repository | No results |
| A distinctive phrase from the decision ledger | No results |
No entry, no partial match, no stale copy. The indexing does not exist and appears never to
have existed.
## Why this is an issue and not a task
The claim is **load-bearing**. Decision 27 separates HQ into its own repository, and the
objection it answers was that a fourth knowledge system repeats the mistake the mesh's
knowledge consolidation was created to fix. The recorded answer is *"indexing, not location"* —
that the split is safe **because** these documents remain searchable alongside everything else.
Without the indexing, the objection stands unanswered and this repository is precisely the
fourth knowledge system it was argued not to be. Either the indexing is built, or decision 27's
reasoning is amended to something that is true.
It is also, exactly, the failure this repository names in its own rules: a document stating a
rule about the mesh must say how the rule is checked. This one stated a mechanism and nobody
checked it — including in the same commit that wrote the rule.
## Open questions
- Where would the indexing run? The operational memory is written through a mesh capability;
is this a periodic sync of a repository into it, or a search surface that reads the
repository directly?
- Which store — the flat symptom-indexed memory, the structured archive, or both? They have
different lifecycles ([`03-DESIGN/00-as-is/07-knowledge.md`](../../03-DESIGN/00-as-is/07-knowledge.md)),
and this content is governed rather than incidental.
- Public repository, private mesh: the sync direction must not become a path for mesh-specific
content to arrive **into** these documents.
+46
View File
@@ -0,0 +1,46 @@
# 04-ISSUES
The front door for "something is wrong" at the level of the mesh's design or governance.
Diagnosis happens here, where the whole mesh is in view; the fix lands in the owning code
repository.
## What belongs here
| Belongs here | Belongs in the knowledge base |
|---|---|
| The design permits a failure to be silent | How to fix one occurrence of it |
| A documented rule is enforced by nothing | A command that works around it |
| A stated invariant is false in practice | A node-specific quirk |
| The owner is unknown and finding it needs the whole mesh in view | Symptom → fix, once the answer is known |
The knowledge base already holds the operational record and is indexed on symptoms. **This
folder is not a second copy of it.** An issue here is a question HQ must *answer*; an entry
there is an incident someone must *clear*. An issue whose answer is a general lesson belongs in
both.
## Structure
```
NNN-short-name/
00-report.md the symptom as observed, with the evidence; status in frontmatter
01-diagnosis.md the investigation trail, dated, including what was ruled out
```
## Frontmatter, on `00-report.md`
```yaml
---
status: open | diagnosing | located | resolved | wontfix
opened: YYYY-MM-DD
located-in: [] # owning repo(s) or module(s), filled by diagnosis
fixed-by: # pull request or commit reference, filled at resolution
amended-design: # design doc path, when the root cause was a design gap
---
```
## Rules
- Anyone may open an issue. No localisation is required to report one.
- The full flow is playbook [`00-META/process/03-issues.md`](../00-META/process/03-issues.md).
- Closed issues are never deleted — they are the mesh's symptom-to-component memory.
- `wontfix` is legitimate and requires a sentence saying why.
+50
View File
@@ -0,0 +1,50 @@
# Agent instructions — Novox HQ
This repository is the source of truth for the HAL mesh's mission, research, design and
decisions. Implementation lives in the code repositories (see
[`00-META/repos.md`](00-META/repos.md)).
Before changing anything here, read the playbooks in
[`00-META/process/`](00-META/process/) — every workflow (research, graduation, design
amendment, issues, build handoff, constitution sync) is documented there, and agents operate
through them. Thin skills in `.claude/skills/` wrap these playbooks for invocation
(`hal-new-research`, `hal-graduate`, `hal-new-issue`, `hal-diagnose`, `hal-amend-design`,
`hal-handoff`, `hal-sync-constitution`, `hal-status`); each defers to its playbook as
authoritative and adds only the mechanical scaffolding.
## Ground rules
- **Markdown only.** No new top-level folders without explicit confirmation.
- **Status lives in YAML frontmatter** — on research overviews (`status`, `became`), design
docs (`layer`, `status`, `code`, `updated`), issue reports (`status`, `located-in`,
`fixed-by`, `amended-design`) and decision records (`status`, `date`, `deciders`).
Never create a central status file; cross-cutting views are generated from frontmatter.
- **`02-DECISIONS/` records are immutable.** Supersede with a new record; never edit meaning. Fixing a
broken link or path is allowed.
- **Design docs are prose and diagrams only** — no code. A manifest field may be named; a
manifest may not be pasted.
- **Two layers, never mixed.** [`03-DESIGN/00-as-is/`](03-DESIGN/00-as-is/) describes the mesh
that exists; [`03-DESIGN/01-to-be/`](03-DESIGN/01-to-be/) describes the one being built
toward. Every design doc says which it is in `layer:`. A statement about the future does not
belong in an as-is document, and an as-is document is never edited to describe an intention.
- **`00-META/how-we-build.md` is the source of the mesh constitution.** The knowledge-base
constitution page is derived from it — see playbook
[`05-constitution-sync.md`](00-META/process/05-constitution-sync.md). Never edit the
derived page directly.
## This repository is public
Nothing here may contain routable addresses, real domain names, hosting providers, node
names, absolute paths, usernames, credentials, or operational detail useful only to an
attacker. Use documentation ranges (RFC 5737, RFC 1918) and role names — `anchor`,
`home-server`, `workstation`, `laptop`, `the build node`, `the broker node`.
The test: would this paragraph still teach a stranger running an entirely different mesh?
If yes, it belongs. If it only makes sense to someone who knows this installation, it is
either a note in the wrong place or a disclosure. The full rule is in
[`README.md`](README.md).
## A rule states how it is checked
If a document states a rule about the mesh, it says how that rule is verified. An unenforced
rule is indistinguishable from a wrong one, and costs more, because people believe it.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
-91
View File
@@ -1,91 +0,0 @@
# Decision ledger
Every decision, in the order it was taken. Append-only — a decision that stops being true is
marked superseded and left in place, because the reasoning that was rejected is the expensive
half to rediscover.
An entry here is a **record**, not the reasoning. Anything architecturally significant carries
its full context, options and consequences in an [ADR](adr/); anything still being worked out
lives in [`01-RESEARCH`](01-RESEARCH/). This file is the index that makes both findable, and
the place small decisions live that never warrant a document of their own.
**Columns.** *Decided* is who made the call. *Where* points at the reasoning. A decision with
no pointer is one small enough that this line is the whole record.
---
## 2026-08-22 — decomposition
| # | Decision | Decided | Where |
|---|---|---|---|
| 1 | The mesh brokers capabilities; nodes host; agents think. Eight bounded contexts replace 33 platform modules. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) |
| 2 | Nodes and agents decouple — a node holds no licence; an agent holds credentials and delivery follows its bindings and modality. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) |
| 3 | noxflow dissolves; `hal/work` inherits tasks and workflows. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) |
| 4 | Third-party modules leave this repository — they run *on* the mesh, not *of* it. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) |
| 5 | ~~Documentation lives inside the code repository under `hq/`.~~ | jochen | **Superseded by #27** |
## 2026-08-22 — the lab
| # | Decision | Decided | Where |
|---|---|---|---|
| 6 | Phase 0 is a **development environment**, not a fixture rig — it is what the host-borrowing dev tooling becomes. | jochen | [002](01-RESEARCH/002-local-mesh/status.md) |
| 7 | The delivery trigger is a **real forge** inside the lab, not a synthetic event — the webhook relay is part of what is under test. | jochen | [002](01-RESEARCH/002-local-mesh/status.md) |
| 8 | ~~A lab node is a **system container**, promotable to a virtual machine.~~ | jochen | **Superseded by #10** |
| 9 | Adoption is a legacy path and is out of scope. Bringing nodes into being is the lab runner's job instead. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 10 | A lab node is a **virtual machine** running the real install. Supersedes #8: the scale argument for system containers was invented rather than required, and a virtual machine dissolves the fidelity question instead of answering it. | jochen | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md) |
| 11 | The environment is called **the lab**. | jochen | — |
| 12 | The lab is driven by `incus` — for virtual machines, snapshots and bridges through one interface, and because it also runs system containers if a scale run is ever needed. | jochen | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md) |
| 13 | The simulated public segment uses **TEST-NET-3** (`203.0.113.0/24`). Not cosmetic: an RFC1918 public segment makes the hub test as unreachable and the mesh silently never forms. | — | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md), [004](01-RESEARCH/004-lab-network/analysis.md) |
| 14 | The lab **issues its own certificates**, keeping production's two-authority split rather than collapsing it. | jochen | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md) |
## 2026-08-22 — what the lab is for
| # | Decision | Decided | Where |
|---|---|---|---|
| 15 | **The module is what is under test; the mesh is the harness.** The loop is: change a module, run it end to end, get a verdict. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 16 | **Nothing new drives delivery.** A lab mesh has its own coordinator; the pipeline that runs is the real one. A second delivery path would be blind to exactly the faults worth catching. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 17 | A **test runner** is a legitimate component *used by* the coordinator — lab lifecycle and assertion execution. The line is the pipeline: a runner that decides what to build is a fork of the coordinator. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 18 | The runner has **two callers** — the coordinator, and a person developing the mesh — so it needs both a run-to-verdict verb and a leave-it-standing verb. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 19 | **A module carries its own assertions**, stated once, in the verification stage the coordinator already dispatches. Running them where failing is free is what makes writing them worth doing. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 20 | A test declares the mesh it needs; **size ranges from one node upward**, chosen by the question rather than by what the mesh happens to have. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
| 21 | The real topology is **one test among many**, not the baseline. Anything only testable there is a gap in the vocabulary. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) |
## 2026-08-22 — working agreements
| # | Decision | Decided | Where |
|---|---|---|---|
| 22 | **Never install packages by hand.** A package is declared in the module manifest and arrives the way every other package does. | jochen | — |
| 23 | **Never open a pull request unprompted.** A permissions list saying it is allowed is not a request. | jochen | — |
| 24 | **Every merge is a human checkpoint**, without exception. | jochen | [`02-DESIGN/00-work-breakdown.md`](02-DESIGN/00-work-breakdown.md) |
| 25 | Work in an **isolated worktree**, never a shared checkout. | jochen | `CLAUDE.md` |
| 26 | Decisions are recorded **here**, thoroughly, as they are taken. | jochen | this file |
| 27 | HQ is **its own repository**, `hal-hq`. Supersedes #5: the original objection was to a fourth knowledge *system*, which indexing answers rather than location. Cadence, reviewers, and a scope wider than one repository all argue for separation. | jochen | [`README`](README.md) |
| 28 | **This repository is public.** Written for a reader who is not its author and has no access to the mesh it describes. No routable addresses, real domains, hosting providers, node names, absolute paths, or operational detail useful only to an attacker. Supersedes the previous rule that research may name instances. | jochen | [`README`](README.md) |
---
## Observations — not decisions, but they should not be lost
Things established by measurement that no decision has yet answered.
| Observed | What it means |
|---|---|
| **A failed package install does not fail the job.** Declaring `incus` produced `error: failed retrieving file … 404` from every mirror, then `-> error installing repo packages`, and the prepare job reported **success**. The package is absent; the pipeline is green. | The first thing the lab was asked to install demonstrated the exact fault the lab exists to catch. A fix is already written and open as PR #848, unmerged since 2026-08-20. |
| **The package database is stale on at least one node.** The install ran without a sync, so it requested a version the mirrors had already superseded — 404 from every mirror. | A declared package can fail purely because the node's index is old, and today that failure is silent. |
| **`scope:` on firewall rules is read by no code.** Declared in five manifests; not part of the rule type. Real scoping is `from:`. | A manifest can appear to restrict a port and restrict nothing. |
| **The reverse proxy sets no `caServer`**, so certificate issuance targets the public authority's production endpoint rather than staging. | Every certificate experiment on a real node consumes production issuance quota. |
| **`test/pipeline/` has been unbuildable since 2026-06-04**, when the npm workspace it depends on was removed. Nothing runs it. | The repository's only end-to-end pipeline test has been silently dead for two and a half months. |
## Deliberately not decided
Recorded so they are not mistaken for oversights.
| Question | Status |
|---|---|
| Which supervision model the mesh adopts — keep the host init system, drop the redundant per-module layer, containerise the daemons, or write a supervisor. | Open. Options costed in [003](01-RESEARCH/003-service-supervision/analysis.md). **No longer gates the lab.** |
| Whether the lab verdict is a workflow guard or an advisory check on the pull request. | Open, and deliberately trivial — a policy detail, changeable in an afternoon, not an architectural choice. |
| `hal/scheduler` — infrastructure or part of `hal/work`. | Open, from ADR 0001. |
| Which context owns the executor. | Open, from ADR 0001. |
| Catalogue destination — one repository or many. | Open, from ADR 0001. Phase 4. |
| What `hal/sdk` keeps after extraction. | Open, from ADR 0001. Phase 3. |
| Where human agent modality is recorded — which user, on which node, a human agent acts as. | Open, from ADR 0001. Required by the model; not yet stored. |
+51 -15
View File
@@ -7,17 +7,42 @@ why. Implementation lives in `modules/`; the reasoning behind it lives here.
| Folder | Purpose |
|--------|---------|
| [`00-GENESIS`](00-GENESIS/) | Mission and foundational context. The northern star for every decision. |
| [`00-META`](00-META/) | Mission, foundational context, the rules that hold across the mesh, the repository map, and the process playbooks. The northern star for every decision. |
| [`01-RESEARCH`](01-RESEARCH/) | Active and historical investigations, before they harden into design. |
| [`02-DESIGN`](02-DESIGN/) | The authoritative specification. Implementation is built against this. |
| [`adr`](adr/) | Numbered architecture decisions — what was chosen, and what was rejected. |
| [`02-DECISIONS`](02-DECISIONS/) | Numbered decision records, in the order the decisions were taken — what was chosen, and what was rejected. |
| [`03-DESIGN`](03-DESIGN/) | The authoritative specification, in two layers: [`00-as-is`](03-DESIGN/00-as-is/) — the mesh that exists — and [`01-to-be`](03-DESIGN/01-to-be/) — the one being built toward. |
| [`04-ISSUES`](04-ISSUES/) | The front door for "something is wrong" at the level of design or governance. |
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
That is why decisions are `02` and design is `03` — a reader following the numbers walks the
process in the order it happens.
## The flow
```
idea ──► 01-RESEARCH ──► decision (02-DECISIONS/) ──► 03-DESIGN/01-to-be ──► built (code repo)
│ │ │
│ │ └─► 03-DESIGN/00-as-is once shipped
│ └────► abandoned (recorded, kept)
└─(small/obvious, still recorded in 02-DECISIONS)──────────► 03-DESIGN directly
symptom ──► 04-ISSUES ──► diagnosis ──► code-repo fix and/or design amendment
00-META/how-we-build.md ──► sync ──► the constitution the mesh injects into design sessions
```
Implementation lives in the code repositories — see
[`00-META/repos.md`](00-META/repos.md). Every workflow is a playbook in
[`00-META/process/`](00-META/process/); agents operate through them and not outside them.
## Rules
- Markdown only.
- No new top-level folders without explicit confirmation.
- Knowledge flows `GENESIS → RESEARCH → DESIGN`. Research graduates into design only
after analysis against GENESIS confirms alignment.
- Knowledge flows `GENESIS → RESEARCH → DESIGN`. Research graduates into design only after
analysis against GENESIS confirms alignment, and only through a recorded decision.
- **Status lives in frontmatter**, never in a central status file. Cross-cutting views are
generated on demand, never hand-maintained.
- **GENESIS and DESIGN are instance-agnostic.** They describe the mesh as a concept — no
machine names, no counts, no topology. A reader must not be able to tell how many nodes
the author happened to have.
@@ -25,9 +50,12 @@ why. Implementation lives in `modules/`; the reasoning behind it lives here.
Evidence is what makes research worth reading, and the shape of a finding survives
anonymisation intact — *a node publicly named but behind a household NAT* carries the whole
lesson without naming anything.
- **The as-is layer records what is, not what should be** — including behaviour nobody would
choose again. A layer that keeps only the good decisions is a brochure.
- A document that states a rule about the mesh should say how that rule is **checked**.
This repo has a rule requiring every tools module to declare `brain` as a dependency;
zero modules do. An unenforced rule is indistinguishable from a wrong one.
This repository has a rule requiring every capability-exposing module to declare the core
runtime as a dependency; zero modules do. An unenforced rule is indistinguishable from a
wrong one.
### This repository is public
@@ -54,10 +82,17 @@ It began inside the code repository, on the reasoning that HAL already has a mes
knowledge store and that adding a fourth knowledge system would repeat the mistake this
folder was created to fix.
**That objection was about a fourth knowledge *system*, and it is answered by indexing, not
by location.** These documents are still indexed into the knowledge base, so
`recall_search` returns them beside everything else. One source, many surfaces — which was
always the actual requirement. Where the source is authored is a separate question.
**That objection was about a fourth knowledge *system*, and the answer offered was indexing
rather than location** — that these documents would be indexed into the knowledge base, so a
symptom search returns them beside everything else. One source, many surfaces. Where the source
is authored is then a separate question.
**That indexing does not exist.** It was checked on 2026-08-23 and returns nothing; it appears
never to have existed. Until it does, the objection stands unanswered and this repository is
the fourth knowledge system it was argued not to be. Recorded as
[`04-ISSUES/006`](04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md), and
left standing here rather than quietly reworded, because a claim that held up a decision and
was never checked is precisely the failure this repository exists to name.
Answered separately, a repository of its own is the better home:
@@ -65,13 +100,14 @@ Answered separately, a repository of its own is the better home:
changes. Tying documents to a code branch means they merge on the code's schedule.
- **The reviewers are different.** A design argument is not reviewed the way an
implementation is, and it should not queue behind a build.
- **The scope is wider than one repository.** ADR 0001 sends most modules out of the
- **The scope is wider than one repository.** [ADR 0015](02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the
monorepo entirely. Documentation that governs several repositories cannot live inside one
of them.
The trade is real and worth naming: a change to a document and the change to the code it
describes can no longer land in one commit. Keeping them honest is a discipline now rather
than a mechanism — which is why [`DECISIONS.md`](DECISIONS.md) records decisions as they are
taken, and why a document that states a rule should say how the rule is checked.
than a mechanism — which is why every decision is recorded in
[`02-DECISIONS`](02-DECISIONS/) as it is taken, and why a document that states a rule should
say how the rule is checked.
Recorded as decision 27 in [`DECISIONS.md`](DECISIONS.md).
Recorded as [ADR 0019](02-DECISIONS/0019-hq-is-its-own-repository.md).
-30
View File
@@ -1,30 +0,0 @@
# Architecture Decision Records
One file per decision, numbered, never deleted. A superseded ADR gets its status changed
and a pointer to what replaced it — the reasoning that was rejected is the expensive half
to rediscover.
## Format
```
# N. Title in plain language
- **Status:** Proposed | Accepted | Superseded by ADR-XXXX
- **Date:** YYYY-MM-DD
- **Deciders:**
## Context what is true today, with evidence
## Considered Options numbered, each with why it was rejected
## Decision what we are doing
## Consequences what follows, including what gets harder
## References code, data, prior art
```
State evidence, not assertion. "Zero of 124 modules declare `brain` as a dependency"
outranks "the dependency rule is not followed".
## Index
| ADR | Title | Status |
|-----|-------|--------|
| [0001](0001-mesh-brokers-nodes-host-agents-think.md) | The mesh brokers capabilities; nodes host; agents think | Accepted |