Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
48 changed files with 265 additions and 1257 deletions
Showing only changes of commit 5e83ac2c22 - Show all commits
+3 -3
View File
@@ -15,14 +15,14 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| Repository | Owns |
|---|---|
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0028](../02-DECISIONS/0028-hq-is-company-scoped.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
## What the mesh becomes
[ADR 0030](../02-DECISIONS/0030-the-repository-structure.md) records the repositories the
[ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the
monorepo decomposes into. **`mesh-lab` and `mesh-host` exist so far** — the lab is built first
([ADR 0029](../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)); the rest are the
([ADR 0016](../02-DECISIONS/0016-the-lab.md)); the rest are the
target, not the present.
| Repository | Tier | Holds |
+1 -1
View File
@@ -2,7 +2,7 @@
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/04-delivery.md, 03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-the-lab.md]
---
# 002 — A mesh that runs locally
+3 -3
View File
@@ -3,9 +3,9 @@ status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/01-mesh-and-transport.md, 03-DESIGN/01-to-be/01-end-to-end-testing.md]
became:
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
- 02-DECISIONS/0033-a-router-is-scenery-not-a-node.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 03-DESIGN/01-to-be/02-scenario-declaration.md
- 03-DESIGN/01-to-be/01-end-to-end-testing.md
---
@@ -68,7 +68,7 @@ against taste:
circle, self-hosted. Personal cloud infrastructure.
8. **Agents make it self-improving and self-healing.**
9. It is **end-to-end testable on one machine**
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Status
@@ -241,7 +241,7 @@ no human checkpoint.
## How this is tested
The lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) raises the
The lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) raises the
tree above on one machine: virtual machines as nodes, a real overlay between them, the real
substrate bundle, the real control plane, the real delivery path.
@@ -96,6 +96,6 @@ and it is not designed.
| What is a **deployed state**, and how does the mesh know it is in one? | Everything follows from this. If a stage reports transport, "deployed" is a claim nobody checked. A desired-state model with reconciliation gives a different answer from a job-completion model. |
| Does the coordinator dispatch **stages**, or converge nodes on a **declaration**? | The current model is a state machine over stages. The alternative is that a node is told what should be true and reports what is. The second makes drift visible; the first cannot see it. |
| How does a change **become** a pipeline, reliably? | Detection has failed for reasons unrelated to the change, silently. |
| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) and to a module carrying its own assertions. |
| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) and to a module carrying its own assertions. |
| How does delivery work **before self-hosting**, and across the transition? | From research 006: source and artifacts start external and are re-bound to internal providers. The coordinator has to be indifferent to which. |
| Does the **three-silo** split survive the artifact/part split? | [ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md) is cardinality-driven, and research 006 renames the thing the cardinality is about. |
+3 -3
View File
@@ -4,7 +4,7 @@ initiated: 2026-08-23
touches:
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
- 03-DESIGN/01-to-be/01-end-to-end-testing.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0016-the-lab.md
- 03-DESIGN/00-as-is/00-overview.md
became: []
---
@@ -28,7 +28,7 @@ Recorded because incremental is the reflex answer and it is wrong in this case.
requirements — none of these can half-apply. Running both models at once means the old one's
assumptions keep constraining the new one, which is how a migration becomes permanent.
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
same risk as one performed for the first time on the real mesh. This is also why the lab is
phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by
@@ -70,7 +70,7 @@ than a discovery.
| Phase | What | Done when |
|---|---|---|
| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| **0** | **Build the lab's bootstrap scenario** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
@@ -3,8 +3,8 @@ status: active
initiated: 2026-08-24
touches:
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
became: []
---
@@ -23,7 +23,7 @@ Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measuremen
## Why it matters
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
[ADR 0016](../../02-DECISIONS/0016-the-lab.md) makes the
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
@@ -13,7 +13,7 @@ image.
| Fact | Value | Consequence |
|---|---|---|
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-the-lab.md) is not paying an emulation penalty |
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
| btrfs kernel module | **available** | the kernel can do it |
@@ -72,7 +72,7 @@ worst, before any of the mesh's own work begins.
**This is too slow for an inner loop**, and the reason is not the design.
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
[ADR 0016](../../02-DECISIONS/0016-the-lab.md) argues that
making the bootstrap path the inner development loop turns the least-exercised code in the
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
@@ -120,7 +120,7 @@ A four-machine reset-and-rerun cycle, the operation the inner loop repeats most:
| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** |
At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds.
[ADR 0016](../../02-DECISIONS/0016-the-lab.md)'s argument holds.
At ninety it did not.
### One honest counter-observation
@@ -1,138 +0,0 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
Two things were measured while establishing what exists
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
- **There is no local mesh.** The dev tooling starts providers through the host's own init
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
It borrows the machine because there is nowhere else to put a mesh.
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
physical machine. Every subsequent question — how faithful is faithful enough, what may be
mocked, which failures remain reachable — follows from that one answer.
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
hardware virtualisation present.
## Considered Options
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
a node as one inverts the thing being modelled: module service stacks then require nested
containers through a privileged daemon, or a shared socket that makes isolation between
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
and the least like a node.
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
is the only option that makes a twenty-node run affordable. It was rejected because **the
scale requirement that justified it was invented rather than required**: the stated goal is
to run the real mesh, which is four nodes, on one computer. And a system container still
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
then has to be answered again for every capability under test.
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
no storage pools, no snapshot management, no network management, no virtual machines.
Snapshots are what make the loop fast, so the saving is not worth what it costs.
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
turns into a node.
## Decision
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
runs the real install, and becomes a node. It is not a model of a node, so no question arises
about how good the model is.
The environment is called **the lab**.
Three things follow directly and are decided here:
### The lab is driven by `incus`
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
bridges between them, through one interface. It also manages system containers, so if a run
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
does.
### The simulated public segment uses TEST-NET-3
`203.0.113.0/24`, reserved by RFC 5737, never routable.
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
peer's underlay address against an RFC1918 regex
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
rather than an addressing mistake.
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
### The lab issues its own certificates
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
single-authority lab would hide any fault living in that split.
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure rather than a mystery.
## Consequences
**The fidelity question disappears, and with it a class of argument.** There is no "how real
is this node" to litigate per capability, because the node is real. What remains not-real is a
short, enumerable list: the model provider, the public internet, and the public certificate
authority.
**The install becomes the thing under test.** A container-shaped lab would have had to skip
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
remains possible at lower fidelity if one is ever genuinely needed.
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
time is noise.
**One change is required before the lab can issue certificates.** The reverse proxy sets no
`caServer`, so it defaults to the public authority's *production* endpoint
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
than staging means every certificate experiment on a real node consumes issuance quota.
## References
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
- RFC 5737 — reserved documentation address blocks
+84
View File
@@ -0,0 +1,84 @@
---
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
consolidates: [0029, 0031, 0032, 0033]
---
# 16. The lab
*Consolidated 2026-08-28 from five records. The lab is one design and was split across five
decisions taken over three days; the reasoning is kept, the fragmentation is not.*
The environment a change is run against before it reaches real machines.
## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
a node**, so no question arises about how good the model is — which is the whole reason for
paying the cost of virtual machines rather than containers.
The lab is driven by **incus**, and a scenario is raised from a declaration.
## A router is scenery, and is therefore a container
**Nothing under test runs on a router.** It is not a participant, holds no identity, has nothing
installed on it by the mesh, and no assertion is ever made about its internals. It exists so that
packets between machines behave the way they behave in the world.
The fidelity argument that makes a node a virtual machine does not reach it: what a router *is*
does not matter, only what it *does to traffic*. So a router is a system container, and the lab
is cheaper for it.
## A scenario declares the underlay, and only the underlay
**What a hosting provider and a home router would have provided**, before any of our software
touched the machine:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and may be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering, no
names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
> A scenario provides what a hosting provider and a home router would provide, and nothing our
> software is responsible for.
## A scenario is a closed address space
Every segment materialises as its own isolated link belonging to one scenario instance. **Two
scenarios raised from the same declaration hold the same addresses and never meet**, because
nothing joins their links. The declaration therefore keeps its literal addresses and they mean
exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the way one would
use a console rather than the network. That is what makes two identical scenarios able to run at
once, and it is why placing anything inside a machine is a hypervisor operation rather than a
network one.
## Two scenario classes, and the first has no pipeline
| | **bootstrap** | **full** |
|---|---|---|
| contains | machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| verdict from | what the host reports about the state it reconciled | a delivery result ending in verification |
| exercises | tiers 0 and 1 | tiers 2 and above, and modules |
**The bootstrap class comes first**, because it is what develops the node host, and because a
full scenario needs tiers that do not exist yet. A lab that could only raise the larger class
would be a lab nobody could use until everything else was built.
## Consequences
- **The lab tests the real code path**, not a reimplementation of it. The network a scenario
produces is generated by the same code production runs.
- **Isolation is what makes it usable in parallel**, and it costs the ability to reach in over
IP. Everything the lab puts inside a machine — a binary, an image, a file — goes through the
hypervisor.
- **A sealed scenario cannot fetch anything**, which is a real limit rather than an inconvenience:
it is why images have to be placed and why a container runtime has to be in the base image.
@@ -0,0 +1,112 @@
---
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
consolidates: [0020, 0021, 0022, 0023, 0024, 0026, 0027, 0028, 0030]
---
# 19. How this repository works
*Consolidated 2026-08-28 from ten records that were one decision seen from ten angles. The
reasoning is kept; the fragmentation is not.*
## The repository
**`novox/hq` is Novox's headquarters, and it is public.**
Company-scoped, not the mesh's. Today almost everything in it is about the mesh, because the
mesh is what Novox is building — a fact about the present rather than a definition. A second
product would live here too.
**Public** means written for a reader who is not its author and has no access to the mesh it
describes. Nothing here may contain routable addresses, real domain names, node names, absolute
paths, usernames or credentials. The test: *would this paragraph still teach a stranger running
an entirely different mesh?*
**Separate from the code** because the cadence differs — a decision changes when thinking
changes, not when code changes — and because a public repository cannot be a private one's
subdirectory.
**The naming rule:** a repository belonging to a product carries that product's prefix. A
company-scoped one does not. So this is `hq` and the mesh's are `mesh-*`.
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the pinned tier-1 services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | the mesh's own domain ([ADR 0065](0065-the-core-library-is-the-meshs-domain.md)) |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this one |
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths. **`Nox` is an identity of Novox**, an agent
participant within the mesh's own model, not a second system.
## The folders, and why they are numbered
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
Following the numbers walks the process in the order it happens.
| | |
|---|---|
| `01-RESEARCH` | an open question, while it is open |
| `02-DECISIONS` | what was decided, and why |
| `03-DESIGN` | what is being built |
| `04-ISSUES` | something wrong at the level of design or governance |
**`03-DESIGN` has two layers and they are never mixed.** `00-as-is/` describes the mesh that
exists, written from the implementation and the operational record. `01-to-be/` describes the one
being built toward. Every document says which it is. A statement about the future does not belong
in an as-is document, and an as-is document is never edited to describe an intention.
**`04-ISSUES` is for design-level faults** — a rule enforced by nothing, a stated invariant that
is false, a failure the design permits to be silent. Not an operational ticket queue.
## What a decision record is, and is not
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is not
recorded.**
That bar has been read too generously. A *finding* is not a decision. A bug is not a decision.
**A record is warranted when there is a genuine fork**: a direction reversed, an alternative
seriously considered and likely to be proposed again, or something contested that needs to stay
settled. Everything else belongs in the design document, where the reasoning is read.
**There is no ledger** — no separate document summarising, ranking or tracking decisions. A
chronological view is generated from frontmatter, which is what a ledger was actually for.
**The design layer is what you read.** These records explain *why* a thing is as it is. They are
not a description of the system, and needing to read them to understand it would mean the design
documents had failed.
## Status, and views over it
**Every document carries its state in YAML frontmatter** — research overviews, design documents,
decision records, issue reports.
**There are no central status files.** Every cross-cutting view — a status matrix, a decision
index, an open-issue list — is generated from frontmatter when asked for, never written to disk.
Two places holding one fact drift, and the written one wins by being closer to hand.
**Prose does not restate status.** One place, and two is one too many.
## Workflows are playbooks
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs it,
steps, outputs. People and agents follow the same ones, and **agents do not act outside them**.
Each is wrapped by a thin skill that defers to the playbook as authoritative and adds only the
mechanical scaffolding — so the process has one definition rather than a document and an
implementation that disagree.
## Consequences
- **A reader has one place per thing.** The design layer describes the system; these records say
why; the playbooks say how work is done.
- **Records will accumulate more slowly**, because the bar is a fork rather than a finding. This
record is itself the correction: ten records became one because they were one decision.
- **The public rule constrains everything written here**, permanently and at every commit. It is
the reason research describes real observations without identifying the mesh it observed.
@@ -1,67 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 20. Design is written in two layers: what is, and what is intended
## Context
HQ held only the intended mesh. Every reader had to already know the running system in order
to understand what the decisions were about, and a statement about current behaviour had
nowhere to live except inside a document describing an intention.
The consequence was invisible until looked for: an as-is claim inside a to-be document is
indistinguishable from the intention around it, so the document silently stops being true as
the system moves — and nobody can tell which half went stale.
The mesh also has a large body of shipped behaviour that nobody would choose again. It is not
design in the sense of "what we decided"; it is design in the sense of "what is there", and it
is exactly the part a person changing the system most needs.
## Considered options
1. **One layer, describing the target.** Rejected — the status quo. The running system goes
undocumented and the target document accumulates unmarked claims about it.
2. **One layer, describing what runs, with intentions only in decision records.** Rejected:
a decision record is an argument, not a specification, and a multi-part intention has
nowhere coherent to live.
3. **Two layers, declared per document, never mixed.** Chosen.
## Decision
`03-DESIGN` holds two layers, and every document declares which it is:
| Layer | Describes | Written from |
|---|---|---|
| `00-as-is/` | The mesh that exists | The implementation and the operational record |
| `01-to-be/` | The mesh being built toward | Decision records |
An as-is document **records what is, not what should be** — including behaviour nobody would
choose again. A layer that keeps only the good decisions is a brochure.
When a to-be design ships it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand: one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning.
Where implementation and intention disagree, the as-is document records the implementation and
says they disagree.
## Consequences
- A reader can tell, from the folder and from one frontmatter field, whether they are reading
a description or a plan. That distinction was previously unavailable at any price.
- Correcting an as-is document requires evidence from the implementation, not agreement — and
needs no decision record, because nothing was decided.
- Two documents must be kept current per subsystem instead of one. This is the cost, and it is
paid on every ship.
- Something that shipped differently from its design becomes a visible divergence rather than
a silently wrong document, and may deserve an issue.
## References
- [`03-DESIGN/README.md`](../03-DESIGN/README.md) — the layer contract and frontmatter schema.
- [`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/) — the first eleven as-is documents, written
2026-08-23 from the monorepo and the operational memory.
@@ -1,57 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 21. Every workflow is a playbook, and agents operate through them
## Context
HQ stated a knowledge flow — research becomes design — and nowhere stated how anything moves
along it. What graduation required, who wrote the decision, what closed an effort, what
happened when something shipped: all of it was convention held in one person's head.
A large share of the work here is done by agents. An unwritten convention is not available to
an agent at all, so each one either invents a procedure or asks. Both produce a repository
whose shape depends on who last touched it.
## Considered options
1. **Convention, learned by reading existing documents.** Rejected — the status quo. It
transmits shape but not rules, and it transmits the mistakes along with the pattern.
2. **One long contributing document.** Rejected: it is read once, and the step someone needs
is never the step they are reading.
3. **A playbook per workflow, each with trigger, steps and outputs, wrapped by a thin skill.**
Chosen.
## Decision
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs
it, steps, outputs. Engineers and agents follow the same playbooks, and **agents must not act
outside them**.
Each playbook is wrapped by a thin skill that defers to it as authoritative and adds only
mechanical scaffolding — next free number, frontmatter block, where the file goes. The
playbook holds the reasoning; the skill holds the steps. When they disagree, the playbook
wins.
## Consequences
- An agent arriving with no context can act correctly, because the procedure is retrievable
rather than remembered.
- The playbooks are themselves reviewable. A bad rule can be found and changed, which is not
true of a convention.
- Duplication between playbook and skill is real, and is managed by making the skill thin and
naming the playbook as authoritative in the skill's first lines. Nothing prevents them
drifting; the constraint is that only one carries reasoning.
- A workflow with no playbook is a workflow agents will get wrong. Adding one is part of
adding the workflow.
## References
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the five playbooks
and the flow they implement.
- Modelled on the process layer in the sibling HQ repository for the PAPA platform, which
arrived at the same shape and the same thin-skill split.
@@ -1,54 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 22. Status lives in frontmatter; cross-cutting views are generated
## Context
Status was carried in prose — a bold line near the top of a document saying what state it was
in — and indexes were maintained by hand. The decision-record index had already drifted from
the folder it described **after a single addition**, which is about as short a demonstration
as the failure mode offers.
A hand-maintained index is a copy of something the filesystem already knows. It is correct
only for as long as everyone remembers it exists, and its being wrong is silent.
## Considered options
1. **Prose status plus hand-maintained indexes.** Rejected — the status quo, already
demonstrably broken.
2. **A central status file.** Rejected. It centralises the drift rather than removing it: the
file and the documents disagree, and the file is the one people read.
3. **Machine-readable frontmatter per document; every cross-cutting view generated on
demand.** Chosen.
## Decision
Every document carries its state in YAML frontmatter — research overviews, design documents,
decision records, issue reports — with a schema stated in the section README.
**There are no central status files.** Every cross-cutting view — a status matrix, the
decision-record index, the open-issue list — is generated from frontmatter when asked for, and
never written to disk.
Prose does not restate status. One place, and two is one too many.
## Consequences
- A view cannot drift from what it describes, because it does not persist.
- Status becomes queryable. Inconsistencies — a closed effort with nothing in `became:`, an
`implemented` design with no owning repository — are findable mechanically, and the
generator reports them as flags rather than silently rendering around them.
- Frontmatter must be valid and paths in it must resolve, which is now something to check.
- A reader browsing the repository on a forge sees no index. That is the trade: the index is
correct and absent rather than present and wrong.
## References
- [`.claude/skills/hq-status/SKILL.md`](../.claude/skills/hq-status/SKILL.md) — the
generator, including the inconsistencies it flags.
- [`02-DECISIONS/README.md`](README.md) — the hand-written index that drifted, and its removal.
@@ -1,60 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 23. Issues have a front door, separate from the operational memory
## Context
Findings that were nobody's task accumulated in a table inside the decision ledger — a package
install reporting success while installing nothing, a manifest key read by no code, an
end-to-end harness dead for months. They were measured, true, and unowned: a table row cannot
be assigned, diagnosed or closed.
The mesh already has an operational memory holding roughly a hundred and thirty entries,
indexed on symptoms. The obvious move — put these there — is wrong, and the reason is the
distinction worth recording.
## Considered options
1. **Leave them in the ledger.** Rejected: a ledger records decisions taken, and these are
the opposite — questions nobody has answered.
2. **Put them in the operational memory.** Rejected. That store answers *how do I fix this
occurrence*; these are *why does the design allow this at all*. Filing them there makes
them findable by symptom and unfindable as open questions, and nothing there has a state
that can be closed.
3. **A numbered issue folder in HQ, deliberately narrow.** Chosen.
## Decision
`04-ISSUES` is the front door for something wrong at the level of **design or governance**:
a rule enforced by nothing, a stated invariant that is false, a failure the design permits to
be silent, or a symptom whose owner cannot be found without the whole mesh in view.
One numbered folder per issue: the report with the symptom as observed and the evidence, and
a diagnosis document carrying the trail, dated, including what was ruled out.
**This is not a second copy of the operational memory.** An issue here is a question HQ must
*answer*; an entry there is an incident someone must *clear*. An issue whose answer is a
general lesson belongs in both — and the operational memory is searched first, because if the
answer is already there this was never an issue.
## Consequences
- A finding gets a number, a state and an owner, and closing it is a visible act.
- The symptom-to-component trail accumulates in a place where the whole mesh is visible, which
is where cross-component diagnosis has to happen.
- The boundary needs judgement on every report, and will sometimes be got wrong. Filing too
narrowly loses a finding; filing too widely rebuilds the operational memory here, which is
the outcome HQ's separation was argued against
([ADR 0019](0019-hq-is-its-own-repository.md)).
- Six issues opened on creation, all previously unowned observations.
## References
- [`04-ISSUES/README.md`](../04-ISSUES/README.md) — the boundary table and the frontmatter
schema.
- [`00-META/process/03-issues.md`](../00-META/process/03-issues.md) — the playbook.
@@ -1,76 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 24. The folder numbering is the flow, and decision records run oldest first
## Context
Two orderings were wrong in ways that only show up when someone new reads the repository.
**The folders.** Decisions lived in an unnumbered folder that sorted after the numbered ones,
so the repository's most load-bearing content read as an annex.
The sibling HQ repository for the PAPA platform had already solved this and solved it
crookedly: its design folder existed from its initial commit, and when its decision folder was
finally promoted it took the **next free number** rather than its place in the sequence. That
repository now reads `01 research → 03 decision → 02 design`. A decision precedes the design
it authorises and is numbered after it. By the time this was visible, the design folder was too
settled to renumber.
**The records.** Fourteen decisions had been taken in implementation and never written down —
the broker, the module abstraction, the mesh database, the artifact, the three silos and the
rest. Meanwhile two records existed, holding numbers 0001 and 0002, for decisions taken last.
## Considered options
1. **Match the sibling repository exactly**, inheriting its ordering. Rejected: structural
parity is worth something, but not the cost of copying a scar the other repository would
not choose again.
2. **Leave the folder unnumbered.** Rejected — the annex problem, and it leaves an unexplained
gap for anyone arriving from the sibling repository.
3. **Number by position in the flow, and renumber the records chronologically.** Chosen, on
the grounds that this repository was four commits old and nothing outside it cited a
number. That is the only window in which either renumbering is free.
## Decision
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
So `01-RESEARCH`, `02-DECISIONS`, `03-DESIGN`, `04-ISSUES`. Following the folder numbers walks
the process in the order it happens.
**Decision records are a chronological ledger.** They run oldest first. The fourteen decisions
already taken in implementation were back-filled as records 0001–0014, each dated from the
history, each carrying `reconstructed: true` and saying so in its first lines, and each citing
the commit, pull request or knowledge-base entry it was recovered from. The two existing
records moved to 0015 and 0016.
A reconstructed record is not a transcript. Where the deliberation is not recoverable it states
what the alternatives were and why the chosen one won on the evidence available — not a
discussion that did not happen. Where a date is not establishable it says so.
The foundational folder is `00-META`, matching the sibling repository.
## Consequences
- The repository reads in process order, and the gap at `03` that a reader coming from the
sibling repository would notice is explained by this record.
- The design documents can cite reasoning instead of asserting rules, because the reasoning now
exists.
- Structural divergence from the sibling repository, deliberately, in exactly one place. It is
recorded here so that the difference reads as a choice rather than an accident.
- **Record numbers are now stable and renumbering is over.** This decision spends the one
window that existed; a future record takes the next free number regardless of its date.
- Reconstructed records carry a standing risk: they are the most confident-sounding documents
in the repository and the least directly witnessed. The `reconstructed` flag exists so that
is never invisible.
## References
- The sibling repository's restructure of 2026-07-13 moved its decision folder in a single
commit of twelve renames with no content change, alongside the same status-into-frontmatter
and playbook changes made here.
- [`02-DECISIONS/README.md`](README.md) — the format, and the note on reconstructed records.
@@ -1,96 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 26. Every decision is a record; there is no ledger
## Context
HQ carried a decision ledger at its root: a chronological table of forty-one numbered
decisions, each with who decided and a pointer to where the reasoning lived. It was created
deliberately, to make decisions findable and to give a home to decisions too small to warrant
a document.
By the time the decision records were back-filled
([ADR 0024](0024-the-numbering-is-the-flow.md)) the ledger had become three things at once,
and only one of them was still needed.
Classified, its forty-one entries were: ten restating a record, eleven restating design
documents, fifteen describing how this repository works — with the reasoning in a README rather
than anywhere citable — three small rules with no home at all, and two superseded stubs.
So the ledger was mostly a copy. Worse, it was a **hand-maintained index**, which
[ADR 0022](0022-status-lives-in-frontmatter.md) had just finished rejecting for the
decision-record index on the grounds that it had drifted after a single addition. Keeping one
copy of that pattern while removing another is not a position.
It had also produced a naming collision that a directory listing makes plain: `DECISIONS.md`
beside `02-DECISIONS/`, holding different things.
## Considered options
1. **Keep the ledger.** Rejected. It duplicates the records, restates status, and is the exact
hand-maintained index this repository decided against elsewhere.
2. **Keep it, renamed, for small decisions only.** Rejected, and this is the option worth
arguing with — it is genuinely useful to record a decision without writing a document. But a
decision small enough to be one table row is almost always a **rule** rather than a
decision, and a rule belongs in [`how-we-build.md`](../00-META/how-we-build.md) where it is
enforced and where its reasoning is kept. That is where the three orphans went.
3. **Every decision is a record; nothing else.** Chosen. This is how the sibling HQ repository
for the PAPA platform works, and it has no ledger of any kind.
## Decision
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is
not recorded.**
`02-DECISIONS` holds every decision, and **there is no ledger** — no separate document in which
a decision is also summarised, ranked or tracked. The chronological view — decisions in the order
they were taken — is *generated* from record frontmatter, which is what the ledger was actually
for.
That generation is not this record's rule. It is
[ADR 0022](0022-status-lives-in-frontmatter.md), which already decides repository-wide that
status lives in frontmatter and every cross-cutting view is generated rather than written. This
record does not restate it — 0022's own words are *prose does not restate status; one place, and
two is one too many*, and an earlier version of this paragraph did exactly that.
Content that was only in the ledger was rehomed rather than dropped:
| Was | Went to |
|---|---|
| Decisions about how this repository works | Records [0019](0019-hq-is-its-own-repository.md)–[0025](0025-hq-is-the-source-of-the-constitution.md) |
| Small rules with no record | [`how-we-build.md`](../00-META/how-we-build.md) — the package rule, and two already there |
| Lab decisions not stated in the design | [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) |
| "Deliberately not decided" | The research effort and design document each question belongs to |
| Unowned observations | [`04-ISSUES`](../04-ISSUES/) ([ADR 0023](0023-issues-have-a-front-door.md)) |
## Consequences
- One place to look, and nothing to keep in sync. The collision between the ledger and the
record folder is gone.
- Structural parity with the sibling repository on decisions, which
[ADR 0024](0024-the-numbering-is-the-flow.md) deliberately broke on folder numbering. The
divergence is now exactly one thing, and it is the one thing that was argued for.
- **Writing a record is now the only way to record a decision, and a record is more work than
a table row.** The real risk is that a small decision goes unrecorded because nobody wanted
to write a document. The mitigation is that a small decision is usually a rule, and
`how-we-build.md` takes rules cheaply — but this is a cost, not a solved problem, and it is
the thing to watch.
- The chronological view now depends on the generator existing and being run. It did not
before.
- Two superseded ledger stubs had no record of their own. The position that documentation
lives inside the code repository is now recorded only as superseded context in
[ADR 0019](0019-hq-is-its-own-repository.md); the system-container position is explained in
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md). Neither is lost.
## References
- The sibling PAPA HQ repository: root holds only agent instructions and a README; every
decision is a numbered record, and its graduation playbook has no path for an unrecorded
decision.
- [ADR 0022](0022-status-lives-in-frontmatter.md) — the hand-maintained-index argument this
applies consistently.
@@ -1,112 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 27. The product is Novox Mesh; Nox is an identity, not a second system
## Context
The name `HAL` was never chosen. This began as a dotfiles repository, the first commits in
February 2026 adopt dotfiles and per-node overrides, and the name arrived with the code — as
recorded in
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md),
most of the current shape is inherited from that origin rather than designed for a mesh. The
name is part of the inheritance.
Three things make it worth changing rather than living with.
**It is borrowed, and borrowed badly.** HAL is the canonical *untrustworthy* machine
intelligence. For infrastructure whose entire proposition is that it manages your machines,
heals itself, and is trusted with credentials, that is an unhelpful flag to fly, and it is not
a name anyone owns.
**There is a name available that is owned.** The company is Novox. A product of Novox should
carry that lineage rather than a film reference.
**A platform and a persona are different things, and one name was doing both.** `HAL` named the
mesh *and*, implicitly, the thing an operator talks to. Those are separate concerns — the
platform is what runs; the persona is who answers.
## Considered options
1. **Keep `HAL`.** Rejected. Every reason to keep it is sunk cost, and the sunk cost is at its
smallest today.
2. **Rename everything to a single new name covering platform and persona.** Rejected: it
repeats the conflation that made `HAL` ambiguous.
3. **Separate the two: a product name and an identity.** Chosen.
## Decision
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths.
**`Nox` is an identity of Novox**, and specifically an **agent identity within the mesh's own
model** — a named participant, exactly as
[ADR 0012](0012-agents-are-persistent-employees.md) defines one. Not a separate product, not a
separate runtime, not a privileged path.
**Nox is the agent of the mesh, not of a node.** This is the part that carries weight:
- **Every node keeps its own identity.** That already exists and stays — a node is a named
participant with its own character, and addressing one directly remains possible and normal.
- **Nox is scoped to the whole mesh.** It is what the mesh is called when the mesh itself
speaks, rather than one machine within it.
- **Nox addresses node identities.** Asking Nox for something that lives on one node is Nox
talking to that node, not a human choosing a machine.
- **A human mostly talks to Nox.** It is the front door.
That last point makes Nox the concrete form of the vision in
[`00-META/mission.md`](../00-META/mission.md): *an agent states an intent and the mesh carries
it out — no console to open, no runbook to follow, no remembering which node holds which
thing.* Nox is who that intent is stated to. The mission described the behaviour; this names
the thing that has it.
Nox holds no private channel. Whatever it can do, it does through the same surfaces every other
agent uses — which is not a naming detail: a persona with its own path would be the one part of
the mesh with no human checkpoint, and the skeleton already rules that out.
`HAL` is retired.
**Timing is the substance of this decision, not an aside.** The skeleton in
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) is not built. Renaming
before it exists costs a search and replace across research documents. Renaming after costs the
same class of migration as everything else this repository is trying to avoid, and would
therefore not happen.
## Consequences
- **The as-is layer keeps `HAL`.** It describes what runs, and what runs is called HAL. The
to-be layer uses `mesh`. The rename is part of the migration, and the two-layer split
([ADR 0020](0020-design-is-written-in-two-layers.md)) is what makes holding both names
coherent rather than confusing.
- **Records 0001–0026 keep `HAL`.** They are immutable and they say what was decided when it
was decided. No record is edited for a name.
- Tier 2 cannot be `mesh-mesh`. The control plane is **`mesh-control`**; `mesh-broker` was
rejected because the substrate already contains a message broker.
- The namespace, environment variable prefix, service names and on-disk paths all change. In
the existing system that is a migration and is not attempted here.
- **`mesh` is a generic word**, and it already means something specific in infrastructure — a
service mesh is a different thing. Recorded as a known trade rather than an oversight: the
full name `Novox Mesh` is distinctive, and the short form is internal.
- The persona has a name before it has behaviour. That is the right order — it is an identity in
a system that already has a model of identities, so it needs no new machinery to exist.
- **Except in one respect, and it is a real gap.**
[ADR 0012](0012-agents-are-persistent-employees.md) binds every agent to a home node, one to
one, with a workspace on that machine. A mesh-scoped agent has no home node by definition, so
the model does not currently have a shape for Nox. Extending it — an agent whose scope is the
mesh rather than a machine — is a decision of its own and is not taken here.
- Two levels of identity now exist where there was one: the node, and the mesh. The distinction
has to stay visible in every surface, or "ask Nox" and "ask a node" collapse into each other
and it stops being clear who is answering.
## References
- [ADR 0012](0012-agents-are-persistent-employees.md) — what an identity is in this system, and
why `Nox` needs no separate mechanism.
- [ADR 0020](0020-design-is-written-in-two-layers.md) — why the as-is and to-be layers can
legitimately use different names for the same system.
- The dotfiles origin, and the naming inheritance it explains:
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md).
-89
View File
@@ -1,89 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 28. HQ is company-scoped; the mesh is its first product
## Context
This repository was `hal-hq` — one product's headquarters, named for the product. Then the
product was renamed ([ADR 0027](0027-the-product-is-novox-mesh.md)), which forced the question
of what the repository is actually the headquarters *of*.
Two facts settled it, and both were checked rather than assumed.
**Novox already delivers other things.** The company's forge organisation holds live projects
beside the mesh, and they are registered as build sources — meaning the mesh already builds and
deploys them. They are not hypothetical future products; they exist and ship today.
**They are tenants, not peers.** They run *on* the mesh. Every one of them is developed,
delivered and hosted by it. So the mesh is not one product among several — it is the ground the
others stand on.
That distinction decides the scope. If the mesh were a product beside others, a per-product HQ
would be right. Because it is the substrate the company operates on, a decision about the mesh
is a decision about how the company works.
## Considered options
1. **`mesh-hq` — one HQ per product.** The safe choice, and the reversible one: a second
product creates its own HQ and shared practice graduates upward later. Rejected, knowingly,
because it models the mesh as a peer of things that are actually its tenants.
2. **A company HQ *and* a product HQ, from the start.** Rejected as ceremony — two repositories
for one operator, and the constitution's own YAGNI rule says not to.
3. **One company-scoped HQ, `novox/hq`, with the mesh as its first product.** Chosen.
## Decision
The repository is **`novox/hq`** — Novox's headquarters, not the mesh's.
It holds the reasoning behind what Novox builds. Today almost all of that is the mesh, because
the mesh is what Novox is building. That is a fact about the present, not a definition of the
repository.
**The scope of each document is fixed now, so the eventual split is mechanical rather than
archaeological:**
| Scope | Documents | Moves if products separate? |
|---|---|---|
| **Company** | [`how-we-build.md`](../00-META/how-we-build.md), [`process/`](../00-META/process/), [`repos.md`](../00-META/repos.md), this record and [0019](0019-hq-is-its-own-repository.md)–[0027](0027-the-product-is-novox-mesh.md) | No — they stay at the top |
| **Product (mesh)** | [`mission.md`](../00-META/mission.md), [`context.md`](../00-META/context.md), [`effect.md`](../00-META/effect.md), `01-RESEARCH`, `03-DESIGN`, `04-ISSUES`, records 0001–0018 | Yes — into a product section |
The folders are **not** restructured now. One product's content under a company name is
correct while there is one product's worth of it, and nesting before there is anything to nest
is the ceremony option 2 was rejected for.
## Consequences
- Engineering practice has a home that does not belong to the mesh. `how-we-build.md` — never
write to production directly, migrations for schema changes, runtime evidence for behavioural
criteria — is true of any Novox project, and its being in a mesh repository was always a
slight mislabelling.
- The constitution derived from it ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md))
can legitimately govern work outside the mesh. Under a product HQ it could not have, without
either duplicating or reaching across repositories.
- **The bet is not entirely forward-looking, and that is worth being honest about.** Novox
already has work that is *not* a mesh tenant — client engagements and at least one product
that is developed outside it. So the company genuinely has a scope wider than the mesh
**today**, which strengthens the case for a company HQ and simultaneously means the split in
the table above is closer than "some day". The table is not a precaution; it is a plan whose
trigger already half-exists.
- What has *not* happened yet is any of that work needing the constitution. That is the actual
trigger ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md)): the moment something
outside the mesh must be governed by the same rules, product-level content moves down a level
and this repository becomes what its name already claims.
- A new repository was created rather than the old one transferred, because the forge predates
the transfer API. The original was verified to contain nothing the new one lacks — every ref
an ancestor, no tags, issues, pull requests, releases or wiki content — and then removed.
- The mesh's own documents now live one conceptual level below the repository they are in. A
reader arriving at `01-RESEARCH` should understand it as the mesh's research, not Novox's.
Nothing in the folder names says so, and that is the cost of not restructuring.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name that forced the question.
- [ADR 0019](0019-hq-is-its-own-repository.md) — why HQ is a repository at all. Unchanged; only
its scope moves.
@@ -1,101 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 29. The lab's first scenario has no pipeline, and the lab comes first
## Context
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
*is* a pipeline result.
That is the right design for testing a module against the mesh that exists. It is unusable for
the thing now being built.
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
all four are tier 2 and do not exist yet.
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
that takes over a machine's packages, services and network — it cannot be developed against a
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
after.
## Considered options
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
over. Developing something that reformats a machine, against a machine that is in use, is
how a machine is lost. And it would leave the bootstrap path exercised only when performed
for real — which is precisely the property that makes the current first-node script the
least-tested code in the system.
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
tiers it is meant to test.
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
## Decision
The lab has **two scenario classes**, and the first has no pipeline in it at all.
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
| Exists to | develop the mesh | test what runs on it |
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
Nothing forks, which is the same rule the existing design already holds itself to.
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
is resequenced accordingly. It is the environment everything else is developed inside.
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
immediately** — something must materialise, snapshot and destroy a mesh before anything else
can be written. **Assertion execution comes later**, with the full scenario, because a
bootstrap scenario's assertions are about the state a single host reconciled and are small
enough to state directly.
## Consequences
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
currently a script that runs when a node is created and is otherwise never touched. Under
this decision it is the inner development loop for every change to tiers 0 and 1.
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
only in what is placed inside the machines.
- The lab acquires a second audience. It was designed for a module author and now also serves
whoever is building the mesh itself — which is the same "one runner, two callers" argument
the design already makes, extended one step.
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
system containers and would not be with virtual machines — the unit choice is what makes the
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
a lab node is a virtual machine, and the scale argument for system containers was found to
have been invented rather than required. The design text did not follow the decision. It does
now.
- The lab's home is `novox/mesh-lab`, recorded in
[ADR 0030](0030-the-repository-structure.md) — written after this record, because this one
needed a repository that no decision had yet named.
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
how a real node is raised, it certifies something that does not happen. That is the same
hazard the existing design names for the full scenario, and the same answer applies —
nothing new drives it, and what runs is the real thing.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
bundle, which is also exactly the bootstrap path.
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
reorders.
@@ -1,99 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 30. The repository structure, and the rule that names them
## Context
The tiers are settled ([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md))
and the product is named ([ADR 0027](0027-the-product-is-novox-mesh.md)), but the repositories
themselves were only ever sketched in research. Two consequences had already appeared.
[ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) makes the lab phase 0 of the entire
migration and **could not say where it lives**, because no record named a repository.
And the research contradicted an accepted record: it listed `mesh-hq` for this repository, while
[ADR 0028](0028-hq-is-company-scoped.md) had decided `novox/hq` and explicitly rejected that
name. A design resting on research is a design resting on something that can change without a
decision.
There is also an implied naming rule that has never been written down. ADR 0027 says
repository names take `mesh`; ADR 0028 gives this repository no prefix at all. Both are right,
for a reason neither states.
## Considered options
Only the naming rule had genuine alternatives; the tier repositories follow from the tiers.
1. **No prefix — `novox/host`, `novox/control`.** The organisation already says Novox, so the
prefix reads as stutter. Rejected once it was established that Novox delivers more than the
mesh: with several products the prefix is not stutter, it is the product namespace doing
real work, and the forge has no nested groups to do it instead.
2. **An organisation per product — `novox-mesh/host`.** Puts the product boundary where the
forge's only real grouping primitive lives, so permissions and teams attach to it. Rejected
for now as premature: no per-product access boundary exists yet, and it costs `novox-`
repeated across every organisation.
3. **Product-prefixed repositories in the company organisation.** Chosen.
## Decision
**The naming rule:** a repository that belongs to a product carries that product's prefix. A
repository that is company-scoped does not.
That is why this one is `hq` and the mesh's are `mesh-*`. Both records were already correct;
the rule connecting them is stated here.
**The repositories:**
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the four pinned services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | ~~contracts shared across tiers: types, not behaviour~~ → **the mesh's own domain** ([ADR 0065](0065-the-core-library-is-the-meshs-domain.md)) |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this repository. Company-scoped ([ADR 0028](0028-hq-is-company-scoped.md)) |
**The lab is its own repository.** Its lifecycle differs from everything else in the list: it
is never shipped to a node, it outlives any single tier, and it drives virtualisation on a
workstation — which nothing else in the mesh does. Putting it inside the host would couple
development tooling to a shipped component; putting it inside the control plane would make the
bootstrap scenario depend on a tier that does not exist when it is needed.
**Tier 4 is deliberately not decided here.** Whether the catalogue is one repository, one per
domain, or one per application remains open from
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) and is blocked on
[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md): how many repositories hold
domains cannot be answered before what the domains are. Recording the gap is the point —
`mesh-catalog` appears in the research sketch and is **not** decided by this record.
## Consequences
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) can name its target. Phase 0 has
a home, which was the immediate blocker.
- The research sketch stops being load-bearing. It remains what it is — a sketch — and the
design layer can now cite a record instead.
- **Seven repositories where there is currently one**, for a mesh that today lives in a single
monorepo. That is the cost, and it is not small: seven release cadences, seven sets of
dependencies, and cross-repository changes that were previously one commit. The offsetting
argument is the tier rule — a boundary that only points downward is enforceable across
repositories and merely conventional inside one.
- The prefix will read as redundant for as long as the mesh is the only product with
repositories. That is accepted deliberately: the alternative is renaming everything at the
moment a second product appears, which is the class of migration this project is trying to
stop performing.
- Nothing is created yet. This records what the repositories *are*; creating them is part of
phase 0 and after.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name the prefix comes from.
- [ADR 0028](0028-hq-is-company-scoped.md) — why this repository has no prefix.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lab, and why it is first.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
sketch this supersedes as a source.
@@ -1,87 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 31. The lab provides the underlay; the mesh builds the overlay
## Context
[Research 004](../01-RESEARCH/004-lab-network/analysis.md) worked out the topology a lab has to
reproduce: a routable segment using documentation addresses, a household segment behind NAT, a
router that forwards exactly one port so a *published-but-NATed* node is real, and a machine
that can attach to either segment or detach entirely.
It also records what makes that topology **mean** something, and this is where a boundary has
to be drawn. Hub election is by convention rather than by flag — the hub is the node whose
profile is server and whose overlay address begins `10.10.0.1`. Direct peering depends on two
nodes sharing a site. Names resolve from mesh configuration on each node.
Those are all facts the *mesh* establishes. The question is whether a scenario declares them.
It is tempting to say yes, because a scenario that hands you a working overlay is a scenario
you can start testing against immediately.
## Considered options
1. **The lab configures the overlay too** — assign the overlay addresses, elect the hub, write
the peer configuration, seed the names. Rejected, and the reason is the whole point of the
lab: **a lab that builds the overlay certifies its own work.** If the mesh's peering logic
is broken, a scenario that pre-built the peering still comes up green. The most valuable
thing the lab can test is precisely the part this would replace.
2. **The lab provides nothing but bare machines** — no addressing, no segments, no NAT. Also
rejected. Then the scenario cannot reproduce *published but behind NAT*, which research 004
identifies as the case that only exists in production today, and the lab loses its reason to
use virtual machines at all.
3. **The lab provides the underlay; the mesh builds the overlay.** Chosen.
## Decision
**A scenario declares the underlay** — the facts a machine would have before any of our
software touched it:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and can be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering,
no names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
The rule stated in one line: **a scenario provides what a hosting provider and a home router
would provide, and nothing our software is responsible for.**
## Consequences
- **The overlay becomes a thing under test rather than a fixture.** Whether peers form,
whether the hub is elected, whether a NATed node's endpoint is learned — all of it is
observed rather than arranged. That is the class of fault research 004 says is discoverable
only in production today.
- The lab stays small, and stays honest. It needs to know about virtualisation, bridges,
addresses and NAT. It never needs to know what a mesh node is.
- A scenario cannot assert "the overlay came up" as a precondition, because it is an outcome.
A bootstrap scenario that wants a working overlay has to wait for one and check, which is
the correct shape.
- **The address ranges are load-bearing, not cosmetic.** The routable segment uses RFC 5737
documentation space specifically because the mesh's own code decides *public versus private*
by matching the address — a private range there makes the hub test as unreachable, and the
mesh silently never forms. Research 004 calls this the single most important fact in the
document, and the declaration format has to make getting it wrong hard.
- The router is a machine the lab materialises without being asked, because NAT requires
somewhere to run. That is an implicit machine in an otherwise explicit declaration, and it is
worth knowing about rather than discovering.
- **Host capability profiles are detected, not declared** — a consequence of
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md), and consistent here: a
scenario does not say what a machine is allowed to do, it provides a machine. Which leaves an
open question: a lab machine is always privileged, so the `user` and `edge` profiles have no
scenario that exercises them yet.
## References
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology, the documentation
ranges, and the hub-election and peering conventions this deliberately does not touch.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the two scenario classes this
declaration has to serve without forking.
@@ -1,71 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 32. A scenario is an isolated address space, and the lab never reaches into it over IP
## Context
A scenario declares literal addresses —
[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has
to be, because reproducing *published but behind NAT* means saying which address the world sees.
That raises a question the declaration left open: **two scenarios at once.** Several agents
working means several scenarios, and the lab design already calls that a requirement. But two
scenarios built from the same declaration want the same addresses, and there are only three
documentation ranges in existence.
## Considered options
1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals.
Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a
specific topology no longer reproduces it; it breaks the RFC-range validation, since
allocated addresses would have to come from somewhere real; and the numbers a person reads
in the file stop being the numbers they will see in a capture.
2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate
an agent has to queue for is a gate that gets bypassed.
3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen.
## Decision
**A scenario is a closed address space.** Every segment materialises as its own isolated link,
belonging to one scenario instance. Two scenarios raised from the same declaration hold the same
addresses and never meet, because nothing joins their links.
The declaration therefore keeps its literal addresses, and they mean exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the same way one
executes a command in a container without the container being routable.
That is not a preference. If the lab reached machines by address, the workstation running it
would need a route into each scenario, and two scenarios carrying the same prefix would give it
two routes to the same destination. Concurrency would be impossible, and it would fail in the
worst available way: not with an error, but by one scenario's traffic arriving in another.
## Consequences
- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit
beyond the machine's capacity.
- The three documentation ranges stop being a scarce resource. Every scenario may use all of
them, because no two scenarios share a link.
- **The lab cannot use IP to check anything**, which is more of a constraint than it first
appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by
executing on a machine, rather than probed from outside. That is the honest way to ask it
anyway — reachability from the workstation is not the question.
- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing
outside holds a reference into it.
- The lab needs a scenario **instance** identity distinct from the scenario name in the
declaration: the declaration is a kind, and several instances of one kind may exist.
- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a
developer wants to reach — a web interface, a database — needs an explicit, deliberate
forward out of the scenario, which is a feature rather than a gap: nothing leaks by default.
## References
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses
this preserves.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes.
@@ -1,81 +0,0 @@
---
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 33. A router is scenery, not a node — so it is a container
## Context
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual
machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install,
so it has to be a real machine or the thing under test is not the thing that ships.
A scenario also needs routers. NAT, port forwarding, policy between segments and mapping
expiry are all things a router does, and until one is materialised a multi-segment scenario
raises isolated islands
([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)).
The declaration already implies them: a gateway is *the one implicit machine in an otherwise
explicit declaration*.
The question is whether ADR 0016 binds those too.
## Considered options
1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency
nobody needs. A router boots in roughly ten seconds against a container's one; a
six-segment scenario wanting three routers spends thirty seconds per raise on scenery.
2. **The hypervisor provides NAT** — bridges with translation switched on, and its own
forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide
what the declaration is supposed to own, and it cannot express a mapping that expires, a
gateway that refuses to forward, or policy between siblings. The model would shrink to fit
the tool.
3. **A router is scenery, and scenery is a container.** Chosen.
## Decision
**ADR 0016 binds nodes. A router is not a node.**
Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh
never installs anything on it, and no assertion is ever made about its internals. It exists so
that packets between machines behave the way they behave in the world — which is the definition
of scenery.
So a router is a **system container**, and the fidelity argument does not reach it: what a
router must reproduce is kernel behaviour — translation, connection tracking, filtering,
forwarding — and a container has the same kernel.
**Verified before deciding, not assumed.** In a plain unprivileged container:
| Needed for | Works |
|---|---|
| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` |
| `nat:` | nftables masquerade, rules accepted and listed back |
| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` |
No privileged mode, no nesting, no capability grants.
## Consequences
- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten,
and a scenario's cost tracks the machines actually under test.
- **The distinction is now load-bearing and has to stay legible.** *Node* means something under
test; *scenery* means something that makes the test real. If anything is ever installed on a
router by the mesh, it has become a node and this decision no longer covers it.
- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra
concept — justified by it being true, rather than by the saving.
- A container shares the host kernel, so a scenario cannot reproduce a router running a
*different* kernel from the workstation. Nothing currently wants that; if something does, that
router becomes a virtual machine and this record needs revisiting rather than bending.
- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and
never names the machine that serves it — which is right, because it is not a machine the
scenario has anything to say about.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged.
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router,
and why *published but behind NAT* only exists in production today.
@@ -92,7 +92,7 @@ difference read off directly.
## References
- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the
- [ADR 0016](0016-the-lab.md) — why the lab must not supply what the
mesh is responsible for; the same instinct, applied to facts rather than to configuration.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production
form.
@@ -93,5 +93,5 @@ fixing a bug.
measured co-change cluster in the catalogue.
- [ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) — what the control plane decides
from.
- [ADR 0030](0030-the-repository-structure.md) — `mesh-host` as tier 0.
- [ADR 0019](0019-how-this-repository-works.md) — `mesh-host` as tier 0.
- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard the direction lint is held to.
@@ -10,7 +10,7 @@ extends: 0037-the-host-applies-it-does-not-decide.md
## Context
[ADR 0030](0030-the-repository-structure.md) calls tier 0 *"the one binary installed by hand"*,
[ADR 0019](0019-how-this-repository-works.md) calls tier 0 *"the one binary installed by hand"*,
and [research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) states the property the
whole tier rests on: *"a binary whose whole argument is that it has no dependencies"*.
@@ -72,7 +72,7 @@ A second language usually costs duplicated logic; here there is none to duplicat
## References
- [ADR 0030](0030-the-repository-structure.md) — *the one binary installed by hand*.
- [ADR 0019](0019-how-this-repository-works.md) — *the one binary installed by hand*.
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host shares no code.
- [ADR 0039](0039-the-link-is-the-security-boundary.md) — why it receives declarations only.
- [`05-the-node-host.md`](../03-DESIGN/01-to-be/05-the-node-host.md) — the design this serves.
+1 -1
View File
@@ -3,7 +3,7 @@ status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
extends: 0030-the-repository-structure.md
extends: 0019-how-this-repository-works.md
---
# 48. The substrate is named
@@ -3,7 +3,7 @@ status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
extends: 0031-the-lab-provides-the-underlay.md
extends: 0016-the-lab.md
---
# 50. Reachability is declared, not inferred from an address
@@ -86,7 +86,7 @@ committed.
*record* that still authorised the arrangement — which was the last thing anybody could have
cited in its defence.
- **Two as-is documents cite 0003 and describe today's behaviour accurately.** They are not
wrong and should not be changed: [ADR 0020](0020-design-is-written-in-two-layers.md) keeps the
wrong and should not be changed: [ADR 0019](0019-how-this-repository-works.md) keeps the
layers separate, and *as-is* describing a superseded decision is exactly what as-is is for. What
changes is the citation's status, not its content.
- **It does not say how today's mesh gets there.** Thirteen tables move, two modules are rewritten,
@@ -10,7 +10,7 @@ extends: 0064-a-build-edge-is-a-third-kind.md
## Context
[ADR 0030](0030-the-repository-structure.md) describes the shared library as *"contracts shared
[ADR 0019](0019-how-this-repository-works.md) describes the shared library as *"contracts shared
across tiers: types, not behaviour"* — a guard against what the current one became, which is a
package holding too much code and, with it, everybody's dependencies.
@@ -62,7 +62,7 @@ than fighting for it.
## Consequences
- **[ADR 0030](0030-the-repository-structure.md)'s description of the shared library is
- **[ADR 0019](0019-how-this-repository-works.md)'s description of the shared library is
superseded.** *Types, not behaviour* is replaced by *the mesh's domain*, and its types move to
the modules that own them. The rest of 0030 — the naming rule and the repository list — stands.
- **A module now publishes its own contract**, which it does not do today. That is real work and
@@ -82,6 +82,6 @@ than fighting for it.
## References
- [ADR 0030](0030-the-repository-structure.md) — the description this replaces.
- [ADR 0019](0019-how-this-repository-works.md) — the description this replaces.
- [ADR 0064](0064-a-build-edge-is-a-third-kind.md) — what makes fan-in measurable.
- [Research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — module, node, assignment.
+1 -1
View File
@@ -15,7 +15,7 @@ The records run in the order the decisions were taken, oldest first.
**Every decision is a record.** There is no ledger and no index file — if a decision is worth
recording it is worth a record, and if it is not worth a record it is not recorded
([ADR 0026](0026-every-decision-is-a-record.md)). A "decision" small enough to be one line is
([ADR 0019](0019-how-this-repository-works.md)). A "decision" small enough to be one line is
almost always a **rule**, and a rule belongs in
[`00-META/how-we-build.md`](../00-META/how-we-build.md), where it is enforced and keeps the
incident that earned it.
+5 -5
View File
@@ -4,11 +4,11 @@ status: implemented
code: [mesh-lab]
updated: 2026-08-25
decisions:
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
- 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
- 02-DECISIONS/0033-a-router-is-scenery-not-a-node.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
---
# The lab, as it stands
+5 -5
View File
@@ -4,9 +4,9 @@ status: in-progress
code: [mesh-lab]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0019-how-this-repository-works.md
---
# End-to-end testing
@@ -37,7 +37,7 @@ today, that is a gap in the vocabulary rather than a reason to privilege that sh
The design below describes a scenario as a complete mesh — forge (Gitea), coordinator,
delivery cascade — because what it tests is a module. **That is the larger of two classes, and
not the first one built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
not the first one built** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
@@ -127,7 +127,7 @@ drifts.
a mesh named by the request instead.
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
at once, each needing its own network and nodes. A lab node is a virtual machine
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)), and snapshots are
what make repetition cheap — restoring a scenario costs far less than building one. The
earlier argument here, that only system containers made this affordable, was superseded: the
scale it assumed was invented rather than required.
@@ -4,9 +4,9 @@ status: in-progress
code: [mesh-lab]
updated: 2026-08-25
decisions:
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
---
# The scenario declaration
@@ -15,7 +15,7 @@ A scenario is a **declaration of an underlay**, plus what to put on it. It is th
everything in the lab hangs off, so it is worth getting small.
It states what a hosting provider and a home router would provide, and nothing the mesh is
responsible for ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
responsible for ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Public networks are unrelated, and routed rather than bridged
@@ -258,7 +258,7 @@ otherwise explicit declaration, and it exists because NAT has to run somewhere.
It is a **container, not a virtual machine** — a router is scenery rather than something under
test, so the fidelity argument that makes a node a virtual machine does not reach it
([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md)). What a router must
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). What a router must
reproduce is kernel behaviour, and a container has the same kernel.
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
@@ -299,7 +299,7 @@ belongs to a router it does not control, and asleep.
Whether the overlay survives that, re-forms, and is noticed to have changed endpoint is
**observed**, never arranged
([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Why the addresses are load-bearing
@@ -326,7 +326,7 @@ file: the failure it prevents is silent, so the check has to be loud.
## The same declaration serves both classes
The bootstrap and full scenarios differ **only in `place:`**
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). Everything
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). Everything
about the underlay is identical, which is what makes one a strict subset of the other rather
than a fork.
@@ -354,7 +354,7 @@ not first.
## What a scenario deliberately cannot say
- **Overlay addresses, the hub, peer configuration.** Outcomes, not inputs
([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
- **What a machine is in mesh terms** — server or workstation, its site, its names. Mesh
configuration, established by the mesh.
- **A host's capability profile.** Detected, never declared.
@@ -543,7 +543,7 @@ cannot yet express.
Nothing here mentions overlay addresses, which node is the hub, who peers with whom, any name,
or any certificate. Research 004 recorded all of those for this topology, and **a scenario must
not state them** ([ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md)): they
not state them** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)): they
are what the mesh does, and a scenario that supplied them would be certifying its own work.
The absence is the point. Given the declaration above, whether a hub is elected, whether the
+6 -6
View File
@@ -4,16 +4,16 @@ status: in-progress
code: [mesh-lab]
updated: 2026-08-25
decisions:
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
- 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
---
# Scenario lifecycle
The first thing the lab must do, and the only thing it must do before anything else can be
written: **materialise a mesh, return it to a known state, and destroy it**
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens
to one.
@@ -39,7 +39,7 @@ The order is not arbitrary — each step needs the one before it to exist:
1. **Segments.** Isolated links, one per declared segment, belonging to this instance and
joined to nothing outside it
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each
distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the
translation, forwarding and mapping-expiry the declaration asked for.
@@ -100,7 +100,7 @@ made after it, and returning undoes it like any other change.
## Reaching in
Everything the lab does to a machine goes through the virtualisation layer, never over IP
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). `exec` runs a
command on a machine and returns its output.
This has one consequence worth stating plainly: **a reachability question is asked from inside**.
+2 -2
View File
@@ -4,7 +4,7 @@ status: designed
code: [mesh-lab]
updated: 2026-08-24
decisions:
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
---
@@ -12,7 +12,7 @@ decisions:
The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity
permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is
built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
built ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
So the lab needs an install path of its own. This describes it, and the shape it has to take is
determined by two failures observed while measuring
+1 -1
View File
@@ -4,7 +4,7 @@ status: in-progress
code: [mesh-host]
updated: 2026-08-27
decisions:
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0019-how-this-repository-works.md
- 02-DECISIONS/0036-a-node-is-a-managed-machine.md
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0038-a-node-joins-by-linking-first.md
+1 -1
View File
@@ -6,7 +6,7 @@ updated: 2026-08-27
decisions:
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0053-one-control-plane-and-no-failover.md
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0019-how-this-repository-works.md
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0055-the-control-plane-is-the-node-coordinating-contexts.md
---
+1 -1
View File
@@ -4,7 +4,7 @@ status: designed
code: []
updated: 2026-08-27
decisions:
- 02-DECISIONS/0030-the-repository-structure.md
- 02-DECISIONS/0019-how-this-repository-works.md
- 02-DECISIONS/0037-the-host-applies-it-does-not-decide.md
- 02-DECISIONS/0038-a-node-joins-by-linking-first.md
- 02-DECISIONS/0046-the-installer-fetches-what-it-pins.md
+3 -3
View File
@@ -10,9 +10,9 @@ document is written and this one's status becomes `implemented`.
| Document | Covers | Rests on |
|---|---|---|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) |
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-the-lab.md), [0029](../../02-DECISIONS/0016-the-lab.md) |
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0016](../../02-DECISIONS/0016-the-lab.md) |
| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0016](../../02-DECISIONS/0016-the-lab.md) |
| [`04-lab-installation.md`](04-lab-installation.md) | Getting the lab onto a clean machine, and why it verifies capability rather than installation | [ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) |
| [`05-the-node-host.md`](05-the-node-host.md) | Tier 0 — the one thing installed by hand, and the only thing that changes a machine | [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) |
| [`06-the-control-plane.md`](06-the-control-plane.md) | Tier 2 — what the term means, and the test for what belongs in it | [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) |
@@ -38,7 +38,7 @@ This is an instance where it was never applied.
## Evidence
- Observed 2026-08-22 while declaring the virtualisation package required by
[ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md).
[ADR 0016](../../02-DECISIONS/0016-the-lab.md).
- A fix is written and open as a pull request, unmerged since 2026-08-20.
## Open questions
@@ -21,7 +21,7 @@ recoverable by retrying — it removes the ability to issue a certificate anyone
The consequence lands hardest on exactly the work most likely to iterate: standing up a new
node, changing how names resolve, or testing the lab's certificate authority split
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Evidence
@@ -59,7 +59,7 @@ checked it — including in the same commit that wrote the rule.
## Proposed direction — Nox is the search
*Added 2026-08-23.* Rather than syncing these documents into the knowledge base, **Nox
([ADR 0027](../../02-DECISIONS/0027-the-product-is-novox-mesh.md)) works from within this
([ADR 0019](../../02-DECISIONS/0019-how-this-repository-works.md)) works from within this
repository and holds its knowledge directly.** Retrieval becomes an agent reading the source,
not a copy living in a second store.
@@ -44,7 +44,7 @@ The distance between the two is the same one the delivery layer already has a na
## Why it matters now
This is the first requirement of the lab
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)), which is
phase 0 of the entire migration. The first capability the new work depends on is present,
declared, and unusable — and would have stayed unusable silently.
@@ -74,7 +74,7 @@ It also removes the lab's export-and-push mechanism rather than fixing it, which
outcome: pushing image tarballs over the hypervisor was always a lab-only invention.
**Not decided here**, because it is design rather than repair: where the registry runs, whether
it is scenery like the router ([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md))
it is scenery like the router ([ADR 0016](../../02-DECISIONS/0016-the-lab.md))
or a placed artifact, and how images get into it.
## Incidental, and already fixed
+1 -1
View File
@@ -1,7 +1,7 @@
# Agent instructions — Novox HQ
This repository is the source of truth for Novox's mission, research, design and decisions —
today almost entirely those of **Novox Mesh**, its first product ([ADR 0028](02-DECISIONS/0028-hq-is-company-scoped.md)). Implementation lives in the code repositories (see
today almost entirely those of **Novox Mesh**, its first product ([ADR 0019](02-DECISIONS/0019-how-this-repository-works.md)). Implementation lives in the code repositories (see
[`00-META/repos.md`](00-META/repos.md)).
Before changing anything here, read the playbooks in