The lab comes first, and its first scenario has no pipeline #4

Merged
jschoubben merged 3 commits from design/lab-bootstrap-scenario into main 2026-08-23 20:24:16 +00:00
8 changed files with 378 additions and 8 deletions
+17
View File
@@ -18,6 +18,23 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0028](../02-DECISIONS/0028-hq-is-company-scoped.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
## What the mesh becomes
[ADR 0030](../02-DECISIONS/0030-the-repository-structure.md) records the repositories the
monorepo decomposes into. **None exist yet** — they are the target, not the present.
| Repository | Tier | Holds |
|---|---|---|
| `mesh-host` | 0 | the node host — the one binary installed by hand |
| `mesh-substrate` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | the control plane and its contexts |
| `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers |
| `mesh-lab` | — | the lab — built first, per [ADR 0029](../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
Tier 4's shape is open, and deliberately so: see ADR 0030 and
[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md).
## What lives where inside the monorepo
Named by role, because the layout is itself part of the as-is design — see
@@ -256,7 +256,7 @@ mesh-surfaces/ TIER 3
mesh-catalog/ TIER 4
<domain>/<module>/ layout as above
mesh-lab/ mesh-sdk/ mesh-hq/
mesh-lab/ mesh-sdk/ hq/ (company-scoped — ADR 0028)
```
## What this does not settle
@@ -84,7 +84,8 @@ mesh-catalog/ TIER 4 — what the mesh hosts
mesh-lab/ the whole mesh, disposable, on one machine
mesh-sdk/ contracts shared across tiers — types, not behaviour
mesh-hq/ this repository
hq/ company-scoped, not a mesh repository — ADR 0028
```
## The dependency rule
+6 -3
View File
@@ -30,7 +30,9 @@ Recorded because incremental is the reflex answer and it is wrong in this case.
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
same risk as one performed for the first time on the real mesh.
same risk as one performed for the first time on the real mesh. This is also why the lab is
phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by
cutover the procedure has been run hundreds of times rather than rehearsed a few.
- **Incremental would carry the rot forward.** The as-is layer documents silent failure paths,
a dead test harness and unenforced rules. A gradual migration preserves them by definition.
@@ -68,8 +70,9 @@ than a discovery.
| Phase | What | Done when |
|---|---|---|
| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
| D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. |
| E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. |
@@ -0,0 +1,101 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 29. The lab's first scenario has no pipeline, and the lab comes first
## Context
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
*is* a pipeline result.
That is the right design for testing a module against the mesh that exists. It is unusable for
the thing now being built.
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
all four are tier 2 and do not exist yet.
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
that takes over a machine's packages, services and network — it cannot be developed against a
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
after.
## Considered options
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
over. Developing something that reformats a machine, against a machine that is in use, is
how a machine is lost. And it would leave the bootstrap path exercised only when performed
for real — which is precisely the property that makes the current first-node script the
least-tested code in the system.
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
tiers it is meant to test.
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
## Decision
The lab has **two scenario classes**, and the first has no pipeline in it at all.
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
| Exists to | develop the mesh | test what runs on it |
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
Nothing forks, which is the same rule the existing design already holds itself to.
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
is resequenced accordingly. It is the environment everything else is developed inside.
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
immediately** — something must materialise, snapshot and destroy a mesh before anything else
can be written. **Assertion execution comes later**, with the full scenario, because a
bootstrap scenario's assertions are about the state a single host reconciled and are small
enough to state directly.
## Consequences
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
currently a script that runs when a node is created and is otherwise never touched. Under
this decision it is the inner development loop for every change to tiers 0 and 1.
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
only in what is placed inside the machines.
- The lab acquires a second audience. It was designed for a module author and now also serves
whoever is building the mesh itself — which is the same "one runner, two callers" argument
the design already makes, extended one step.
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
system containers and would not be with virtual machines — the unit choice is what makes the
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
a lab node is a virtual machine, and the scale argument for system containers was found to
have been invented rather than required. The design text did not follow the decision. It does
now.
- The lab's home is `novox/mesh-lab`, recorded in
[ADR 0030](0030-the-repository-structure.md) — written after this record, because this one
needed a repository that no decision had yet named.
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
how a real node is raised, it certifies something that does not happen. That is the same
hazard the existing design names for the full scenario, and the same answer applies —
nothing new drives it, and what runs is the real thing.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
bundle, which is also exactly the bootstrap path.
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
reorders.
@@ -0,0 +1,99 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 30. The repository structure, and the rule that names them
## Context
The tiers are settled ([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md))
and the product is named ([ADR 0027](0027-the-product-is-novox-mesh.md)), but the repositories
themselves were only ever sketched in research. Two consequences had already appeared.
[ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) makes the lab phase 0 of the entire
migration and **could not say where it lives**, because no record named a repository.
And the research contradicted an accepted record: it listed `mesh-hq` for this repository, while
[ADR 0028](0028-hq-is-company-scoped.md) had decided `novox/hq` and explicitly rejected that
name. A design resting on research is a design resting on something that can change without a
decision.
There is also an implied naming rule that has never been written down. ADR 0027 says
repository names take `mesh`; ADR 0028 gives this repository no prefix at all. Both are right,
for a reason neither states.
## Considered options
Only the naming rule had genuine alternatives; the tier repositories follow from the tiers.
1. **No prefix — `novox/host`, `novox/control`.** The organisation already says Novox, so the
prefix reads as stutter. Rejected once it was established that Novox delivers more than the
mesh: with several products the prefix is not stutter, it is the product namespace doing
real work, and the forge has no nested groups to do it instead.
2. **An organisation per product — `novox-mesh/host`.** Puts the product boundary where the
forge's only real grouping primitive lives, so permissions and teams attach to it. Rejected
for now as premature: no per-product access boundary exists yet, and it costs `novox-`
repeated across every organisation.
3. **Product-prefixed repositories in the company organisation.** Chosen.
## Decision
**The naming rule:** a repository that belongs to a product carries that product's prefix. A
repository that is company-scoped does not.
That is why this one is `hq` and the mesh's are `mesh-*`. Both records were already correct;
the rule connecting them is stated here.
**The repositories:**
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the four pinned services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | contracts shared across tiers: types, not behaviour |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this repository. Company-scoped ([ADR 0028](0028-hq-is-company-scoped.md)) |
**The lab is its own repository.** Its lifecycle differs from everything else in the list: it
is never shipped to a node, it outlives any single tier, and it drives virtualisation on a
workstation — which nothing else in the mesh does. Putting it inside the host would couple
development tooling to a shipped component; putting it inside the control plane would make the
bootstrap scenario depend on a tier that does not exist when it is needed.
**Tier 4 is deliberately not decided here.** Whether the catalogue is one repository, one per
domain, or one per application remains open from
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) and is blocked on
[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md): how many repositories hold
domains cannot be answered before what the domains are. Recording the gap is the point —
`mesh-catalog` appears in the research sketch and is **not** decided by this record.
## Consequences
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) can name its target. Phase 0 has
a home, which was the immediate blocker.
- The research sketch stops being load-bearing. It remains what it is — a sketch — and the
design layer can now cite a record instead.
- **Seven repositories where there is currently one**, for a mesh that today lives in a single
monorepo. That is the cost, and it is not small: seven release cadences, seven sets of
dependencies, and cross-repository changes that were previously one commit. The offsetting
argument is the tier rule — a boundary that only points downward is enforceable across
repositories and merely conventional inside one.
- The prefix will read as redundant for as long as the mesh is the only product with
repositories. That is accepted deliberately: the alternative is renaming everything at the
moment a second product appears, which is the class of migration this project is trying to
stop performing.
- Nothing is created yet. This records what the repositories *are*; creating them is part of
phase 0 and after.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name the prefix comes from.
- [ADR 0028](0028-hq-is-company-scoped.md) — why this repository has no prefix.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lab, and why it is first.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
sketch this supersedes as a source.
+35 -3
View File
@@ -30,6 +30,36 @@ its question needs. If something can only be tested against the shape the mesh h
today, that is a gap in the vocabulary rather than a reason to privilege that shape.
## Two classes of scenario
The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade
— because what it tests is a module. **That is the larger of two classes, and not the first one
built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | the node host and the substrate | the control plane and everything above it |
| Exists to | **develop the mesh** | **test what runs on it** |
The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same
lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in
the way work happens"* onward describes the full scenario, and applies once there is a
coordinator to describe.
**The bootstrap scenario is built first, ahead of everything it will later test.** The
component that takes over a machine's packages, services and network cannot be developed
against a machine anyone needs, and raising a node from nothing is today the least-exercised
path in the system precisely because it only ever runs for real. Making it the inner
development loop inverts that.
Of the runner's two jobs, this settles their order: **scenario lifecycle first** — something
must materialise, snapshot and destroy a mesh before anything else can be written — and
**assertion execution second**, with the full scenario, since a bootstrap scenario's assertions
concern the state one host reconciled and are small enough to state directly.
## Where this sits in the way work happens
Work reaches the mesh along one path today:
@@ -93,9 +123,11 @@ drifts.
nodes that have the module assigned. It needs to be able to run the same pipeline against
a mesh named by the request instead.
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
at once, each needing its own network and nodes. This is affordable with system containers
and would not be with virtual machines — the unit choice is what makes the gate possible
at all.
at once, each needing its own network and nodes. A lab node is a virtual machine
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are
what make repetition cheap — restoring a scenario costs far less than building one. The
earlier argument here, that only system containers made this affordable, was superseded: the
scale it assumed was invented rather than required.
- **The gate is only as good as the verification behind it.** A module with no assertions
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
the number that matters**, and it starts at approximately zero.
@@ -0,0 +1,117 @@
---
status: diagnosing
opened: 2026-08-23
located-in: [hal]
fixed-by:
- "the instance only: PR #962 — incus hook. Merged, ran on one node in 6s of a 60s budget; pipeline #6832 green in 48s. The class remains open."
amended-design:
---
# 007 — An installed package is not an available capability
## Symptom
A module declares the virtualisation package the lab needs. The package is installed —
version 7.3.0-1, recorded as explicitly installed. The client binary runs.
The capability does not exist:
| Checked | State |
|---|---|
| `incus.service` | disabled, inactive |
| `incus.socket` | disabled, inactive |
| `incus-user.socket` | disabled, inactive |
| the operator's group membership | not a member of any incus group |
| the client | reports **`Server version: unreachable`** |
Nothing failed. Nothing reported anything. The declaration was satisfied exactly as written,
and the thing it was declared for cannot be used.
## Why this is not issue 001 again
[Issue 001](../001-failed-package-install-reports-success/00-report.md) is *the install failed
and the job reported success*. This is the opposite and arguably worse: **the install
succeeded, and success was not the point.**
A package is a set of files. A capability is a running service, an enabled socket, and an
identity permitted to reach it. The module model declares the first and has no vocabulary for
the second, so the gap between them is invisible — there is no state in which the mesh believes
this node has virtualisation and is wrong, because the mesh was never asked to believe it.
The distance between the two is the same one the delivery layer already has a name for:
**transport versus effect.** A package install reports that files arrived, which is transport.
## Why it matters now
This is the first requirement of the lab
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is
phase 0 of the entire migration. The first capability the new work depends on is present,
declared, and unusable — and would have stayed unusable silently.
It also generalises. Every module that declares a package needing a unit enabled, a group
joined, a kernel module loaded, or a socket activated has this gap. Post-install work lives in
hooks, and hooks have their own recorded failure mode: thirteen were found that had never run.
## Evidence
- Package recorded as installed 2026-08-23 00:02, explicitly.
- Both units disabled and inactive; the operator in no incus group; client reports the server
unreachable.
## Open questions
- Should a module be able to declare a **capability** — a unit that must be enabled, a group
the operator must be in — rather than only the package that provides it?
- If that is what hooks are for, why is the gap invisible when a hook does not run? A hook that
never fires and a hook that fires and does nothing are indistinguishable today.
- Is this what the verify stage should be asserting? It exists, and a module's own assertions
are meant to test outcomes rather than steps — "the socket accepts a connection" is exactly
that shape.
- How many other declared packages are in this state? Nothing currently reports it, which means
the answer is unknown rather than zero.
## The instance is fixed; the class is what this issue is now about
*Updated 2026-08-23.* A hook now does the post-install work — group membership, subordinate id
ranges, enabling both units, creating the storage pool, the bridge, and the default profile.
Verified independently afterwards: the group exists with the operator in it, both id files
carry the range, the service is active, and the storage pool reports `CREATED`.
**So the mechanism was never missing.** Hooks are exactly the right place for this, and used
properly they work. The gap is narrower and worse than "there is no way to do it":
> The hook did six things. **Six checks were then performed by a human, by hand.** Nothing in
> the pipeline asserted any of them, and a pipeline that dispatched a hook which silently never
> fired would have been green in the same 48 seconds.
That is not hypothetical — a hook named for a feature its module does not carry is skipped
without complaint, and thirteen such hooks were found at once in the past. The distance between
*the hook ran and did six things* and *the hook was dispatched* is invisible from the outside,
and it is the whole of this issue.
### The verification that was done by hand is the assertion set
The six checks performed after the merge are, almost word for word, what the module's own
verification should assert — outcomes, not steps, exactly as the lab design requires:
| Asserted by hand | As a module assertion |
|---|---|
| operator in the admin group | the group exists and contains the operator |
| subordinate id range in both files | both files carry the range |
| both units enabled, service active | the service is active, and still active shortly after |
| storage pool present | the pool exists and reports created |
| bridge present with an address | the bridge exists and holds its address |
| default profile wired to both | the profile references the bridge and the pool |
They currently live in a chat message. Moved into the module, they would run on every delivery
to every node, and the difference between a hook that worked and a hook that was merely
dispatched would stop being something a person has to notice.
### What this issue now asks
- Does a module gain a way to declare a **capability** — the outcome — separately from the
package that provides it, or is "write assertions" the whole answer?
- The verify stage exists and asserts almost nothing. Is this simply the first module that
should use it properly, making the issue a coverage problem rather than a design gap?
- **How many other declared packages are in the state this one was in?** Still unknown, still
unreported by anything, and now demonstrably worth asking.