diff --git a/00-META/repos.md b/00-META/repos.md index 08a7b41..2999eaa 100644 --- a/00-META/repos.md +++ b/00-META/repos.md @@ -18,6 +18,23 @@ and a forge address is an operational detail (see [`README`](../README.md)). | `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0028](../02-DECISIONS/0028-hq-is-company-scoped.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. | | *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. | +## What the mesh becomes + +[ADR 0030](../02-DECISIONS/0030-the-repository-structure.md) records the repositories the +monorepo decomposes into. **None exist yet** — they are the target, not the present. + +| Repository | Tier | Holds | +|---|---|---| +| `mesh-host` | 0 | the node host — the one binary installed by hand | +| `mesh-substrate` | 1 | the four pinned services, as declarations | +| `mesh-control` | 2 | the control plane and its contexts | +| `mesh-surfaces` | 3 | tools, web, cli | +| `mesh-sdk` | — | contracts shared across tiers | +| `mesh-lab` | — | the lab — built first, per [ADR 0029](../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | + +Tier 4's shape is open, and deliberately so: see ADR 0030 and +[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md). + ## What lives where inside the monorepo Named by role, because the layout is itself part of the as-is design — see diff --git a/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md b/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md index fad0bb5..87398b4 100644 --- a/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md +++ b/01-RESEARCH/006-mesh-from-scratch/code-skeleton.md @@ -256,7 +256,7 @@ mesh-surfaces/ TIER 3 mesh-catalog/ TIER 4 // layout as above -mesh-lab/ mesh-sdk/ mesh-hq/ +mesh-lab/ mesh-sdk/ hq/ (company-scoped — ADR 0028) ``` ## What this does not settle diff --git a/01-RESEARCH/006-mesh-from-scratch/skeleton.md b/01-RESEARCH/006-mesh-from-scratch/skeleton.md index 3cb355d..ecd6868 100644 --- a/01-RESEARCH/006-mesh-from-scratch/skeleton.md +++ b/01-RESEARCH/006-mesh-from-scratch/skeleton.md @@ -84,7 +84,8 @@ mesh-catalog/ TIER 4 — what the mesh hosts mesh-lab/ the whole mesh, disposable, on one machine mesh-sdk/ contracts shared across tiers — types, not behaviour -mesh-hq/ this repository + +hq/ company-scoped, not a mesh repository — ADR 0028 ``` ## The dependency rule diff --git a/01-RESEARCH/009-migration/00-overview.md b/01-RESEARCH/009-migration/00-overview.md index f72efb4..832d76a 100644 --- a/01-RESEARCH/009-migration/00-overview.md +++ b/01-RESEARCH/009-migration/00-overview.md @@ -30,7 +30,9 @@ Recorded because incremental is the reflex answer and it is wrong in this case. - **Nothing external depends on it.** No users outside the operator, no service level to hold. - **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)). A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the - same risk as one performed for the first time on the real mesh. + same risk as one performed for the first time on the real mesh. This is also why the lab is + phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by + cutover the procedure has been run hundreds of times rather than rehearsed a few. - **Incremental would carry the rot forward.** The as-is layer documents silent failure paths, a dead test harness and unenforced rules. A gradual migration preserves them by definition. @@ -68,8 +70,9 @@ than a discovery. | Phase | What | Done when | |---|---|---| -| A | Build tier 0. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. | -| B | Build tier 1 and 2. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. | +| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. | +| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. | +| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. | | C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. | | D | Declare the existing workloads against the new model. | The lab runs them, with copies of real data shapes. | | E | **Rehearse the cutover in the lab** against a mesh built to resemble the real one. | Repeatable, and repeatably reversible. | diff --git a/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md b/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md new file mode 100644 index 0000000..0ca6c40 --- /dev/null +++ b/02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md @@ -0,0 +1,101 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +extends: 0016-a-lab-node-is-a-virtual-machine.md +--- + +# 29. The lab's first scenario has no pipeline, and the lab comes first + +## Context + +[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under +test is a module; the mesh is the harness"*, and everything follows from that: a scenario has +its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict +*is* a pipeline result. + +That is the right design for testing a module against the mesh that exists. It is unusable for +the thing now being built. + +**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle +([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that +requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because +all four are tier 2 and do not exist yet. + +**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md) +placed the lab at phase B, as verification of tiers already built. But tier 0 is the component +that takes over a machine's packages, services and network — it cannot be developed against a +machine anyone needs. It needs somewhere disposable to exist **before** it is written, not +after. + +## Considered options + +1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice + over. Developing something that reformats a machine, against a machine that is in use, is + how a machine is lost. And it would leave the bootstrap path exercised only when performed + for real — which is precisely the property that makes the current first-node script the + least-tested code in the system. +2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a + coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the + tiers it is meant to test. +3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen. + +## Decision + +The lab has **two scenario classes**, and the first has no pipeline in it at all. + +| | **Bootstrap scenario** | **Full scenario** | +|---|---|---| +| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules | +| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify | +| Exercises | tiers 0 and 1 | tiers 2 and above, and modules | +| Exists to | develop the mesh | test what runs on it | + +The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the +same networking, the same scenario lifecycle, simply stopping before a control plane exists. +Nothing forks, which is the same rule the existing design already holds itself to. + +**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md) +is resequenced accordingly. It is the environment everything else is developed inside. + +Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed +immediately** — something must materialise, snapshot and destroy a mesh before anything else +can be written. **Assertion execution comes later**, with the full scenario, because a +bootstrap scenario's assertions are about the state a single host reconciled and are small +enough to state directly. + +## Consequences + +- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is + currently a script that runs when a node is created and is otherwise never touched. Under + this decision it is the inner development loop for every change to tiers 0 and 1. +- The first thing built is small: virtualisation, a network, a way to place a binary, and a way + to snapshot and reset. No forge, no coordinator, no pipeline, no modules. +- The full scenario becomes reachable by *addition* rather than by rework, because it differs + only in what is placed inside the machines. +- The lab acquires a second audience. It was designed for a module author and now also serves + whoever is building the mesh itself — which is the same "one runner, two callers" argument + the design already makes, extended one step. +- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with + system containers and would not be with virtual machines — the unit choice is what makes the + gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that: + a lab node is a virtual machine, and the scale argument for system containers was found to + have been invented rather than required. The design text did not follow the decision. It does + now. +- The lab's home is `novox/mesh-lab`, recorded in + [ADR 0030](0030-the-repository-structure.md) — written after this record, because this one + needed a repository that no decision had yet named. +- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from + how a real node is raised, it certifies something that does not happen. That is the same + hazard the existing design names for the full scenario, and the same answer applies — + nothing new drives it, and what runs is the real thing. + +## References + +- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine. +- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the + observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned + bundle, which is also exactly the bootstrap path. +- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this + reorders. diff --git a/02-DECISIONS/0030-the-repository-structure.md b/02-DECISIONS/0030-the-repository-structure.md new file mode 100644 index 0000000..4dca3ae --- /dev/null +++ b/02-DECISIONS/0030-the-repository-structure.md @@ -0,0 +1,99 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +--- + +# 30. The repository structure, and the rule that names them + +## Context + +The tiers are settled ([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)) +and the product is named ([ADR 0027](0027-the-product-is-novox-mesh.md)), but the repositories +themselves were only ever sketched in research. Two consequences had already appeared. + +[ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) makes the lab phase 0 of the entire +migration and **could not say where it lives**, because no record named a repository. + +And the research contradicted an accepted record: it listed `mesh-hq` for this repository, while +[ADR 0028](0028-hq-is-company-scoped.md) had decided `novox/hq` and explicitly rejected that +name. A design resting on research is a design resting on something that can change without a +decision. + +There is also an implied naming rule that has never been written down. ADR 0027 says +repository names take `mesh`; ADR 0028 gives this repository no prefix at all. Both are right, +for a reason neither states. + +## Considered options + +Only the naming rule had genuine alternatives; the tier repositories follow from the tiers. + +1. **No prefix — `novox/host`, `novox/control`.** The organisation already says Novox, so the + prefix reads as stutter. Rejected once it was established that Novox delivers more than the + mesh: with several products the prefix is not stutter, it is the product namespace doing + real work, and the forge has no nested groups to do it instead. +2. **An organisation per product — `novox-mesh/host`.** Puts the product boundary where the + forge's only real grouping primitive lives, so permissions and teams attach to it. Rejected + for now as premature: no per-product access boundary exists yet, and it costs `novox-` + repeated across every organisation. +3. **Product-prefixed repositories in the company organisation.** Chosen. + +## Decision + +**The naming rule:** a repository that belongs to a product carries that product's prefix. A +repository that is company-scoped does not. + +That is why this one is `hq` and the mesh's are `mesh-*`. Both records were already correct; +the rule connecting them is stated here. + +**The repositories:** + +| Repository | Tier | Holds | +|---|---|---| +| `novox/mesh-host` | 0 | the node host — the one binary installed by hand | +| `novox/mesh-substrate` | 1 | the four pinned services, as declarations | +| `novox/mesh-control` | 2 | the control plane and its contexts | +| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic | +| `novox/mesh-sdk` | — | contracts shared across tiers: types, not behaviour | +| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement | +| `novox/hq` | — | this repository. Company-scoped ([ADR 0028](0028-hq-is-company-scoped.md)) | + +**The lab is its own repository.** Its lifecycle differs from everything else in the list: it +is never shipped to a node, it outlives any single tier, and it drives virtualisation on a +workstation — which nothing else in the mesh does. Putting it inside the host would couple +development tooling to a shipped component; putting it inside the control plane would make the +bootstrap scenario depend on a tier that does not exist when it is needed. + +**Tier 4 is deliberately not decided here.** Whether the catalogue is one repository, one per +domain, or one per application remains open from +[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) and is blocked on +[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md): how many repositories hold +domains cannot be answered before what the domains are. Recording the gap is the point — +`mesh-catalog` appears in the research sketch and is **not** decided by this record. + +## Consequences + +- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) can name its target. Phase 0 has + a home, which was the immediate blocker. +- The research sketch stops being load-bearing. It remains what it is — a sketch — and the + design layer can now cite a record instead. +- **Seven repositories where there is currently one**, for a mesh that today lives in a single + monorepo. That is the cost, and it is not small: seven release cadences, seven sets of + dependencies, and cross-repository changes that were previously one commit. The offsetting + argument is the tier rule — a boundary that only points downward is enforceable across + repositories and merely conventional inside one. +- The prefix will read as redundant for as long as the mesh is the only product with + repositories. That is accepted deliberately: the alternative is renaming everything at the + moment a second product appears, which is the class of migration this project is trying to + stop performing. +- Nothing is created yet. This records what the repositories *are*; creating them is part of + phase 0 and after. + +## References + +- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name the prefix comes from. +- [ADR 0028](0028-hq-is-company-scoped.md) — why this repository has no prefix. +- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lab, and why it is first. +- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the + sketch this supersedes as a source. diff --git a/03-DESIGN/01-to-be/01-end-to-end-testing.md b/03-DESIGN/01-to-be/01-end-to-end-testing.md index f8cdad1..b550a47 100644 --- a/03-DESIGN/01-to-be/01-end-to-end-testing.md +++ b/03-DESIGN/01-to-be/01-end-to-end-testing.md @@ -30,6 +30,36 @@ its question needs. If something can only be tested against the shape the mesh h today, that is a gap in the vocabulary rather than a reason to privilege that shape. +## Two classes of scenario + +The design below describes a scenario as a complete mesh — forge, coordinator, delivery cascade +— because what it tests is a module. **That is the larger of two classes, and not the first one +built** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +| | **Bootstrap scenario** | **Full scenario** | +|---|---|---| +| Contains | virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules | +| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify | +| Exercises | the node host and the substrate | the control plane and everything above it | +| Exists to | **develop the mesh** | **test what runs on it** | + +The bootstrap scenario is a **strict subset**: same virtualisation, same networking, same +lifecycle — it simply stops before a control plane exists. Everything from *"Where this sits in +the way work happens"* onward describes the full scenario, and applies once there is a +coordinator to describe. + +**The bootstrap scenario is built first, ahead of everything it will later test.** The +component that takes over a machine's packages, services and network cannot be developed +against a machine anyone needs, and raising a node from nothing is today the least-exercised +path in the system precisely because it only ever runs for real. Making it the inner +development loop inverts that. + +Of the runner's two jobs, this settles their order: **scenario lifecycle first** — something +must materialise, snapshot and destroy a mesh before anything else can be written — and +**assertion execution second**, with the full scenario, since a bootstrap scenario's assertions +concern the state one host reconciled and are small enough to state directly. + + ## Where this sits in the way work happens Work reaches the mesh along one path today: @@ -93,9 +123,11 @@ drifts. nodes that have the module assigned. It needs to be able to run the same pipeline against a mesh named by the request instead. - **Scenarios must be concurrent and cheap.** Several agents working means several scenarios - at once, each needing its own network and nodes. This is affordable with system containers - and would not be with virtual machines — the unit choice is what makes the gate possible - at all. + at once, each needing its own network and nodes. A lab node is a virtual machine + ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)), and snapshots are + what make repetition cheap — restoring a scenario costs far less than building one. The + earlier argument here, that only system containers made this affordable, was superseded: the + scale it assumed was invented rather than required. - **The gate is only as good as the verification behind it.** A module with no assertions gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes the number that matters**, and it starts at approximately zero. diff --git a/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md b/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md new file mode 100644 index 0000000..646bb1d --- /dev/null +++ b/04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md @@ -0,0 +1,117 @@ +--- +status: diagnosing +opened: 2026-08-23 +located-in: [hal] +fixed-by: + - "the instance only: PR #962 — incus hook. Merged, ran on one node in 6s of a 60s budget; pipeline #6832 green in 48s. The class remains open." +amended-design: +--- + +# 007 — An installed package is not an available capability + +## Symptom + +A module declares the virtualisation package the lab needs. The package is installed — +version 7.3.0-1, recorded as explicitly installed. The client binary runs. + +The capability does not exist: + +| Checked | State | +|---|---| +| `incus.service` | disabled, inactive | +| `incus.socket` | disabled, inactive | +| `incus-user.socket` | disabled, inactive | +| the operator's group membership | not a member of any incus group | +| the client | reports **`Server version: unreachable`** | + +Nothing failed. Nothing reported anything. The declaration was satisfied exactly as written, +and the thing it was declared for cannot be used. + +## Why this is not issue 001 again + +[Issue 001](../001-failed-package-install-reports-success/00-report.md) is *the install failed +and the job reported success*. This is the opposite and arguably worse: **the install +succeeded, and success was not the point.** + +A package is a set of files. A capability is a running service, an enabled socket, and an +identity permitted to reach it. The module model declares the first and has no vocabulary for +the second, so the gap between them is invisible — there is no state in which the mesh believes +this node has virtualisation and is wrong, because the mesh was never asked to believe it. + +The distance between the two is the same one the delivery layer already has a name for: +**transport versus effect.** A package install reports that files arrived, which is transport. + +## Why it matters now + +This is the first requirement of the lab +([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)), which is +phase 0 of the entire migration. The first capability the new work depends on is present, +declared, and unusable — and would have stayed unusable silently. + +It also generalises. Every module that declares a package needing a unit enabled, a group +joined, a kernel module loaded, or a socket activated has this gap. Post-install work lives in +hooks, and hooks have their own recorded failure mode: thirteen were found that had never run. + +## Evidence + +- Package recorded as installed 2026-08-23 00:02, explicitly. +- Both units disabled and inactive; the operator in no incus group; client reports the server + unreachable. + +## Open questions + +- Should a module be able to declare a **capability** — a unit that must be enabled, a group + the operator must be in — rather than only the package that provides it? +- If that is what hooks are for, why is the gap invisible when a hook does not run? A hook that + never fires and a hook that fires and does nothing are indistinguishable today. +- Is this what the verify stage should be asserting? It exists, and a module's own assertions + are meant to test outcomes rather than steps — "the socket accepts a connection" is exactly + that shape. +- How many other declared packages are in this state? Nothing currently reports it, which means + the answer is unknown rather than zero. + +## The instance is fixed; the class is what this issue is now about + +*Updated 2026-08-23.* A hook now does the post-install work — group membership, subordinate id +ranges, enabling both units, creating the storage pool, the bridge, and the default profile. +Verified independently afterwards: the group exists with the operator in it, both id files +carry the range, the service is active, and the storage pool reports `CREATED`. + +**So the mechanism was never missing.** Hooks are exactly the right place for this, and used +properly they work. The gap is narrower and worse than "there is no way to do it": + +> The hook did six things. **Six checks were then performed by a human, by hand.** Nothing in +> the pipeline asserted any of them, and a pipeline that dispatched a hook which silently never +> fired would have been green in the same 48 seconds. + +That is not hypothetical — a hook named for a feature its module does not carry is skipped +without complaint, and thirteen such hooks were found at once in the past. The distance between +*the hook ran and did six things* and *the hook was dispatched* is invisible from the outside, +and it is the whole of this issue. + +### The verification that was done by hand is the assertion set + +The six checks performed after the merge are, almost word for word, what the module's own +verification should assert — outcomes, not steps, exactly as the lab design requires: + +| Asserted by hand | As a module assertion | +|---|---| +| operator in the admin group | the group exists and contains the operator | +| subordinate id range in both files | both files carry the range | +| both units enabled, service active | the service is active, and still active shortly after | +| storage pool present | the pool exists and reports created | +| bridge present with an address | the bridge exists and holds its address | +| default profile wired to both | the profile references the bridge and the pool | + +They currently live in a chat message. Moved into the module, they would run on every delivery +to every node, and the difference between a hook that worked and a hook that was merely +dispatched would stop being something a person has to notice. + +### What this issue now asks + +- Does a module gain a way to declare a **capability** — the outcome — separately from the + package that provides it, or is "write assertions" the whole answer? +- The verify stage exists and asserts almost nothing. Is this simply the first module that + should use it properly, making the issue a coverage problem rather than a design gap? +- **How many other declared packages are in the state this one was in?** Still unknown, still + unreported by anything, and now demonstrably worth asking.