Skills take the hq- prefix: they are HQ process workflows, not mesh workflows, and HQ is company-scoped now. hq-new-research, hq-graduate, hq-new-issue, hq-diagnose, hq-amend-design, hq-handoff, hq-sync-constitution, hq-status. Forward-looking prose becomes Novox Mesh or simply the mesh — the root README, AGENTS.md, the 00-META README, the mission's module example, and one to-be document that addressed 'someone working on HAL'. Three categories deliberately keep HAL, per ADR 0027: The monorepo is still called hal on the forge. repos.md, every code: field and every located-in: field name a repository that exists under that name, and renaming them in prose would make them false. The as-is layer and the research that measured it describe the system that runs, and that system is called HAL. 124 modules, 9 daemons, a dead containerised node — those are observations, not intentions. Records 0001-0026 are immutable. A record says what was decided when it was decided, and no record is edited for a name. Also repoints ADR 0022's link at the renamed skill — a path fix, which the immutability rule permits, not a change of meaning.
382 lines
16 KiB
Markdown
382 lines
16 KiB
Markdown
---
|
||
layer: to-be
|
||
status: designed
|
||
code: [hal]
|
||
updated: 2026-08-23
|
||
decisions: [02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
|
||
---
|
||
|
||
# End-to-end testing
|
||
|
||
**What is under test is a module.** The mesh is the harness.
|
||
|
||
You change a module or write a new one, run it end to end, and get a verdict before it goes
|
||
anywhere near production. That loop is the product of this design; everything else exists to
|
||
make it fast and honest.
|
||
|
||
This is the Phase 0 prerequisite from [`00-work-breakdown.md`](00-work-breakdown.md).
|
||
|
||
---
|
||
|
||
## Two boundaries this design sets
|
||
|
||
**Adoption is out of scope.** Bringing a node into being is the lab runner's job. Adopting a
|
||
machine that already exists, with its own configuration, is a legacy path and the lab does not
|
||
reproduce it — a scenario starts from nothing every time, which is what makes it a fixture
|
||
rather than a snapshot.
|
||
|
||
**The real topology is one test among many, not the baseline.** A scenario declares the mesh
|
||
its question needs. If something can only be tested against the shape the mesh happens to have
|
||
today, that is a gap in the vocabulary rather than a reason to privilege that shape.
|
||
|
||
|
||
## Where this sits in the way work happens
|
||
|
||
Work reaches the mesh along one path today:
|
||
|
||
```
|
||
a change is made a human in a session, or an agent given work
|
||
│
|
||
▼
|
||
a pull request appears
|
||
│
|
||
▼
|
||
a human reads the diff and merges ← the gate
|
||
│
|
||
▼
|
||
the coordinator delivers to the real nodes
|
||
│
|
||
▼
|
||
production reports whether it worked ← the test
|
||
```
|
||
|
||
**The gate is a human reading a diff, and the test is production.** That is workable at a
|
||
change a day and it is the constraint at ten. For autonomous work it is worse than a
|
||
constraint: an agent's output arrives as a diff that *looks* right, carrying no evidence
|
||
that it runs, and the only reviewer is the condition `00-META/context.md` calls mandatory
|
||
— *human agents are few, often one, and usually asleep.*
|
||
|
||
The missing step goes between the pull request and the merge:
|
||
|
||
```
|
||
a pull request appears
|
||
│
|
||
▼
|
||
the coordinator delivers the branch to a SCENARIO mesh ← the missing step
|
||
and runs the same stages, ending in verify
|
||
│
|
||
▼
|
||
the verdict is attached to the pull request
|
||
│
|
||
▼
|
||
a human merges evidence rather than hope
|
||
│
|
||
▼
|
||
the coordinator delivers to the real nodes — same verify, now loud
|
||
```
|
||
|
||
### One pipeline, two targets
|
||
|
||
| target | triggered by | what a failure means |
|
||
|---|---|---|
|
||
| a scenario mesh | a branch, or a pull request | the change is not finished; it should not merge |
|
||
| the real mesh | a merge | a red delivery, loudly, before anything is built on it |
|
||
|
||
Same coordinator, same cascade, same stages, same verification. **Only the target mesh
|
||
differs.** This is the through-line of the whole design — one verification statement with two
|
||
jobs, one runner with two callers, one pipeline with two targets. Nothing forks, so nothing
|
||
drifts.
|
||
|
||
### What it demands
|
||
|
||
- **The coordinator must accept a target mesh.** Today a pipeline's targets are derived: the
|
||
nodes that have the module assigned. It needs to be able to run the same pipeline against
|
||
a mesh named by the request instead.
|
||
- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios
|
||
at once, each needing its own network and nodes. This is affordable with system containers
|
||
and would not be with virtual machines — the unit choice is what makes the gate possible
|
||
at all.
|
||
- **The gate is only as good as the verification behind it.** A module with no assertions
|
||
gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes
|
||
the number that matters**, and it starts at approximately zero.
|
||
- **It has to be fast enough to wait for.** A gate an agent cannot wait on is a report nobody
|
||
reads.
|
||
|
||
---
|
||
|
||
## The coordinator drives it
|
||
|
||
The temptation is to build a framework that delivers a module and checks it. That would be a
|
||
**second delivery path**, and a second delivery path is worthless — the faults worth catching
|
||
live in the real one.
|
||
|
||
So the rule is not "no new components". It is: **nothing new drives delivery.** A scenario is
|
||
a complete mesh with its own coordinator. Push the working tree to that mesh's forge; its
|
||
coordinator does exactly what a coordinator does — works out the cascade, dispatches build,
|
||
install, configure, start, and then **verify** — and its meshware executes on its nodes. The
|
||
result of that pipeline *is* the verdict.
|
||
|
||
A **test runner** is a legitimate component within that, used *by* the coordinator rather
|
||
than instead of it. Two jobs plausibly belong to it, and their boundary is worth settling
|
||
before either is built:
|
||
|
||
- **scenario lifecycle** — materialise the mesh a test needs, restore it to a snapshot, tear
|
||
it down. Something must do this before a coordinator exists to drive anything.
|
||
- **assertion execution** — give verification more than "run a script and check the exit
|
||
code": setup and teardown, timeouts, retry-until-true for things that settle, and results
|
||
structured enough to report rather than grep.
|
||
|
||
The line to hold is the pipeline itself. A runner that stands up a mesh and executes
|
||
assertions is a component. A runner that decides what to build, in what order, and ships it
|
||
to a node is a fork of the coordinator.
|
||
|
||
### It has two callers, and they want different things
|
||
|
||
The runner serves **the coordinator** and **a person developing the mesh**, and its interface
|
||
has to suit both:
|
||
|
||
| caller | wants |
|
||
|---|---|
|
||
| the coordinator | non-interactive, structured results it can record against a pipeline, a clean teardown, no prompts and no colour |
|
||
| someone working on the mesh | readable output, the failing mesh **left standing** to open a shell into, and a way to re-run one assertion without repeating the whole delivery |
|
||
|
||
Hence at least two verbs: one that runs to a verdict and tears down, and one that stands a
|
||
scenario up and leaves it there. The second is how a developer works *inside* a mesh —
|
||
which is the thing the current host-borrowing tooling is really for, and the reason this
|
||
replaces it rather than sitting beside it.
|
||
|
||
That the same runner serves both is deliberate, and it is the same argument as the module's
|
||
own assertions serving both development and delivery: **one statement, two jobs.** Anything
|
||
that only the developer path can do is a divergence, and it will drift.
|
||
|
||
```
|
||
edit a module in the working tree
|
||
│
|
||
▼
|
||
push to the scenario's forge
|
||
│
|
||
▼
|
||
the scenario's COORDINATOR runs a pipeline ← existing machinery, unchanged
|
||
│
|
||
├── cascade: which modules are affected
|
||
├── build → install → configure → start
|
||
└── verify: the module's own assertions ← existing stage, dispatched today
|
||
│
|
||
▼
|
||
the pipeline result is the verdict
|
||
```
|
||
|
||
This is the same property `00-META/mission.md` asks for: *the mesh's own components ship
|
||
through the same machinery as anything else it carries — if they need an exception, the
|
||
machinery is not finished.* A test that needed its own delivery path would be that exception.
|
||
|
||
**It tests the working tree**, because the forge is inside the scenario. A loop that requires
|
||
pushing to production and waiting is not a loop. Real path, local code, nothing shared with
|
||
production.
|
||
|
||
---
|
||
|
||
## What it catches
|
||
|
||
This list is the specification. These are the ways a module change fails today, and each one
|
||
currently reaches production or wastes a pipeline run:
|
||
|
||
| failure | why it survives today |
|
||
|---|---|
|
||
| a file never reached the artifact | absence and "declares nothing here" are indistinguishable, so it ships green |
|
||
| a migration compiled to nothing, or never ran | the stage reports success when there is nothing to run |
|
||
| the manifest is wrong — bad package list, wrong paths | validated shallowly, if at all |
|
||
| a capability was never provisioned, or its credential never arrived | delivery reports transport, not effect |
|
||
| an environment value was not generated | the module starts and reads a default |
|
||
| the service did not come up, or came up and crashed | nothing asserts it is still running a minute later |
|
||
| the dependency cascade did not include the module | a green pipeline that rebuilt the wrong set |
|
||
| it works on a fresh install but breaks on upgrade | almost never exercised — see below |
|
||
| it works on one node and not another | only one node is ever tried |
|
||
|
||
A test that only proves "the pipeline went green" reproduces the exact blindness this is
|
||
meant to remove.
|
||
|
||
---
|
||
|
||
## Fresh install and upgrade are different tests
|
||
|
||
The most common shape of a module bug is: works from scratch, breaks on the machine that
|
||
already had the previous version. Existing state outranks new state, a file is added but
|
||
never removed, a migration assumes a column that an older node lacks.
|
||
|
||
Snapshots make both cheap, so both are default:
|
||
|
||
- **fresh** — restore a mesh that has never seen the module, deliver, assert
|
||
- **upgrade** — restore a mesh running the *previous released version*, deliver the working
|
||
tree over it, assert
|
||
|
||
Same assertions, different starting state. A module that passes one and fails the other is
|
||
the normal case, not an edge case.
|
||
|
||
---
|
||
|
||
## A module carries its own assertions
|
||
|
||
**A module states what must be true about it, and it states it once.** That statement is the
|
||
module's verification — the stage the coordinator already dispatches at the end of every
|
||
delivery, with a working handler, implemented today by essentially nothing.
|
||
|
||
It asserts **outcomes**, never that a step ran: the unit is active and still active shortly
|
||
after, the schema has the column, the name resolves, the credential authenticates, the file
|
||
on the node holds what the mesh believes it holds, the endpoint answers.
|
||
|
||
The same statement serves both places, which is the point:
|
||
|
||
| where it runs | what a failure means |
|
||
|---|---|
|
||
| in a scenario, during development | your change is not finished — cheap, fast, nobody affected |
|
||
| on delivery to production | the deploy is red, loudly, before anyone builds on it |
|
||
|
||
**This is what finally makes verification worth writing.** Today it can only ever cost you a
|
||
deploy, which is precisely why almost no module has one. Give it a second job — telling a
|
||
developer whether their change works — and writing it stops being an act of discipline and
|
||
starts being the fastest way to get an answer.
|
||
|
||
Where a module has no verification yet, the coordinator asserts the generic invariants it can
|
||
know on its own: the artifact contained what the manifest declared, the migrations that were
|
||
pending ran, the declared capabilities were provisioned, the declared services are up.
|
||
|
||
---
|
||
|
||
## The mesh shape is a parameter
|
||
|
||
A test declares the mesh it needs, and the default is the smallest one that can exercise the
|
||
module:
|
||
|
||
```yaml
|
||
module: a-web-service
|
||
mesh:
|
||
my-cool-node: { role: published, publishes: my-cool-node.com }
|
||
assert:
|
||
- https://my-cool-node.com answers 200
|
||
- the certificate presented is valid for that name
|
||
```
|
||
|
||
One node, because one node is enough to answer that question. Standing up a mesh to test one
|
||
module is the same mistake as starting the application to test a function.
|
||
|
||
More nodes when the module's behaviour is *between* nodes:
|
||
|
||
```yaml
|
||
module: a-module-requiring-a-database
|
||
mesh:
|
||
store: { role: anchor }
|
||
consumer-a: { role: resident }
|
||
consumer-b: { role: mobile }
|
||
assert:
|
||
- both consumers authenticate against the database
|
||
- after redeploying the provider, both still do
|
||
```
|
||
|
||
Node names and domains in a test are **invented**. The mesh under test is whatever the test
|
||
says it is — which is also how this document stays free of any particular installation.
|
||
|
||
### Scale, when scale is the question
|
||
|
||
Size is chosen by what is being asked, and ranges from one node to twenty or more.
|
||
|
||
| size | what only this size answers |
|
||
|---|---|
|
||
| one | does it install, migrate, provision and run at all — the fastest loop |
|
||
| two–three | anything *between* nodes: delivery, provisioning, rotation, absence |
|
||
| ten–twenty | whether a fan-out reaches *every* node, whether the cascade converges, whether something is quietly quadratic |
|
||
|
||
A fan-out reaching three of four nodes reads as a flake; at twenty it is a diagnosis, and
|
||
which nodes were missed tells you why. **This is what the unit choice below bought** — twenty
|
||
virtual machines do not fit on a workstation, and twenty system containers do.
|
||
|
||
---
|
||
|
||
## Two kinds of test
|
||
|
||
**Module tests** are the daily case and the reason this exists: does my change work, end to
|
||
end.
|
||
|
||
**Mesh tests** use the same machinery to ask whether the mesh itself behaves — that a
|
||
credential rotation reaches every consumer, that delivery to an absent node is reported as
|
||
pending rather than done, that a returning node catches up. These are fewer and change
|
||
rarely, but they are where the known production faults get encoded so they stay fixed.
|
||
|
||
The known faults become mesh tests that fail today. That is the Phase 0 checkpoint.
|
||
|
||
---
|
||
|
||
## The substrate
|
||
|
||
### A node is a system container
|
||
|
||
An OS userspace with its own init, its own network interface, its own filesystem, sharing
|
||
the host kernel.
|
||
|
||
**A node's job is to run containers**, so modelling a node *as* an application container
|
||
inverts the thing being modelled: it forces nested containers through a privileged daemon or
|
||
a shared socket, and a shared socket makes isolation between nodes cosmetic. A system
|
||
container has no such problem — init runs as PID 1, so units and timers work as written;
|
||
containers nest properly, so module stacks run as they do anywhere. **Module code and mesh
|
||
code both run unmodified**, which is the property that makes a verdict trustworthy. Anything
|
||
needing a special case locally is a divergence that will hide a fault.
|
||
|
||
Boot is around a second and snapshots are cheap, which is what makes this an inner loop
|
||
rather than an errand.
|
||
|
||
**Any node can be a full virtual machine instead**, through the same tooling and the same
|
||
test file — when a question needs a kernel to answer it, or when the point is that nodes are
|
||
*not* identical.
|
||
|
||
### The network
|
||
|
||
Two segments and an overlay, because some module behaviour is only visible across a real
|
||
network boundary:
|
||
|
||
- **wan** — a published node holds an address here, and an authoritative resolver maps its
|
||
name to it, so a public touchpoint is real enough to exercise routing, virtual hosts and
|
||
certificates
|
||
- **local** — behind translation, as a home network is
|
||
- **the overlay** — the mechanism production uses; the mesh addresses peers by mesh name and
|
||
never learns which segment anyone is on
|
||
|
||
A node can be moved between segments or detached entirely, mid-test.
|
||
|
||
### What is not real
|
||
|
||
- **the model provider** — thinking is stubbed, so tests cost nothing to run
|
||
- **the public internet** — a bridge, with an authoritative resolver rather than delegation
|
||
- **the public certificate authority** — the lab runs **its own ACME issuer** on the public
|
||
segment, so issuance, challenge and renewal are genuinely exercised rather than stubbed.
|
||
Internal names keep the mesh CA, so the lab preserves production's two-authority split
|
||
rather than collapsing it into one.
|
||
|
||
Everything a node itself does is real, because a node is a real machine.
|
||
|
||
---
|
||
|
||
## Consequences
|
||
|
||
**Bringing a node into being is part of the framework.** A test creates its own nodes — one
|
||
for the smallest, twenty for the largest — repeatably, unattended, and cheaply enough to do
|
||
it twenty times in a row. Those are the constraints a real mesh wants, so the mechanism the
|
||
runner needs is the one the mesh should keep.
|
||
|
||
**Module verification becomes worth writing**, because it is the thing that gives a
|
||
developer a verdict, not just a stricter deploy.
|
||
|
||
**Host-borrowing ends.** Today's tooling starts providers on the host's own init system and
|
||
reads credentials from host paths, because there is nowhere else to put a mesh. Once there
|
||
is, a workstation stops being collateral.
|
||
|
||
**The supervision question stops gating anything.**
|
||
[`01-RESEARCH/003-service-supervision`](../../01-RESEARCH/003-service-supervision/analysis.md)
|
||
remains open on its own merits — and once this exists, its options are cheap to try rather
|
||
than expensive to argue about.
|
||
|
||
## Deliberately not decided
|
||
|
||
**Whether the lab verdict is a workflow guard or an advisory check on the pull request.** Open,
|
||
and deliberately trivial — a policy detail, changeable in an afternoon, not an architectural
|
||
choice. Recorded so it is not mistaken for an oversight.
|