Files
hq/03-DESIGN/01-to-be/01-end-to-end-testing.md
T
jschoubben 5dcb9dfbf6 ADR 0089: a bed reads the catalogue it proves; issue 073 diagnosed; issues 074 and 075 opened
The end-to-end design held 'the run rebuilds what it tests' for binaries and images
and not for manifests. The beds' inline copies fell into three kinds; only the first
is a stale copy. The other two are named: a mesh test wearing a catalogue module's
name (074) and a stocked runtime image the run never rebuilds (075).
2026-09-21 14:35:08 +02:00

22 KiB
Raw Blame History

layer, status, code, updated, decisions
layer status code updated decisions
to-be in-progress
mesh-lab
2026-09-21
02-DECISIONS/0016-the-lab.md
02-DECISIONS/0019-how-this-repository-works.md
02-DECISIONS/0089-a-bed-reads-the-catalogue-it-proves.md

End-to-end testing

What is under test is a module. The mesh is the harness.

You change a module or write a new one, run it end to end, and get a verdict before it goes anywhere near production. That loop is the product of this design; everything else exists to make it fast and honest.

This is the Phase 0 prerequisite from 00-work-breakdown.md.


Two boundaries this design sets

Adoption is out of scope. Bringing a node into being is the lab runner's job. Adopting a machine that already exists, with its own configuration, is a legacy path and the lab does not reproduce it — a scenario starts from nothing every time, which is what makes it a fixture rather than a snapshot.

The real topology is one test among many, not the baseline. A scenario declares the mesh its question needs. If something can only be tested against the shape the mesh happens to have today, that is a gap in the vocabulary rather than a reason to privilege that shape.

Two classes of scenario

The design below describes a scenario as a complete mesh — forge (Gitea), coordinator, delivery cascade — because what it tests is a module. That is the larger of two classes, and not the first one built (ADR 0016).

Bootstrap scenario Full scenario
Contains virtual machines, the host binary, a pinned foundation bundle a complete mesh: forge, coordinator, delivery, modules
Verdict from what the host reports about the state it reconciled a pipeline result ending in verify
Exercises the node host and the foundation the controller and everything above it
Exists to develop the mesh test what runs on it

The bootstrap scenario is a strict subset: same virtualisation, same networking, same lifecycle — it simply stops before a controller exists. Everything from "Where this sits in the way work happens" onward describes the full scenario, and applies once there is a coordinator to describe.

The bootstrap scenario is built first, ahead of everything it will later test. The component that takes over a machine's packages, services and network cannot be developed against a machine anyone needs, and raising a node from nothing is today the least-exercised path in the system precisely because it only ever runs for real. Making it the inner development loop inverts that.

Of the runner's two jobs, this settles their order: scenario lifecycle first — something must materialise, snapshot and destroy a mesh before anything else can be written — and assertion execution second, with the full scenario, since a bootstrap scenario's assertions concern the state one host reconciled and are small enough to state directly.

Where this sits in the way work happens

Work reaches the mesh along one path today:

  a change is made                    a human in a session, or an agent given work
        │
        ▼
  a pull request appears
        │
        ▼
  a human reads the diff and merges   ← the gate
        │
        ▼
  the coordinator delivers to the real nodes
        │
        ▼
  production reports whether it worked   ← the test

The gate is a human reading a diff, and the test is production. That is workable at a change a day and it is the constraint at ten. For autonomous work it is worse than a constraint: an agent's output arrives as a diff that looks right, carrying no evidence that it runs, and the only reviewer is the condition 00-META/context.md calls mandatory — human agents are few, often one, and usually asleep.

The missing step goes between the pull request and the merge:

  a pull request appears
        │
        ▼
  the coordinator delivers the branch to a SCENARIO mesh       ← the missing step
  and runs the same stages, ending in verify
        │
        ▼
  the verdict is attached to the pull request
        │
        ▼
  a human merges evidence rather than hope
        │
        ▼
  the coordinator delivers to the real nodes — same verify, now loud

One pipeline, two targets

target triggered by what a failure means
a scenario mesh a branch, or a pull request the change is not finished; it should not merge
the real mesh a merge a red delivery, loudly, before anything is built on it

Same coordinator, same cascade, same stages, same verification. Only the target mesh differs. This is the through-line of the whole design — one verification statement with two jobs, one runner with two callers, one pipeline with two targets. Nothing forks, so nothing drifts.

What it demands

  • The coordinator must accept a target mesh. Today a pipeline's targets are derived: the nodes that have the module assigned. It needs to be able to run the same pipeline against a mesh named by the request instead.
  • Scenarios must be concurrent and cheap. Several agents working means several scenarios at once, each needing its own network and nodes. A lab node is a virtual machine (ADR 0016), and snapshots are what make repetition cheap — restoring a scenario costs far less than building one. The earlier argument here, that only system containers made this affordable, was superseded: the scale it assumed was invented rather than required.
  • The gate is only as good as the verification behind it. A module with no assertions gets a weak gate: delivery succeeded, nothing checked. So verification coverage becomes the number that matters, and it starts at approximately zero.
  • It has to be fast enough to wait for. A gate an agent cannot wait on is a report nobody reads.

The coordinator drives it

The temptation is to build a framework that delivers a module and checks it. That would be a second delivery path, and a second delivery path is worthless — the faults worth catching live in the real one.

So the rule is not "no new components". It is: nothing new drives delivery. A scenario is a complete mesh with its own coordinator. Push the working tree to that mesh's forge; its coordinator does exactly what a coordinator does — works out the cascade, dispatches build, install, configure, start, and then verify — and its meshware executes on its nodes. The result of that pipeline is the verdict.

A test runner is a legitimate component within that, used by the coordinator rather than instead of it. Two jobs plausibly belong to it, and their boundary is worth settling before either is built:

  • scenario lifecycle — materialise the mesh a test needs, restore it to a snapshot, tear it down. Something must do this before a coordinator exists to drive anything.
  • assertion execution — give verification more than "run a script and check the exit code": setup and teardown, timeouts, retry-until-true for things that settle, and results structured enough to report rather than grep.

The line to hold is the pipeline itself. A runner that stands up a mesh and executes assertions is a component. A runner that decides what to build, in what order, and ships it to a node is a fork of the coordinator.

It has two callers, and they want different things

The runner serves the coordinator and a person developing the mesh, and its interface has to suit both:

caller wants
the coordinator non-interactive, structured results it can record against a pipeline, a clean teardown, no prompts and no colour
someone working on the mesh readable output, the failing mesh left standing to open a shell into, and a way to re-run one assertion without repeating the whole delivery

Hence at least two verbs: one that runs to a verdict and tears down, and one that stands a scenario up and leaves it there. The second is how a developer works inside a mesh — which is the thing the current host-borrowing tooling is really for, and the reason this replaces it rather than sitting beside it.

That the same runner serves both is deliberate, and it is the same argument as the module's own assertions serving both development and delivery: one statement, two jobs. Anything that only the developer path can do is a divergence, and it will drift.

  edit a module in the working tree
        │
        ▼
  push to the scenario's forge
        │
        ▼
  the scenario's COORDINATOR runs a pipeline          ← existing machinery, unchanged
        │
        ├── cascade: which modules are affected
        ├── build → install → configure → start
        └── verify: the module's own assertions        ← existing stage, dispatched today
        │
        ▼
  the pipeline result is the verdict

This is the same property 00-META/mission.md asks for: the mesh's own components ship through the same machinery as anything else it carries — if they need an exception, the machinery is not finished. A test that needed its own delivery path would be that exception.

It tests the working tree, because the forge is inside the scenario. A loop that requires pushing to production and waiting is not a loop. Real path, local code, nothing shared with production.


What it catches

This list is the specification. These are the ways a module change fails today, and each one currently reaches production or wastes a pipeline run:

failure why it survives today
a file never reached the artifact absence and "declares nothing here" are indistinguishable, so it ships green
a migration compiled to nothing, or never ran the stage reports success when there is nothing to run
the manifest is wrong — bad package list, wrong paths validated shallowly, if at all
a capability was never provisioned, or its credential never arrived delivery reports transport, not effect
an environment value was not generated the module starts and reads a default
the service did not come up, or came up and crashed nothing asserts it is still running a minute later
the dependency cascade did not include the module a green pipeline that rebuilt the wrong set
it works on a fresh install but breaks on upgrade almost never exercised — see below
it works on one node and not another only one node is ever tried

A test that only proves "the pipeline went green" reproduces the exact blindness this is meant to remove.


Fresh install and upgrade are different tests

The most common shape of a module bug is: works from scratch, breaks on the machine that already had the previous version. Existing state outranks new state, a file is added but never removed, a migration assumes a column that an older node lacks.

Snapshots make both cheap, so both are default:

  • fresh — restore a mesh that has never seen the module, deliver, assert
  • upgrade — restore a mesh running the previous released version, deliver the working tree over it, assert

Same assertions, different starting state. A module that passes one and fails the other is the normal case, not an edge case.


A module carries its own assertions

A module states what must be true about it, and it states it once. That statement is the module's verification — the stage the coordinator already dispatches at the end of every delivery, with a working handler, implemented today by essentially nothing.

It asserts outcomes, never that a step ran: the unit is active and still active shortly after, the schema has the column, the name resolves, the credential authenticates, the file on the node holds what the mesh believes it holds, the endpoint answers.

The same statement serves both places, which is the point:

where it runs what a failure means
in a scenario, during development your change is not finished — cheap, fast, nobody affected
on delivery to production the deploy is red, loudly, before anyone builds on it

This is what finally makes verification worth writing. Today it can only ever cost you a deploy, which is precisely why almost no module has one. Give it a second job — telling a developer whether their change works — and writing it stops being an act of discipline and starts being the fastest way to get an answer.

Where a module has no verification yet, the coordinator asserts the generic invariants it can know on its own: the artifact contained what the manifest declared, the migrations that were pending ran, the declared capabilities were provisioned, the declared services are up.


The mesh shape is a parameter

A test declares the mesh it needs, and the default is the smallest one that can exercise the module:

module: a-web-service
mesh:
  my-cool-node: { role: published, publishes: my-cool-node.com }
assert:
  - https://my-cool-node.com answers 200
  - the certificate presented is valid for that name

One node, because one node is enough to answer that question. Standing up a mesh to test one module is the same mistake as starting the application to test a function.

More nodes when the module's behaviour is between nodes:

module: a-module-requiring-a-database
mesh:
  store:      { role: anchor }
  consumer-a: { role: resident }
  consumer-b: { role: mobile }
assert:
  - both consumers authenticate against the database
  - after redeploying the provider, both still do

Node names and domains in a test are invented. The mesh under test is whatever the test says it is — which is also how this document stays free of any particular installation.

Scale, when scale is the question

Size is chosen by what is being asked, and ranges from one node to twenty or more.

size what only this size answers
one does it install, migrate, provision and run at all — the fastest loop
two–three anything between nodes: delivery, provisioning, rotation, absence
ten–twenty whether a fan-out reaches every node, whether the cascade converges, whether something is quietly quadratic

A fan-out reaching three of four nodes reads as a flake; at twenty it is a diagnosis, and which nodes were missed tells you why. This is what the unit choice below bought — twenty virtual machines do not fit on a workstation, and twenty system containers do.


Two kinds of test

Module tests are the daily case and the reason this exists: does my change work, end to end.

Mesh tests use the same machinery to ask whether the mesh itself behaves — that a credential rotation reaches every consumer, that delivery to an absent node is reported as pending rather than done, that a returning node catches up. These are fewer and change rarely, but they are where the known production faults get encoded so they stay fixed.

The known faults become mesh tests that fail today. Phase 0 of 00-work-breakdown.md is now complete on this basis — twenty-two assertions on real machines — and what it does not cover is recorded there: no module from the existing system has run against any of it yet.


The foundation

A node is a system container

An OS userspace with its own init, its own network interface, its own filesystem, sharing the host kernel.

A node's job is to run containers, so modelling a node as an application container inverts the thing being modelled: it forces nested containers through a privileged daemon or a shared socket, and a shared socket makes isolation between nodes cosmetic. A system container has no such problem — init runs as PID 1, so units and timers work as written; containers nest properly, so module stacks run as they do anywhere. Module code and mesh code both run unmodified, which is the property that makes a verdict trustworthy. Anything needing a special case locally is a divergence that will hide a fault.

Boot is around a second and snapshots are cheap, which is what makes this an inner loop rather than an errand.

Any node can be a full virtual machine instead, through the same tooling and the same test file — when a question needs a kernel to answer it, or when the point is that nodes are not identical.

The network

Two segments and an overlay, because some module behaviour is only visible across a real network boundary:

  • wan — a published node holds an address here, and an authoritative resolver maps its name to it, so a public touchpoint is real enough to exercise routing, virtual hosts and certificates
  • local — behind translation, as a home network is
  • the overlay — the mechanism production uses; the mesh addresses peers by mesh name and never learns which segment anyone is on

A node can be moved between segments or detached entirely, mid-test.

What is not real

  • the model provider — thinking is stubbed, so tests cost nothing to run
  • the public internet — a bridge, with an authoritative resolver rather than delegation
  • the public certificate authority — the lab runs its own ACME issuer on the public segment, so issuance, challenge and renewal are genuinely exercised rather than stubbed. Internal names keep the mesh CA, so the lab preserves production's two-authority split rather than collapsing it into one.

Everything a node itself does is real, because a node is a real machine.


A suite too expensive to run on every push says when it last ran

Written 2026-08-31, from resolving 04-ISSUES/005.

This suite needs a machine with a hypervisor. It therefore cannot run on every push, and a suite that does not run on every push runs when somebody remembers. Remembering is not a mechanism, and the harness this one replaces proves it: it had not built for two and a half months, nothing said so, and the coverage was assumed rather than checked.

The danger is not that the suite breaks. It is that nobody notices it stopped running — and that danger belongs to this design, not to the harness it retired.

So three rules, each held by a test:

A run leaves a receipt — when, what passed, what it ran, and the commit each repository was at. Kept outside version control: the question is has this machine run it, and a receipt in git would be a claim about everybody's machine made by whoever committed last.

A receipt says why it does not count. Old, failed, taken against commits the repositories have moved past, or a run that never raised a machine. Something can be asked, and answers non-zero. A receipt that says nothing about something is not a receipt that clears it — including a receipt written before it recorded a given fact, which claims nothing rather than everything.

The run rebuilds what it tests. The suite consumes artifacts from other repositories, and an artifact rebuilt from memory is one rebuilt sometimes. A stale binary reporting success against rules that have since changed is the same fault wearing different clothes.

The general rule, which outlives this suite: silence and success must never look alike. It is the same rule the host follows about a service that does not exist (ADR 0004) — absence must be distinguishable from a failure to answer — applied to coverage instead of to a machine.

A bed reads the catalogue it proves

Written 2026-09-21, from resolving 04-ISSUES/073; decided in ADR 0089.

The rule above — the run rebuilds what it tests — was held for binaries and images and not for manifests. Beds built the manifests they install inline, as literals copied from the catalogue when each bed was written; the copies did not move when the catalogue did, and a catalogue change was proven by no bed at all, while every bed stayed green against its copy.

A bed that installs a catalogue module reads that module's manifest from the catalogue the run was pointed at. It rewrites what the lab must — a build artifact becomes the image the machine holds, an image is pinned, a host port is remapped where one machine carries colliding modules, an address may point at a stand-in the bed raises — and nothing else.

A bed that needs less than the module declares is not testing that module. No upstream server, a secret in the environment, a requirement edge cut so no second provider is needed: that is a mesh test, and it carries a fixture with a name of its own, never a catalogue module's.

The receipt names the catalogue's commit with the others', so a run taken before a manifest changed says so — the same rule as for the binaries, for the same reason.

How it is checked: a unit test in the lab refuses an inline manifest literal that names a catalogue module unless the bed is declared, with its reason, in the test's own list, and refuses a declaration for a copy that is gone; the receipt test asserts the catalogue is claimed whenever the run is pointed at one.

Consequences

Bringing a node into being is part of the framework. A test creates its own nodes — one for the smallest, twenty for the largest — repeatably, unattended, and cheaply enough to do it twenty times in a row. Those are the constraints a real mesh wants, so the mechanism the runner needs is the one the mesh should keep.

Module verification becomes worth writing, because it is the thing that gives a developer a verdict, not just a stricter deploy.

Host-borrowing ends. Today's tooling starts providers on the host's own init system and reads credentials from host paths, because there is nowhere else to put a mesh. Once there is, a workstation stops being collateral.

The supervision question stops gating anything. 01-RESEARCH/003-service-supervision remains open on its own merits — and once this exists, its options are cheap to try rather than expensive to argue about.

Deliberately not decided

Whether the lab verdict is a workflow guard or an advisory check on the pull request. Open, and deliberately trivial — a policy detail, changeable in an afternoon, not an architectural choice. Recorded so it is not mistaken for an oversight.