diff --git a/03-DESIGN/01-to-be/01-end-to-end-testing.md b/03-DESIGN/01-to-be/01-end-to-end-testing.md index 192379e..a7e0574 100644 --- a/03-DESIGN/01-to-be/01-end-to-end-testing.md +++ b/03-DESIGN/01-to-be/01-end-to-end-testing.md @@ -2,9 +2,8 @@ layer: to-be status: in-progress code: [mesh-lab] -updated: 2026-08-28 +updated: 2026-08-31 decisions: - - 02-DECISIONS/0016-the-lab.md - 02-DECISIONS/0016-the-lab.md - 02-DECISIONS/0019-how-this-repository-works.md --- @@ -390,6 +389,38 @@ Everything a node itself does is real, because a node is a real machine. --- +## A suite too expensive to run on every push says when it last ran + +*Written 2026-08-31, from resolving [04-ISSUES/005](../../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).* + +This suite needs a machine with a hypervisor. It therefore cannot run on every push, and a suite +that does not run on every push runs **when somebody remembers**. Remembering is not a mechanism, +and the harness this one replaces proves it: it had not built for two and a half months, nothing +said so, and the coverage was assumed rather than checked. + +**The danger is not that the suite breaks. It is that nobody notices it stopped running** — and +that danger belongs to *this* design, not to the harness it retired. + +So three rules, each held by a test: + +**A run leaves a receipt** — when, what passed, what it ran, and the commit each repository was +at. Kept **outside version control**: the question is *has this machine run it*, and a receipt in +git would be a claim about everybody's machine made by whoever committed last. + +**A receipt says why it does not count.** Old, failed, taken against commits the repositories have +moved past, or a run that never raised a machine. Something can be asked, and answers non-zero. +**A receipt that says nothing about something is not a receipt that clears it** — including a +receipt written before it recorded a given fact, which claims nothing rather than everything. + +**The run rebuilds what it tests.** The suite consumes artifacts from other repositories, and an +artifact rebuilt from memory is one rebuilt sometimes. A stale binary reporting success against +rules that have since changed is the same fault wearing different clothes. + +**The general rule, which outlives this suite:** *silence and success must never look alike.* +It is the same rule the host follows about a service that does not exist +([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) — absence must be distinguishable +from a failure to answer — applied to coverage instead of to a machine. + ## Consequences **Bringing a node into being is part of the framework.** A test creates its own nodes — one diff --git a/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md b/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md index e2879c1..def605b 100644 --- a/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md +++ b/04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md @@ -1,9 +1,9 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [hal] -fixed-by: -amended-design: +located-in: [hal, mesh-lab] +fixed-by: mesh-lab — a run leaves a receipt, and the receipt says what it covered +amended-design: 03-DESIGN/01-to-be/01-end-to-end-testing.md --- # 005 — The end-to-end pipeline harness has not built since the workspace was removed @@ -45,3 +45,52 @@ unmentioned: until the lab exists, this is the coverage the pipeline is presumed - Repair, or retire in favour of the lab? Leaving it in the repository unbuilt is the one option that keeps the false impression of coverage. - Was anything relying on it, or had it already stopped running before the workspace removal? + +## Resolution + +*2026-08-31.* **Retired in favour of the lab, and the reason it went unnoticed was fixed +separately from the harness itself.** + +The old harness is not repaired. What replaced it is the end-to-end suite on a lab mesh, which +raises real machines and proves the pipeline against them. That answers the first open question. + +The second finding is the one worth keeping. *Nothing runs it, and nothing reports that nothing +runs it* is not a fact about that harness — it is a fact about **any** suite too expensive to run +on every push, and the lab suite is exactly that: it needs a machine with a hypervisor, so it runs +when somebody remembers. **Remembering is not a mechanism**, and the replacement inherited the +fault it was replacing. + +So three things now hold, each checked by a test that was confirmed to fail without it: + +- **A run leaves a receipt** — when it ran, what passed, and the commit each repository was at. + Kept outside version control, because the question is *has this machine run it*, and a receipt in + git would be a claim about everybody's machine made by whoever committed last. +- **The receipt can be judged, and says why it does not count.** Old, failed, taken against code + the repositories have since moved past, or a run that never raised a machine — each reads + differently, and only the last of those is new. **A receipt that says nothing about something is + not a receipt that clears it.** +- **The artifacts are rebuilt by the run, not beside it.** The suite consumes three artifacts from + two repositories. They were rebuilt by hand, from memory, and a rename that needed two of them + got one — leaving a binary eleven hours old refusing a field the mesh had just renamed, found by + a full run. That step now lives in the repository rather than in a terminal history. + +### What this issue taught twice + +**The fix reintroduced the fault, in miniature, and the second time was caught by running it.** + +The suite takes paths, so it can be pointed at one quick unit file — and the receipt from that run +was, at first, indistinguishable from a receipt for the real thing. A green record standing for a +run that raised no machines is this issue's own symptom, rebuilt inside its remedy. The receipt now +records what it ran, and a run that did not include the end-to-end file is not coverage. + +Separately, the code that decides *no receipt rather than a guessed one* — the rule that keeps the +record meaning something — was first written where no test could reach it. Writing "0 failed" +because nothing said otherwise is how a green record comes to mean nothing. + +And the counting itself **passed every test while reading nothing**: the test runner colours its +summary even into a pipe, so the anchored pattern never matched, and the fixtures it was checked +against were output that had been imagined rather than captured. **A fixture that agrees with the +mistake proves the mistake.** It is now checked against the runner's real bytes. + +Each of these was found by running the thing, not by reading it — which is the same argument this +issue makes about the pipeline.