From a2495e4d8ee79e7b8b39b5cc9bc685eb236d60e0 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 22:22:19 +0200 Subject: [PATCH] ADR 0034 (proposed): a test defends a decision MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit how-we-build 5 already says that if a document states a rule about the mesh, it says how the rule is verified — an unenforced rule being indistinguishable from a wrong one, and costing more because people believe it. That has never been applied to decisions, and a decision record states the same kind of claim. The gap was found by review: the lab reached 2,128 lines with 1,072 untested and no stated rule broken, because there is no testing posture in how-we-build at all. Every decision the lab embodies was verified by hand and none of it survives the terminal it ran in — which is 04-ISSUES/005 in miniature, coverage assumed rather than checked. Rejected a coverage percentage: it measures how much code a test touched, not whether anything important is defended, and would have been satisfied by testing the parser harder while the hypervisor integration stayed unasserted. Rejected test-driven development as a hard rule, and not because it is wrong in general. Half this implementation was discovery — that the hypervisor CLI reads a definition from stdin and hangs, that it assigns a MAC without recording it, that a stock image's networking flushes a static address. A test written first against undiscovered behaviour asserts a guess. So: structure and logic tested first, behaviour against a real system tested alongside, mocking the boundary forbidden, and a blocking gate as the definition of done. A test names the decision it defends, which is what makes the pairing checkable — a decision without one can be found rather than noticed. Stated as proposed rather than adopted: 6 requires review by someone who is not the proposer. Records 0001-0033 predate it and are not retroactively invalid, but each should acquire a test or an explicit note that it cannot have one, and until then the rule is aspirational for them — which is the state 5 warns about, recorded rather than hidden. --- 00-META/how-we-build.md | 20 ++++ .../0034-a-test-defends-a-decision.md | 91 +++++++++++++++++++ 2 files changed, 111 insertions(+) create mode 100644 02-DECISIONS/0034-a-test-defends-a-decision.md diff --git a/00-META/how-we-build.md b/00-META/how-we-build.md index 9926fd4..a353a70 100644 --- a/00-META/how-we-build.md +++ b/00-META/how-we-build.md @@ -152,6 +152,26 @@ finds one names the evidence required and returns the work. The reason is the mesh's most consistent failure shape: a green result proves transport, not effect. Absence reads as success unless something looked. +### A test defends a decision + +*Proposed — [ADR 0034](../02-DECISIONS/0034-a-test-defends-a-decision.md), pending review.* + +The rule above applies to prose. It applies to **decisions** too: a decision record states +something that must be true, and a test asserts it. A decision with no test is one that will +quietly stop being true, and nobody will learn that from a document. + +- Structure and logic — what is accepted, what is refused, how a value is derived — is tested + **first**, because the behaviour is knowable before the code. +- Behaviour against a real system is tested **alongside**, because it is discovered rather than + known. +- **Mocking the boundary is forbidden.** A test that fakes the system under integration asserts + that the fake behaves as expected. +- **The gate is blocking.** Green is the definition of done; a change that has not run its tests + is not finished, whatever the diff looks like. + +Not test-driven development as a blanket rule — a test written first against undiscovered +behaviour asserts a guess. The obligation is that every decision has a defender. + ### Search the record before forming a hypothesis The first action on any error message, failing service or unexpected behaviour is to search the diff --git a/02-DECISIONS/0034-a-test-defends-a-decision.md b/02-DECISIONS/0034-a-test-defends-a-decision.md new file mode 100644 index 0000000..93ef48f --- /dev/null +++ b/02-DECISIONS/0034-a-test-defends-a-decision.md @@ -0,0 +1,91 @@ +--- +status: proposed +date: 2026-08-24 +deciders: jochen +reconstructed: false +--- + +# 34. A test defends a decision + +## Context + +[`how-we-build.md`](../00-META/how-we-build.md) §5 already says that **if a document states a +rule about the mesh, it says how the rule is verified**, on the grounds that an unenforced rule +is indistinguishable from a wrong one and costs more, because people believe it. + +That rule is applied to prose and to acceptance criteria. It has never been applied to +**decisions**, and it should be — a decision record states something that must be true, which +is the same kind of claim. + +The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and +**no stated rule was broken.** There is no testing posture in `how-we-build.md` at all: no +expectation, no gate, no definition of done. Every decision the lab embodies — the underlay +boundary, the closed address space, routers as scenery, waiting for usable rather than for a +call to return — was verified by hand, by running scenarios and reading output, and none of +that survives the terminal it was run in. + +Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has +not built since 2026-06-04, and nothing said so +([`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). Coverage +assumed rather than checked. + +## Considered options + +1. **A coverage percentage.** Rejected. It measures how much code a test touched, not whether + anything important is defended, and it is satisfied by tests that assert nothing. A number + would have been met by testing the parser harder while the hypervisor integration stayed + unasserted. +2. **Test-driven development as a hard rule.** Rejected, and not because it is wrong in general. + Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from + stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time + networking flushes a static address. A test written first against undiscovered behaviour + asserts a guess. +3. **A test defends a decision.** Chosen. + +## Decision + +**Every decision record states something that must be true. A test asserts it.** + +A decision with no test is a decision that will quietly stop being true, and nobody will find +out from a document. Concretely: + +- Where a decision is about **structure or logic** — what a declaration may say, what is + refused, how a name is derived — the test is a unit test, and it is **written first**, because + the behaviour is knowable before the code. +- Where a decision is about **behaviour against a real system** — a hypervisor, a broker, a + daemon — the test runs against the real thing, and is written **alongside**, because the + behaviour is discovered rather than known. +- **Mocking the boundary is forbidden.** A test that fakes a hypervisor asserts that the fake + behaves as expected, which is the shape of test this whole effort exists to stop shipping. +- **The gate is blocking, and green is the definition of done.** A change that has not run its + tests is not finished, whatever its diff looks like. + +A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable, +so a decision without one can be *found* rather than noticed. + +## Consequences + +- The question *"which tests matter"* has an answer that is not a number. The decisions are the + list, and they are already written down. +- **A new decision costs a test.** That is the intended friction — a decision nobody will assert + is one worth reconsidering. +- Integration tests need real infrastructure and are slow. That cost is accepted: a fast test + suite that mocks the boundary would tell us nothing about the boundary, which is where every + interesting fault in this session actually was. +- Some decisions are not mechanically assertable — *the mesh brokers capabilities; nodes host; + agents think* is a shape, not a predicate. Those should say so in the record rather than being + quietly exempt, so the exemption is visible. +- Records 0001–0033 were made before this rule. They are not retroactively invalid, but each + should acquire a test or an explicit note that it cannot have one, and until then this rule + is aspirational for them — which is exactly the state §5 warns about, recorded rather than + hidden. + +## References + +- [`how-we-build.md`](../00-META/how-we-build.md) §5 — the rule this extends from prose to + decisions. +- [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md) — coverage + assumed rather than checked, for two and a half months. +- The sibling HQ repository for the PAPA platform states the same boundary rule — the contract + is tested against the real system, mocking the client is forbidden, and a blocking gate is the + definition of done. This record adopts that posture and adds the decision pairing.