Files
hq/02-DECISIONS/0034-a-test-defends-a-decision.md
T
jschoubben a2495e4d8e ADR 0034 (proposed): a test defends a decision
how-we-build 5 already says that if a document states a rule about the
mesh, it says how the rule is verified — an unenforced rule being
indistinguishable from a wrong one, and costing more because people
believe it. That has never been applied to decisions, and a decision
record states the same kind of claim.

The gap was found by review: the lab reached 2,128 lines with 1,072
untested and no stated rule broken, because there is no testing posture in
how-we-build at all. Every decision the lab embodies was verified by hand
and none of it survives the terminal it ran in — which is
04-ISSUES/005 in miniature, coverage assumed rather than checked.

Rejected a coverage percentage: it measures how much code a test touched,
not whether anything important is defended, and would have been satisfied
by testing the parser harder while the hypervisor integration stayed
unasserted.

Rejected test-driven development as a hard rule, and not because it is
wrong in general. Half this implementation was discovery — that the
hypervisor CLI reads a definition from stdin and hangs, that it assigns a
MAC without recording it, that a stock image's networking flushes a static
address. A test written first against undiscovered behaviour asserts a
guess.

So: structure and logic tested first, behaviour against a real system
tested alongside, mocking the boundary forbidden, and a blocking gate as
the definition of done. A test names the decision it defends, which is what
makes the pairing checkable — a decision without one can be found rather
than noticed.

Stated as proposed rather than adopted: 6 requires review by someone who
is not the proposer. Records 0001-0033 predate it and are not retroactively
invalid, but each should acquire a test or an explicit note that it cannot
have one, and until then the rule is aspirational for them — which is the
state 5 warns about, recorded rather than hidden.
2026-08-24 22:22:19 +02:00

4.8 KiB
Raw Blame History

status, date, deciders, reconstructed
status date deciders reconstructed
proposed 2026-08-24 jochen false

34. A test defends a decision

Context

how-we-build.md §5 already says that if a document states a rule about the mesh, it says how the rule is verified, on the grounds that an unenforced rule is indistinguishable from a wrong one and costs more, because people believe it.

That rule is applied to prose and to acceptance criteria. It has never been applied to decisions, and it should be — a decision record states something that must be true, which is the same kind of claim.

The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and no stated rule was broken. There is no testing posture in how-we-build.md at all: no expectation, no gate, no definition of done. Every decision the lab embodies — the underlay boundary, the closed address space, routers as scenery, waiting for usable rather than for a call to return — was verified by hand, by running scenarios and reading output, and none of that survives the terminal it was run in.

Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has not built since 2026-06-04, and nothing said so (04-ISSUES/005). Coverage assumed rather than checked.

Considered options

  1. A coverage percentage. Rejected. It measures how much code a test touched, not whether anything important is defended, and it is satisfied by tests that assert nothing. A number would have been met by testing the parser harder while the hypervisor integration stayed unasserted.
  2. Test-driven development as a hard rule. Rejected, and not because it is wrong in general. Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time networking flushes a static address. A test written first against undiscovered behaviour asserts a guess.
  3. A test defends a decision. Chosen.

Decision

Every decision record states something that must be true. A test asserts it.

A decision with no test is a decision that will quietly stop being true, and nobody will find out from a document. Concretely:

  • Where a decision is about structure or logic — what a declaration may say, what is refused, how a name is derived — the test is a unit test, and it is written first, because the behaviour is knowable before the code.
  • Where a decision is about behaviour against a real system — a hypervisor, a broker, a daemon — the test runs against the real thing, and is written alongside, because the behaviour is discovered rather than known.
  • Mocking the boundary is forbidden. A test that fakes a hypervisor asserts that the fake behaves as expected, which is the shape of test this whole effort exists to stop shipping.
  • The gate is blocking, and green is the definition of done. A change that has not run its tests is not finished, whatever its diff looks like.

A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable, so a decision without one can be found rather than noticed.

Consequences

  • The question "which tests matter" has an answer that is not a number. The decisions are the list, and they are already written down.
  • A new decision costs a test. That is the intended friction — a decision nobody will assert is one worth reconsidering.
  • Integration tests need real infrastructure and are slow. That cost is accepted: a fast test suite that mocks the boundary would tell us nothing about the boundary, which is where every interesting fault in this session actually was.
  • Some decisions are not mechanically assertable — the mesh brokers capabilities; nodes host; agents think is a shape, not a predicate. Those should say so in the record rather than being quietly exempt, so the exemption is visible.
  • Records 0001–0033 were made before this rule. They are not retroactively invalid, but each should acquire a test or an explicit note that it cannot have one, and until then this rule is aspirational for them — which is exactly the state §5 warns about, recorded rather than hidden.

References

  • how-we-build.md §5 — the rule this extends from prose to decisions.
  • 04-ISSUES/005 — coverage assumed rather than checked, for two and a half months.
  • The sibling HQ repository for the PAPA platform states the same boundary rule — the contract is tested against the real system, mocking the client is forbidden, and a blocking gate is the definition of done. This record adopts that posture and adds the decision pairing.