Tier 0: the questions answered, the decisions taken, and the design #9

Merged
jschoubben merged 16 commits from design/close-the-record into main 2026-08-25 22:56:07 +00:00
2 changed files with 111 additions and 0 deletions
Showing only changes of commit a2495e4d8e - Show all commits
+20
View File
@@ -152,6 +152,26 @@ finds one names the evidence required and returns the work.
The reason is the mesh's most consistent failure shape: a green result proves transport, not
effect. Absence reads as success unless something looked.
### A test defends a decision
*Proposed — [ADR 0034](../02-DECISIONS/0034-a-test-defends-a-decision.md), pending review.*
The rule above applies to prose. It applies to **decisions** too: a decision record states
something that must be true, and a test asserts it. A decision with no test is one that will
quietly stop being true, and nobody will learn that from a document.
- Structure and logic — what is accepted, what is refused, how a value is derived — is tested
**first**, because the behaviour is knowable before the code.
- Behaviour against a real system is tested **alongside**, because it is discovered rather than
known.
- **Mocking the boundary is forbidden.** A test that fakes the system under integration asserts
that the fake behaves as expected.
- **The gate is blocking.** Green is the definition of done; a change that has not run its tests
is not finished, whatever the diff looks like.
Not test-driven development as a blanket rule — a test written first against undiscovered
behaviour asserts a guess. The obligation is that every decision has a defender.
### Search the record before forming a hypothesis
The first action on any error message, failing service or unexpected behaviour is to search the
@@ -0,0 +1,91 @@
---
status: proposed
date: 2026-08-24
deciders: jochen
reconstructed: false
---
# 34. A test defends a decision
## Context
[`how-we-build.md`](../00-META/how-we-build.md) §5 already says that **if a document states a
rule about the mesh, it says how the rule is verified**, on the grounds that an unenforced rule
is indistinguishable from a wrong one and costs more, because people believe it.
That rule is applied to prose and to acceptance criteria. It has never been applied to
**decisions**, and it should be — a decision record states something that must be true, which
is the same kind of claim.
The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and
**no stated rule was broken.** There is no testing posture in `how-we-build.md` at all: no
expectation, no gate, no definition of done. Every decision the lab embodies — the underlay
boundary, the closed address space, routers as scenery, waiting for usable rather than for a
call to return — was verified by hand, by running scenarios and reading output, and none of
that survives the terminal it was run in.
Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has
not built since 2026-06-04, and nothing said so
([`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). Coverage
assumed rather than checked.
## Considered options
1. **A coverage percentage.** Rejected. It measures how much code a test touched, not whether
anything important is defended, and it is satisfied by tests that assert nothing. A number
would have been met by testing the parser harder while the hypervisor integration stayed
unasserted.
2. **Test-driven development as a hard rule.** Rejected, and not because it is wrong in general.
Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from
stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time
networking flushes a static address. A test written first against undiscovered behaviour
asserts a guess.
3. **A test defends a decision.** Chosen.
## Decision
**Every decision record states something that must be true. A test asserts it.**
A decision with no test is a decision that will quietly stop being true, and nobody will find
out from a document. Concretely:
- Where a decision is about **structure or logic** — what a declaration may say, what is
refused, how a name is derived — the test is a unit test, and it is **written first**, because
the behaviour is knowable before the code.
- Where a decision is about **behaviour against a real system** — a hypervisor, a broker, a
daemon — the test runs against the real thing, and is written **alongside**, because the
behaviour is discovered rather than known.
- **Mocking the boundary is forbidden.** A test that fakes a hypervisor asserts that the fake
behaves as expected, which is the shape of test this whole effort exists to stop shipping.
- **The gate is blocking, and green is the definition of done.** A change that has not run its
tests is not finished, whatever its diff looks like.
A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable,
so a decision without one can be *found* rather than noticed.
## Consequences
- The question *"which tests matter"* has an answer that is not a number. The decisions are the
list, and they are already written down.
- **A new decision costs a test.** That is the intended friction — a decision nobody will assert
is one worth reconsidering.
- Integration tests need real infrastructure and are slow. That cost is accepted: a fast test
suite that mocks the boundary would tell us nothing about the boundary, which is where every
interesting fault in this session actually was.
- Some decisions are not mechanically assertable — *the mesh brokers capabilities; nodes host;
agents think* is a shape, not a predicate. Those should say so in the record rather than being
quietly exempt, so the exemption is visible.
- Records 0001–0033 were made before this rule. They are not retroactively invalid, but each
should acquire a test or an explicit note that it cannot have one, and until then this rule
is aspirational for them — which is exactly the state §5 warns about, recorded rather than
hidden.
## References
- [`how-we-build.md`](../00-META/how-we-build.md) §5 — the rule this extends from prose to
decisions.
- [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md) — coverage
assumed rather than checked, for two and a half months.
- The sibling HQ repository for the PAPA platform states the same boundary rule — the contract
is tested against the real system, mocking the client is forbidden, and a blocking gate is the
definition of done. This record adopts that posture and adds the decision pairing.