The router is scenery, and a test defends a decision #8

Manually merged
jschoubben merged 4 commits from design/tests-defend-decisions into main 2026-08-25 22:56:07 +00:00
6 changed files with 329 additions and 0 deletions
+46
View File
@@ -152,6 +152,52 @@ finds one names the evidence required and returns the work.
The reason is the mesh's most consistent failure shape: a green result proves transport, not The reason is the mesh's most consistent failure shape: a green result proves transport, not
effect. Absence reads as success unless something looked. effect. Absence reads as success unless something looked.
### A test defends a decision
*Proposed — [ADR 0034](../02-DECISIONS/0034-a-test-defends-a-decision.md), pending review.*
The rule above applies to prose. It applies to **decisions** too: a decision record states
something that must be true, and a test asserts it. A decision with no test is one that will
quietly stop being true, and nobody will learn that from a document.
- Structure and logic — what is accepted, what is refused, how a value is derived — is tested
**first**, because the behaviour is knowable before the code.
- Behaviour against a real system is tested **alongside**, because it is discovered rather than
known.
- **Mocking the boundary is forbidden.** A test that fakes the system under integration asserts
that the fake behaves as expected.
- **The gate is blocking.** Green is the definition of done; a change that has not run its tests
is not finished, whatever the diff looks like.
Not test-driven development as a blanket rule — a test written first against undiscovered
behaviour asserts a guess. The obligation is that every decision has a defender.
### A report is read from the system, never from what asked for it
*Proposed — [ADR 0035](../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md), pending
review.*
The same rule as the two above, pointed at reporting rather than at verification. **Anything
that describes the state of the mesh — a status view, an inventory, a diagram, a health
check — is assembled from the running system.** Assembling it from the intended state produces
a report that always agrees with itself and can never disagree with reality, which is not a
report.
Where the system does not natively hold a fact the report needs, **the thing that applied the
fact records it** — and:
> **A record of behaviour is written after the behaviour works, never when the resource is
> created.**
Written up front it restates the request in a new place and inherits none of the authority of
having happened. A failed run leaves its wreckage standing, and a report of that wreckage must
not describe what the wreckage was supposed to be.
This is the production form of the mesh's most expensive fault: a firewall key declared in five
manifests and read by no code
([04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)). The
declaration was never wrong. Nothing ever asked the system.
### Search the record before forming a hypothesis ### Search the record before forming a hypothesis
The first action on any error message, failing service or unexpected behaviour is to search the The first action on any error message, failing service or unexpected behaviour is to search the
@@ -0,0 +1,81 @@
---
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 33. A router is scenery, not a node — so it is a container
## Context
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual
machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install,
so it has to be a real machine or the thing under test is not the thing that ships.
A scenario also needs routers. NAT, port forwarding, policy between segments and mapping
expiry are all things a router does, and until one is materialised a multi-segment scenario
raises isolated islands
([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)).
The declaration already implies them: a gateway is *the one implicit machine in an otherwise
explicit declaration*.
The question is whether ADR 0016 binds those too.
## Considered options
1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency
nobody needs. A router boots in roughly ten seconds against a container's one; a
six-segment scenario wanting three routers spends thirty seconds per raise on scenery.
2. **The hypervisor provides NAT** — bridges with translation switched on, and its own
forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide
what the declaration is supposed to own, and it cannot express a mapping that expires, a
gateway that refuses to forward, or policy between siblings. The model would shrink to fit
the tool.
3. **A router is scenery, and scenery is a container.** Chosen.
## Decision
**ADR 0016 binds nodes. A router is not a node.**
Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh
never installs anything on it, and no assertion is ever made about its internals. It exists so
that packets between machines behave the way they behave in the world — which is the definition
of scenery.
So a router is a **system container**, and the fidelity argument does not reach it: what a
router must reproduce is kernel behaviour — translation, connection tracking, filtering,
forwarding — and a container has the same kernel.
**Verified before deciding, not assumed.** In a plain unprivileged container:
| Needed for | Works |
|---|---|
| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` |
| `nat:` | nftables masquerade, rules accepted and listed back |
| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` |
No privileged mode, no nesting, no capability grants.
## Consequences
- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten,
and a scenario's cost tracks the machines actually under test.
- **The distinction is now load-bearing and has to stay legible.** *Node* means something under
test; *scenery* means something that makes the test real. If anything is ever installed on a
router by the mesh, it has become a node and this decision no longer covers it.
- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra
concept — justified by it being true, rather than by the saving.
- A container shares the host kernel, so a scenario cannot reproduce a router running a
*different* kernel from the workstation. Nothing currently wants that; if something does, that
router becomes a virtual machine and this record needs revisiting rather than bending.
- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and
never names the machine that serves it — which is right, because it is not a machine the
scenario has anything to say about.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged.
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router,
and why *published but behind NAT* only exists in production today.
@@ -0,0 +1,91 @@
---
status: proposed
date: 2026-08-24
deciders: jochen
reconstructed: false
---
# 34. A test defends a decision
## Context
[`how-we-build.md`](../00-META/how-we-build.md) §5 already says that **if a document states a
rule about the mesh, it says how the rule is verified**, on the grounds that an unenforced rule
is indistinguishable from a wrong one and costs more, because people believe it.
That rule is applied to prose and to acceptance criteria. It has never been applied to
**decisions**, and it should be — a decision record states something that must be true, which
is the same kind of claim.
The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and
**no stated rule was broken.** There is no testing posture in `how-we-build.md` at all: no
expectation, no gate, no definition of done. Every decision the lab embodies — the underlay
boundary, the closed address space, routers as scenery, waiting for usable rather than for a
call to return — was verified by hand, by running scenarios and reading output, and none of
that survives the terminal it was run in.
Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has
not built since 2026-06-04, and nothing said so
([`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). Coverage
assumed rather than checked.
## Considered options
1. **A coverage percentage.** Rejected. It measures how much code a test touched, not whether
anything important is defended, and it is satisfied by tests that assert nothing. A number
would have been met by testing the parser harder while the hypervisor integration stayed
unasserted.
2. **Test-driven development as a hard rule.** Rejected, and not because it is wrong in general.
Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from
stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time
networking flushes a static address. A test written first against undiscovered behaviour
asserts a guess.
3. **A test defends a decision.** Chosen.
## Decision
**Every decision record states something that must be true. A test asserts it.**
A decision with no test is a decision that will quietly stop being true, and nobody will find
out from a document. Concretely:
- Where a decision is about **structure or logic** — what a declaration may say, what is
refused, how a name is derived — the test is a unit test, and it is **written first**, because
the behaviour is knowable before the code.
- Where a decision is about **behaviour against a real system** — a hypervisor, a broker, a
daemon — the test runs against the real thing, and is written **alongside**, because the
behaviour is discovered rather than known.
- **Mocking the boundary is forbidden.** A test that fakes a hypervisor asserts that the fake
behaves as expected, which is the shape of test this whole effort exists to stop shipping.
- **The gate is blocking, and green is the definition of done.** A change that has not run its
tests is not finished, whatever its diff looks like.
A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable,
so a decision without one can be *found* rather than noticed.
## Consequences
- The question *"which tests matter"* has an answer that is not a number. The decisions are the
list, and they are already written down.
- **A new decision costs a test.** That is the intended friction — a decision nobody will assert
is one worth reconsidering.
- Integration tests need real infrastructure and are slow. That cost is accepted: a fast test
suite that mocks the boundary would tell us nothing about the boundary, which is where every
interesting fault in this session actually was.
- Some decisions are not mechanically assertable — *the mesh brokers capabilities; nodes host;
agents think* is a shape, not a predicate. Those should say so in the record rather than being
quietly exempt, so the exemption is visible.
- Records 0001–0033 were made before this rule. They are not retroactively invalid, but each
should acquire a test or an explicit note that it cannot have one, and until then this rule
is aspirational for them — which is exactly the state §5 warns about, recorded rather than
hidden.
## References
- [`how-we-build.md`](../00-META/how-we-build.md) §5 — the rule this extends from prose to
decisions.
- [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md) — coverage
assumed rather than checked, for two and a half months.
- The sibling HQ repository for the PAPA platform states the same boundary rule — the contract
is tested against the real system, mocking the client is forbidden, and a blocking gate is the
definition of done. This record adopts that posture and adds the decision pairing.
@@ -0,0 +1,100 @@
---
status: proposed
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0034-a-test-defends-a-decision.md
---
# 35. A picture of a system is read from the system, never from what asked for it
## Context
A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets.
The two are supposed to correspond, and the entire value of the lab rests on noticing when
they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a
claim that will quietly stop being true.
Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not.
A diagram of a running system can be produced two ways: parse the declaration and lay it out,
or interrogate the system and lay *that* out. The first is far easier — the declaration is
already parsed, already validated, already in memory.
It is also worthless for the only question worth asking of such a picture: *is what is running
what I asked for?* A drawing built from the request and captioned **as raised** answers that
question with the request, which always agrees with itself.
This is the same fault as
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a firewall key declared in
five manifests and read by no code, so a manifest appears to restrict a port and restricts
nothing. The declaration was never wrong. Nothing ever asked the system.
## Considered options
1. **Draw the declaration, and label it honestly.** Cheap, and useful for review before
anything is raised. Insufficient alone: it can never disagree with itself.
2. **Draw the system, inferring the rest from the declaration where the system is silent.**
The tempting middle. Rejected — a picture where some facts are observed and some are
assumed has no honest caption, and the assumed ones are exactly the interesting ones.
3. **Two pictures, one layout, neither borrowing from the other.** Chosen.
## Decision
**A picture captioned *as raised* reads only the running system.** It never opens the
declaration, not even for a fact the system happens not to record.
Where the hypervisor does not natively hold a fact the picture needs — whether a segment is
public, what a gateway translates, whether a machine refuses inbound — **the raise records it
on the resource** as metadata, and the picture reads it back from there.
That recording carries its own rule, which is the substance of this decision rather than an
implementation note:
> **A behavioural tag is written after the behaviour works, never when the resource is
> created.**
Written at creation, a tag restates the request in a new location and inherits none of the
authority of having happened. A failed raise deliberately leaves its wreckage standing, so a
tag written up front would let a picture of that wreckage badge translation the gateway was
never configured to do — reproducing, inside the tool built to catch the fault, exactly the
fault.
So: the gateway is tagged after its ruleset applies; the machine after the read-back proves
its firewall loaded.
Both pictures render through **one layout**, so they can be put side by side and the
difference read off directly.
## Consequences
- **It earned itself on the first comparison.** Drawn side by side, every virtual machine in
the live picture held no addresses at all. A container's interface carries the name of the
device it was configured as; a virtual machine names its own — so joining addresses to
devices by name attached every address to a container and none to a VM. Nothing failed;
a whole class of machine silently lost its addresses. The two pictures disagreed, so it was
visible in seconds. It is now joined on MAC.
- Raise does more work, and writes metadata it does not itself consume. Accepted: the cost is
a few config keys, and it is what makes a raised instance self-describing.
- A resource raised before a tag existed is missing it. The reader says so rather than filling
the gap from the declaration — an untagged link draws as unknown, not as what the file said
it should be.
- **The rule generalises past diagrams.** Anything reporting on the mesh — a status view, an
inventory, a health check — is subject to it. A report assembled from the intended state is
not a report.
- The declared picture stays, and stays useful: it is review before raising, and it is one half
of the comparison. It carries no runtime status, because it cannot know any.
- **The constitution sync is not done and must not be.** `how-we-build.md` §5 carries this rule
marked *proposed*, and playbook
[05](../00-META/process/05-constitution-sync.md) publishes the derived page only for rules the
mesh should enforce now. A rule the mesh enforces before a second person has agreed to it is
the failure mode §6 exists to prevent. The sync happens when this record is accepted, and this
line is what makes the gap visible rather than silent.
## References
- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the
mesh is responsible for; the same instinct, applied to facts rather than to configuration.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production
form.
@@ -256,6 +256,11 @@ reachability, and it does not.
The lab materialises a machine to be the gateway. That is the one implicit machine in an The lab materialises a machine to be the gateway. That is the one implicit machine in an
otherwise explicit declaration, and it exists because NAT has to run somewhere. otherwise explicit declaration, and it exists because NAT has to run somewhere.
It is a **container, not a virtual machine** — a router is scenery rather than something under
test, so the fidelity argument that makes a node a virtual machine does not reach it
([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md)). What a router must
reproduce is kernel behaviour, and a container has the same kernel.
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several **`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node
with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on
@@ -133,6 +133,12 @@ decision rather than a second implementation.
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md). [research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
- **Instance naming.** A declaration is a kind and instances are many; how they are named - **Instance naming.** A declaration is a kind and instances are many; how they are named
decides whether a person can find the one they left standing yesterday. decides whether a person can find the one they left standing yesterday.
- ~~**Does a scenario snapshot need the machines stopped?**~~ **Answered by the integration
test on its first run: no, but they must be flushed.** A snapshot captures disk and not
memory, so a write still in the guest's page cache is absent from it — not stale, absent. A
file written seconds before a snapshot did not survive the restore. Flushing first buys
write-durability; it does not buy application-consistency, and anything mid-transaction is
still captured mid-transaction.
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying - **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
the instance must not destroy them. the instance must not destroy them.
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and - **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and