The router is scenery, and a test defends a decision #8
@@ -152,6 +152,52 @@ finds one names the evidence required and returns the work.
|
|||||||
The reason is the mesh's most consistent failure shape: a green result proves transport, not
|
The reason is the mesh's most consistent failure shape: a green result proves transport, not
|
||||||
effect. Absence reads as success unless something looked.
|
effect. Absence reads as success unless something looked.
|
||||||
|
|
||||||
|
### A test defends a decision
|
||||||
|
|
||||||
|
*Proposed — [ADR 0034](../02-DECISIONS/0034-a-test-defends-a-decision.md), pending review.*
|
||||||
|
|
||||||
|
The rule above applies to prose. It applies to **decisions** too: a decision record states
|
||||||
|
something that must be true, and a test asserts it. A decision with no test is one that will
|
||||||
|
quietly stop being true, and nobody will learn that from a document.
|
||||||
|
|
||||||
|
- Structure and logic — what is accepted, what is refused, how a value is derived — is tested
|
||||||
|
**first**, because the behaviour is knowable before the code.
|
||||||
|
- Behaviour against a real system is tested **alongside**, because it is discovered rather than
|
||||||
|
known.
|
||||||
|
- **Mocking the boundary is forbidden.** A test that fakes the system under integration asserts
|
||||||
|
that the fake behaves as expected.
|
||||||
|
- **The gate is blocking.** Green is the definition of done; a change that has not run its tests
|
||||||
|
is not finished, whatever the diff looks like.
|
||||||
|
|
||||||
|
Not test-driven development as a blanket rule — a test written first against undiscovered
|
||||||
|
behaviour asserts a guess. The obligation is that every decision has a defender.
|
||||||
|
|
||||||
|
### A report is read from the system, never from what asked for it
|
||||||
|
|
||||||
|
*Proposed — [ADR 0035](../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md), pending
|
||||||
|
review.*
|
||||||
|
|
||||||
|
The same rule as the two above, pointed at reporting rather than at verification. **Anything
|
||||||
|
that describes the state of the mesh — a status view, an inventory, a diagram, a health
|
||||||
|
check — is assembled from the running system.** Assembling it from the intended state produces
|
||||||
|
a report that always agrees with itself and can never disagree with reality, which is not a
|
||||||
|
report.
|
||||||
|
|
||||||
|
Where the system does not natively hold a fact the report needs, **the thing that applied the
|
||||||
|
fact records it** — and:
|
||||||
|
|
||||||
|
> **A record of behaviour is written after the behaviour works, never when the resource is
|
||||||
|
> created.**
|
||||||
|
|
||||||
|
Written up front it restates the request in a new place and inherits none of the authority of
|
||||||
|
having happened. A failed run leaves its wreckage standing, and a report of that wreckage must
|
||||||
|
not describe what the wreckage was supposed to be.
|
||||||
|
|
||||||
|
This is the production form of the mesh's most expensive fault: a firewall key declared in five
|
||||||
|
manifests and read by no code
|
||||||
|
([04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)). The
|
||||||
|
declaration was never wrong. Nothing ever asked the system.
|
||||||
|
|
||||||
### Search the record before forming a hypothesis
|
### Search the record before forming a hypothesis
|
||||||
|
|
||||||
The first action on any error message, failing service or unexpected behaviour is to search the
|
The first action on any error message, failing service or unexpected behaviour is to search the
|
||||||
|
|||||||
@@ -0,0 +1,81 @@
|
|||||||
|
---
|
||||||
|
status: accepted
|
||||||
|
date: 2026-08-24
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
extends: 0016-a-lab-node-is-a-virtual-machine.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 33. A router is scenery, not a node — so it is a container
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual
|
||||||
|
machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install,
|
||||||
|
so it has to be a real machine or the thing under test is not the thing that ships.
|
||||||
|
|
||||||
|
A scenario also needs routers. NAT, port forwarding, policy between segments and mapping
|
||||||
|
expiry are all things a router does, and until one is materialised a multi-segment scenario
|
||||||
|
raises isolated islands
|
||||||
|
([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)).
|
||||||
|
The declaration already implies them: a gateway is *the one implicit machine in an otherwise
|
||||||
|
explicit declaration*.
|
||||||
|
|
||||||
|
The question is whether ADR 0016 binds those too.
|
||||||
|
|
||||||
|
## Considered options
|
||||||
|
|
||||||
|
1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency
|
||||||
|
nobody needs. A router boots in roughly ten seconds against a container's one; a
|
||||||
|
six-segment scenario wanting three routers spends thirty seconds per raise on scenery.
|
||||||
|
2. **The hypervisor provides NAT** — bridges with translation switched on, and its own
|
||||||
|
forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide
|
||||||
|
what the declaration is supposed to own, and it cannot express a mapping that expires, a
|
||||||
|
gateway that refuses to forward, or policy between siblings. The model would shrink to fit
|
||||||
|
the tool.
|
||||||
|
3. **A router is scenery, and scenery is a container.** Chosen.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**ADR 0016 binds nodes. A router is not a node.**
|
||||||
|
|
||||||
|
Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh
|
||||||
|
never installs anything on it, and no assertion is ever made about its internals. It exists so
|
||||||
|
that packets between machines behave the way they behave in the world — which is the definition
|
||||||
|
of scenery.
|
||||||
|
|
||||||
|
So a router is a **system container**, and the fidelity argument does not reach it: what a
|
||||||
|
router must reproduce is kernel behaviour — translation, connection tracking, filtering,
|
||||||
|
forwarding — and a container has the same kernel.
|
||||||
|
|
||||||
|
**Verified before deciding, not assumed.** In a plain unprivileged container:
|
||||||
|
|
||||||
|
| Needed for | Works |
|
||||||
|
|---|---|
|
||||||
|
| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` |
|
||||||
|
| `nat:` | nftables masquerade, rules accepted and listed back |
|
||||||
|
| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` |
|
||||||
|
|
||||||
|
No privileged mode, no nesting, no capability grants.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten,
|
||||||
|
and a scenario's cost tracks the machines actually under test.
|
||||||
|
- **The distinction is now load-bearing and has to stay legible.** *Node* means something under
|
||||||
|
test; *scenery* means something that makes the test real. If anything is ever installed on a
|
||||||
|
router by the mesh, it has become a node and this decision no longer covers it.
|
||||||
|
- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra
|
||||||
|
concept — justified by it being true, rather than by the saving.
|
||||||
|
- A container shares the host kernel, so a scenario cannot reproduce a router running a
|
||||||
|
*different* kernel from the workstation. Nothing currently wants that; if something does, that
|
||||||
|
router becomes a virtual machine and this record needs revisiting rather than bending.
|
||||||
|
- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and
|
||||||
|
never names the machine that serves it — which is right, because it is not a machine the
|
||||||
|
scenario has anything to say about.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged.
|
||||||
|
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router,
|
||||||
|
and why *published but behind NAT* only exists in production today.
|
||||||
@@ -0,0 +1,91 @@
|
|||||||
|
---
|
||||||
|
status: proposed
|
||||||
|
date: 2026-08-24
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
---
|
||||||
|
|
||||||
|
# 34. A test defends a decision
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
[`how-we-build.md`](../00-META/how-we-build.md) §5 already says that **if a document states a
|
||||||
|
rule about the mesh, it says how the rule is verified**, on the grounds that an unenforced rule
|
||||||
|
is indistinguishable from a wrong one and costs more, because people believe it.
|
||||||
|
|
||||||
|
That rule is applied to prose and to acceptance criteria. It has never been applied to
|
||||||
|
**decisions**, and it should be — a decision record states something that must be true, which
|
||||||
|
is the same kind of claim.
|
||||||
|
|
||||||
|
The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and
|
||||||
|
**no stated rule was broken.** There is no testing posture in `how-we-build.md` at all: no
|
||||||
|
expectation, no gate, no definition of done. Every decision the lab embodies — the underlay
|
||||||
|
boundary, the closed address space, routers as scenery, waiting for usable rather than for a
|
||||||
|
call to return — was verified by hand, by running scenarios and reading output, and none of
|
||||||
|
that survives the terminal it was run in.
|
||||||
|
|
||||||
|
Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has
|
||||||
|
not built since 2026-06-04, and nothing said so
|
||||||
|
([`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). Coverage
|
||||||
|
assumed rather than checked.
|
||||||
|
|
||||||
|
## Considered options
|
||||||
|
|
||||||
|
1. **A coverage percentage.** Rejected. It measures how much code a test touched, not whether
|
||||||
|
anything important is defended, and it is satisfied by tests that assert nothing. A number
|
||||||
|
would have been met by testing the parser harder while the hypervisor integration stayed
|
||||||
|
unasserted.
|
||||||
|
2. **Test-driven development as a hard rule.** Rejected, and not because it is wrong in general.
|
||||||
|
Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from
|
||||||
|
stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time
|
||||||
|
networking flushes a static address. A test written first against undiscovered behaviour
|
||||||
|
asserts a guess.
|
||||||
|
3. **A test defends a decision.** Chosen.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**Every decision record states something that must be true. A test asserts it.**
|
||||||
|
|
||||||
|
A decision with no test is a decision that will quietly stop being true, and nobody will find
|
||||||
|
out from a document. Concretely:
|
||||||
|
|
||||||
|
- Where a decision is about **structure or logic** — what a declaration may say, what is
|
||||||
|
refused, how a name is derived — the test is a unit test, and it is **written first**, because
|
||||||
|
the behaviour is knowable before the code.
|
||||||
|
- Where a decision is about **behaviour against a real system** — a hypervisor, a broker, a
|
||||||
|
daemon — the test runs against the real thing, and is written **alongside**, because the
|
||||||
|
behaviour is discovered rather than known.
|
||||||
|
- **Mocking the boundary is forbidden.** A test that fakes a hypervisor asserts that the fake
|
||||||
|
behaves as expected, which is the shape of test this whole effort exists to stop shipping.
|
||||||
|
- **The gate is blocking, and green is the definition of done.** A change that has not run its
|
||||||
|
tests is not finished, whatever its diff looks like.
|
||||||
|
|
||||||
|
A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable,
|
||||||
|
so a decision without one can be *found* rather than noticed.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- The question *"which tests matter"* has an answer that is not a number. The decisions are the
|
||||||
|
list, and they are already written down.
|
||||||
|
- **A new decision costs a test.** That is the intended friction — a decision nobody will assert
|
||||||
|
is one worth reconsidering.
|
||||||
|
- Integration tests need real infrastructure and are slow. That cost is accepted: a fast test
|
||||||
|
suite that mocks the boundary would tell us nothing about the boundary, which is where every
|
||||||
|
interesting fault in this session actually was.
|
||||||
|
- Some decisions are not mechanically assertable — *the mesh brokers capabilities; nodes host;
|
||||||
|
agents think* is a shape, not a predicate. Those should say so in the record rather than being
|
||||||
|
quietly exempt, so the exemption is visible.
|
||||||
|
- Records 0001–0033 were made before this rule. They are not retroactively invalid, but each
|
||||||
|
should acquire a test or an explicit note that it cannot have one, and until then this rule
|
||||||
|
is aspirational for them — which is exactly the state §5 warns about, recorded rather than
|
||||||
|
hidden.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [`how-we-build.md`](../00-META/how-we-build.md) §5 — the rule this extends from prose to
|
||||||
|
decisions.
|
||||||
|
- [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md) — coverage
|
||||||
|
assumed rather than checked, for two and a half months.
|
||||||
|
- The sibling HQ repository for the PAPA platform states the same boundary rule — the contract
|
||||||
|
is tested against the real system, mocking the client is forbidden, and a blocking gate is the
|
||||||
|
definition of done. This record adopts that posture and adds the decision pairing.
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
---
|
||||||
|
status: proposed
|
||||||
|
date: 2026-08-24
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
extends: 0034-a-test-defends-a-decision.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 35. A picture of a system is read from the system, never from what asked for it
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets.
|
||||||
|
The two are supposed to correspond, and the entire value of the lab rests on noticing when
|
||||||
|
they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a
|
||||||
|
claim that will quietly stop being true.
|
||||||
|
|
||||||
|
Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not.
|
||||||
|
A diagram of a running system can be produced two ways: parse the declaration and lay it out,
|
||||||
|
or interrogate the system and lay *that* out. The first is far easier — the declaration is
|
||||||
|
already parsed, already validated, already in memory.
|
||||||
|
|
||||||
|
It is also worthless for the only question worth asking of such a picture: *is what is running
|
||||||
|
what I asked for?* A drawing built from the request and captioned **as raised** answers that
|
||||||
|
question with the request, which always agrees with itself.
|
||||||
|
|
||||||
|
This is the same fault as
|
||||||
|
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a firewall key declared in
|
||||||
|
five manifests and read by no code, so a manifest appears to restrict a port and restricts
|
||||||
|
nothing. The declaration was never wrong. Nothing ever asked the system.
|
||||||
|
|
||||||
|
## Considered options
|
||||||
|
|
||||||
|
1. **Draw the declaration, and label it honestly.** Cheap, and useful for review before
|
||||||
|
anything is raised. Insufficient alone: it can never disagree with itself.
|
||||||
|
2. **Draw the system, inferring the rest from the declaration where the system is silent.**
|
||||||
|
The tempting middle. Rejected — a picture where some facts are observed and some are
|
||||||
|
assumed has no honest caption, and the assumed ones are exactly the interesting ones.
|
||||||
|
3. **Two pictures, one layout, neither borrowing from the other.** Chosen.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**A picture captioned *as raised* reads only the running system.** It never opens the
|
||||||
|
declaration, not even for a fact the system happens not to record.
|
||||||
|
|
||||||
|
Where the hypervisor does not natively hold a fact the picture needs — whether a segment is
|
||||||
|
public, what a gateway translates, whether a machine refuses inbound — **the raise records it
|
||||||
|
on the resource** as metadata, and the picture reads it back from there.
|
||||||
|
|
||||||
|
That recording carries its own rule, which is the substance of this decision rather than an
|
||||||
|
implementation note:
|
||||||
|
|
||||||
|
> **A behavioural tag is written after the behaviour works, never when the resource is
|
||||||
|
> created.**
|
||||||
|
|
||||||
|
Written at creation, a tag restates the request in a new location and inherits none of the
|
||||||
|
authority of having happened. A failed raise deliberately leaves its wreckage standing, so a
|
||||||
|
tag written up front would let a picture of that wreckage badge translation the gateway was
|
||||||
|
never configured to do — reproducing, inside the tool built to catch the fault, exactly the
|
||||||
|
fault.
|
||||||
|
|
||||||
|
So: the gateway is tagged after its ruleset applies; the machine after the read-back proves
|
||||||
|
its firewall loaded.
|
||||||
|
|
||||||
|
Both pictures render through **one layout**, so they can be put side by side and the
|
||||||
|
difference read off directly.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- **It earned itself on the first comparison.** Drawn side by side, every virtual machine in
|
||||||
|
the live picture held no addresses at all. A container's interface carries the name of the
|
||||||
|
device it was configured as; a virtual machine names its own — so joining addresses to
|
||||||
|
devices by name attached every address to a container and none to a VM. Nothing failed;
|
||||||
|
a whole class of machine silently lost its addresses. The two pictures disagreed, so it was
|
||||||
|
visible in seconds. It is now joined on MAC.
|
||||||
|
- Raise does more work, and writes metadata it does not itself consume. Accepted: the cost is
|
||||||
|
a few config keys, and it is what makes a raised instance self-describing.
|
||||||
|
- A resource raised before a tag existed is missing it. The reader says so rather than filling
|
||||||
|
the gap from the declaration — an untagged link draws as unknown, not as what the file said
|
||||||
|
it should be.
|
||||||
|
- **The rule generalises past diagrams.** Anything reporting on the mesh — a status view, an
|
||||||
|
inventory, a health check — is subject to it. A report assembled from the intended state is
|
||||||
|
not a report.
|
||||||
|
- The declared picture stays, and stays useful: it is review before raising, and it is one half
|
||||||
|
of the comparison. It carries no runtime status, because it cannot know any.
|
||||||
|
|
||||||
|
- **The constitution sync is not done and must not be.** `how-we-build.md` §5 carries this rule
|
||||||
|
marked *proposed*, and playbook
|
||||||
|
[05](../00-META/process/05-constitution-sync.md) publishes the derived page only for rules the
|
||||||
|
mesh should enforce now. A rule the mesh enforces before a second person has agreed to it is
|
||||||
|
the failure mode §6 exists to prevent. The sync happens when this record is accepted, and this
|
||||||
|
line is what makes the gap visible rather than silent.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
|
||||||
|
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the
|
||||||
|
mesh is responsible for; the same instinct, applied to facts rather than to configuration.
|
||||||
|
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production
|
||||||
|
form.
|
||||||
@@ -256,6 +256,11 @@ reachability, and it does not.
|
|||||||
The lab materialises a machine to be the gateway. That is the one implicit machine in an
|
The lab materialises a machine to be the gateway. That is the one implicit machine in an
|
||||||
otherwise explicit declaration, and it exists because NAT has to run somewhere.
|
otherwise explicit declaration, and it exists because NAT has to run somewhere.
|
||||||
|
|
||||||
|
It is a **container, not a virtual machine** — a router is scenery rather than something under
|
||||||
|
test, so the fidelity argument that makes a node a virtual machine does not reach it
|
||||||
|
([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md)). What a router must
|
||||||
|
reproduce is kernel behaviour, and a container has the same kernel.
|
||||||
|
|
||||||
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
|
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
|
||||||
segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node
|
segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node
|
||||||
with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on
|
with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on
|
||||||
|
|||||||
@@ -133,6 +133,12 @@ decision rather than a second implementation.
|
|||||||
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
||||||
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||||
decides whether a person can find the one they left standing yesterday.
|
decides whether a person can find the one they left standing yesterday.
|
||||||
|
- ~~**Does a scenario snapshot need the machines stopped?**~~ **Answered by the integration
|
||||||
|
test on its first run: no, but they must be flushed.** A snapshot captures disk and not
|
||||||
|
memory, so a write still in the guest's page cache is absent from it — not stale, absent. A
|
||||||
|
file written seconds before a snapshot did not survive the restore. Flushing first buys
|
||||||
|
write-durability; it does not buy application-consistency, and anything mid-transaction is
|
||||||
|
still captured mid-transaction.
|
||||||
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||||
the instance must not destroy them.
|
the instance must not destroy them.
|
||||||
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
||||||
|
|||||||
Reference in New Issue
Block a user