The router is scenery, and a test defends a decision #8
@@ -152,6 +152,52 @@ finds one names the evidence required and returns the work.
|
||||
The reason is the mesh's most consistent failure shape: a green result proves transport, not
|
||||
effect. Absence reads as success unless something looked.
|
||||
|
||||
### A test defends a decision
|
||||
|
||||
*Proposed — [ADR 0034](../02-DECISIONS/0034-a-test-defends-a-decision.md), pending review.*
|
||||
|
||||
The rule above applies to prose. It applies to **decisions** too: a decision record states
|
||||
something that must be true, and a test asserts it. A decision with no test is one that will
|
||||
quietly stop being true, and nobody will learn that from a document.
|
||||
|
||||
- Structure and logic — what is accepted, what is refused, how a value is derived — is tested
|
||||
**first**, because the behaviour is knowable before the code.
|
||||
- Behaviour against a real system is tested **alongside**, because it is discovered rather than
|
||||
known.
|
||||
- **Mocking the boundary is forbidden.** A test that fakes the system under integration asserts
|
||||
that the fake behaves as expected.
|
||||
- **The gate is blocking.** Green is the definition of done; a change that has not run its tests
|
||||
is not finished, whatever the diff looks like.
|
||||
|
||||
Not test-driven development as a blanket rule — a test written first against undiscovered
|
||||
behaviour asserts a guess. The obligation is that every decision has a defender.
|
||||
|
||||
### A report is read from the system, never from what asked for it
|
||||
|
||||
*Proposed — [ADR 0035](../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md), pending
|
||||
review.*
|
||||
|
||||
The same rule as the two above, pointed at reporting rather than at verification. **Anything
|
||||
that describes the state of the mesh — a status view, an inventory, a diagram, a health
|
||||
check — is assembled from the running system.** Assembling it from the intended state produces
|
||||
a report that always agrees with itself and can never disagree with reality, which is not a
|
||||
report.
|
||||
|
||||
Where the system does not natively hold a fact the report needs, **the thing that applied the
|
||||
fact records it** — and:
|
||||
|
||||
> **A record of behaviour is written after the behaviour works, never when the resource is
|
||||
> created.**
|
||||
|
||||
Written up front it restates the request in a new place and inherits none of the authority of
|
||||
having happened. A failed run leaves its wreckage standing, and a report of that wreckage must
|
||||
not describe what the wreckage was supposed to be.
|
||||
|
||||
This is the production form of the mesh's most expensive fault: a firewall key declared in five
|
||||
manifests and read by no code
|
||||
([04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)). The
|
||||
declaration was never wrong. Nothing ever asked the system.
|
||||
|
||||
### Search the record before forming a hypothesis
|
||||
|
||||
The first action on any error message, failing service or unexpected behaviour is to search the
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-24
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0016-a-lab-node-is-a-virtual-machine.md
|
||||
---
|
||||
|
||||
# 33. A router is scenery, not a node — so it is a container
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual
|
||||
machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install,
|
||||
so it has to be a real machine or the thing under test is not the thing that ships.
|
||||
|
||||
A scenario also needs routers. NAT, port forwarding, policy between segments and mapping
|
||||
expiry are all things a router does, and until one is materialised a multi-segment scenario
|
||||
raises isolated islands
|
||||
([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)).
|
||||
The declaration already implies them: a gateway is *the one implicit machine in an otherwise
|
||||
explicit declaration*.
|
||||
|
||||
The question is whether ADR 0016 binds those too.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency
|
||||
nobody needs. A router boots in roughly ten seconds against a container's one; a
|
||||
six-segment scenario wanting three routers spends thirty seconds per raise on scenery.
|
||||
2. **The hypervisor provides NAT** — bridges with translation switched on, and its own
|
||||
forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide
|
||||
what the declaration is supposed to own, and it cannot express a mapping that expires, a
|
||||
gateway that refuses to forward, or policy between siblings. The model would shrink to fit
|
||||
the tool.
|
||||
3. **A router is scenery, and scenery is a container.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**ADR 0016 binds nodes. A router is not a node.**
|
||||
|
||||
Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh
|
||||
never installs anything on it, and no assertion is ever made about its internals. It exists so
|
||||
that packets between machines behave the way they behave in the world — which is the definition
|
||||
of scenery.
|
||||
|
||||
So a router is a **system container**, and the fidelity argument does not reach it: what a
|
||||
router must reproduce is kernel behaviour — translation, connection tracking, filtering,
|
||||
forwarding — and a container has the same kernel.
|
||||
|
||||
**Verified before deciding, not assumed.** In a plain unprivileged container:
|
||||
|
||||
| Needed for | Works |
|
||||
|---|---|
|
||||
| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` |
|
||||
| `nat:` | nftables masquerade, rules accepted and listed back |
|
||||
| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` |
|
||||
|
||||
No privileged mode, no nesting, no capability grants.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten,
|
||||
and a scenario's cost tracks the machines actually under test.
|
||||
- **The distinction is now load-bearing and has to stay legible.** *Node* means something under
|
||||
test; *scenery* means something that makes the test real. If anything is ever installed on a
|
||||
router by the mesh, it has become a node and this decision no longer covers it.
|
||||
- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra
|
||||
concept — justified by it being true, rather than by the saving.
|
||||
- A container shares the host kernel, so a scenario cannot reproduce a router running a
|
||||
*different* kernel from the workstation. Nothing currently wants that; if something does, that
|
||||
router becomes a virtual machine and this record needs revisiting rather than bending.
|
||||
- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and
|
||||
never names the machine that serves it — which is right, because it is not a machine the
|
||||
scenario has anything to say about.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged.
|
||||
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router,
|
||||
and why *published but behind NAT* only exists in production today.
|
||||
@@ -0,0 +1,91 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-24
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
---
|
||||
|
||||
# 34. A test defends a decision
|
||||
|
||||
## Context
|
||||
|
||||
[`how-we-build.md`](../00-META/how-we-build.md) §5 already says that **if a document states a
|
||||
rule about the mesh, it says how the rule is verified**, on the grounds that an unenforced rule
|
||||
is indistinguishable from a wrong one and costs more, because people believe it.
|
||||
|
||||
That rule is applied to prose and to acceptance criteria. It has never been applied to
|
||||
**decisions**, and it should be — a decision record states something that must be true, which
|
||||
is the same kind of claim.
|
||||
|
||||
The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and
|
||||
**no stated rule was broken.** There is no testing posture in `how-we-build.md` at all: no
|
||||
expectation, no gate, no definition of done. Every decision the lab embodies — the underlay
|
||||
boundary, the closed address space, routers as scenery, waiting for usable rather than for a
|
||||
call to return — was verified by hand, by running scenarios and reading output, and none of
|
||||
that survives the terminal it was run in.
|
||||
|
||||
Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has
|
||||
not built since 2026-06-04, and nothing said so
|
||||
([`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). Coverage
|
||||
assumed rather than checked.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **A coverage percentage.** Rejected. It measures how much code a test touched, not whether
|
||||
anything important is defended, and it is satisfied by tests that assert nothing. A number
|
||||
would have been met by testing the parser harder while the hypervisor integration stayed
|
||||
unasserted.
|
||||
2. **Test-driven development as a hard rule.** Rejected, and not because it is wrong in general.
|
||||
Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from
|
||||
stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time
|
||||
networking flushes a static address. A test written first against undiscovered behaviour
|
||||
asserts a guess.
|
||||
3. **A test defends a decision.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**Every decision record states something that must be true. A test asserts it.**
|
||||
|
||||
A decision with no test is a decision that will quietly stop being true, and nobody will find
|
||||
out from a document. Concretely:
|
||||
|
||||
- Where a decision is about **structure or logic** — what a declaration may say, what is
|
||||
refused, how a name is derived — the test is a unit test, and it is **written first**, because
|
||||
the behaviour is knowable before the code.
|
||||
- Where a decision is about **behaviour against a real system** — a hypervisor, a broker, a
|
||||
daemon — the test runs against the real thing, and is written **alongside**, because the
|
||||
behaviour is discovered rather than known.
|
||||
- **Mocking the boundary is forbidden.** A test that fakes a hypervisor asserts that the fake
|
||||
behaves as expected, which is the shape of test this whole effort exists to stop shipping.
|
||||
- **The gate is blocking, and green is the definition of done.** A change that has not run its
|
||||
tests is not finished, whatever its diff looks like.
|
||||
|
||||
A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable,
|
||||
so a decision without one can be *found* rather than noticed.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The question *"which tests matter"* has an answer that is not a number. The decisions are the
|
||||
list, and they are already written down.
|
||||
- **A new decision costs a test.** That is the intended friction — a decision nobody will assert
|
||||
is one worth reconsidering.
|
||||
- Integration tests need real infrastructure and are slow. That cost is accepted: a fast test
|
||||
suite that mocks the boundary would tell us nothing about the boundary, which is where every
|
||||
interesting fault in this session actually was.
|
||||
- Some decisions are not mechanically assertable — *the mesh brokers capabilities; nodes host;
|
||||
agents think* is a shape, not a predicate. Those should say so in the record rather than being
|
||||
quietly exempt, so the exemption is visible.
|
||||
- Records 0001–0033 were made before this rule. They are not retroactively invalid, but each
|
||||
should acquire a test or an explicit note that it cannot have one, and until then this rule
|
||||
is aspirational for them — which is exactly the state §5 warns about, recorded rather than
|
||||
hidden.
|
||||
|
||||
## References
|
||||
|
||||
- [`how-we-build.md`](../00-META/how-we-build.md) §5 — the rule this extends from prose to
|
||||
decisions.
|
||||
- [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md) — coverage
|
||||
assumed rather than checked, for two and a half months.
|
||||
- The sibling HQ repository for the PAPA platform states the same boundary rule — the contract
|
||||
is tested against the real system, mocking the client is forbidden, and a blocking gate is the
|
||||
definition of done. This record adopts that posture and adds the decision pairing.
|
||||
@@ -0,0 +1,100 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-24
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0034-a-test-defends-a-decision.md
|
||||
---
|
||||
|
||||
# 35. A picture of a system is read from the system, never from what asked for it
|
||||
|
||||
## Context
|
||||
|
||||
A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets.
|
||||
The two are supposed to correspond, and the entire value of the lab rests on noticing when
|
||||
they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a
|
||||
claim that will quietly stop being true.
|
||||
|
||||
Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not.
|
||||
A diagram of a running system can be produced two ways: parse the declaration and lay it out,
|
||||
or interrogate the system and lay *that* out. The first is far easier — the declaration is
|
||||
already parsed, already validated, already in memory.
|
||||
|
||||
It is also worthless for the only question worth asking of such a picture: *is what is running
|
||||
what I asked for?* A drawing built from the request and captioned **as raised** answers that
|
||||
question with the request, which always agrees with itself.
|
||||
|
||||
This is the same fault as
|
||||
[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a firewall key declared in
|
||||
five manifests and read by no code, so a manifest appears to restrict a port and restricts
|
||||
nothing. The declaration was never wrong. Nothing ever asked the system.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Draw the declaration, and label it honestly.** Cheap, and useful for review before
|
||||
anything is raised. Insufficient alone: it can never disagree with itself.
|
||||
2. **Draw the system, inferring the rest from the declaration where the system is silent.**
|
||||
The tempting middle. Rejected — a picture where some facts are observed and some are
|
||||
assumed has no honest caption, and the assumed ones are exactly the interesting ones.
|
||||
3. **Two pictures, one layout, neither borrowing from the other.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A picture captioned *as raised* reads only the running system.** It never opens the
|
||||
declaration, not even for a fact the system happens not to record.
|
||||
|
||||
Where the hypervisor does not natively hold a fact the picture needs — whether a segment is
|
||||
public, what a gateway translates, whether a machine refuses inbound — **the raise records it
|
||||
on the resource** as metadata, and the picture reads it back from there.
|
||||
|
||||
That recording carries its own rule, which is the substance of this decision rather than an
|
||||
implementation note:
|
||||
|
||||
> **A behavioural tag is written after the behaviour works, never when the resource is
|
||||
> created.**
|
||||
|
||||
Written at creation, a tag restates the request in a new location and inherits none of the
|
||||
authority of having happened. A failed raise deliberately leaves its wreckage standing, so a
|
||||
tag written up front would let a picture of that wreckage badge translation the gateway was
|
||||
never configured to do — reproducing, inside the tool built to catch the fault, exactly the
|
||||
fault.
|
||||
|
||||
So: the gateway is tagged after its ruleset applies; the machine after the read-back proves
|
||||
its firewall loaded.
|
||||
|
||||
Both pictures render through **one layout**, so they can be put side by side and the
|
||||
difference read off directly.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **It earned itself on the first comparison.** Drawn side by side, every virtual machine in
|
||||
the live picture held no addresses at all. A container's interface carries the name of the
|
||||
device it was configured as; a virtual machine names its own — so joining addresses to
|
||||
devices by name attached every address to a container and none to a VM. Nothing failed;
|
||||
a whole class of machine silently lost its addresses. The two pictures disagreed, so it was
|
||||
visible in seconds. It is now joined on MAC.
|
||||
- Raise does more work, and writes metadata it does not itself consume. Accepted: the cost is
|
||||
a few config keys, and it is what makes a raised instance self-describing.
|
||||
- A resource raised before a tag existed is missing it. The reader says so rather than filling
|
||||
the gap from the declaration — an untagged link draws as unknown, not as what the file said
|
||||
it should be.
|
||||
- **The rule generalises past diagrams.** Anything reporting on the mesh — a status view, an
|
||||
inventory, a health check — is subject to it. A report assembled from the intended state is
|
||||
not a report.
|
||||
- The declared picture stays, and stays useful: it is review before raising, and it is one half
|
||||
of the comparison. It carries no runtime status, because it cannot know any.
|
||||
|
||||
- **The constitution sync is not done and must not be.** `how-we-build.md` §5 carries this rule
|
||||
marked *proposed*, and playbook
|
||||
[05](../00-META/process/05-constitution-sync.md) publishes the derived page only for rules the
|
||||
mesh should enforce now. A rule the mesh enforces before a second person has agreed to it is
|
||||
the failure mode §6 exists to prevent. The sync happens when this record is accepted, and this
|
||||
line is what makes the gap visible rather than silent.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
|
||||
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the
|
||||
mesh is responsible for; the same instinct, applied to facts rather than to configuration.
|
||||
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production
|
||||
form.
|
||||
@@ -256,6 +256,11 @@ reachability, and it does not.
|
||||
The lab materialises a machine to be the gateway. That is the one implicit machine in an
|
||||
otherwise explicit declaration, and it exists because NAT has to run somewhere.
|
||||
|
||||
It is a **container, not a virtual machine** — a router is scenery rather than something under
|
||||
test, so the fidelity argument that makes a node a virtual machine does not reach it
|
||||
([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md)). What a router must
|
||||
reproduce is kernel behaviour, and a container has the same kernel.
|
||||
|
||||
**`machines[].at`** — segment and addresses, or a **list** of them for a machine on several
|
||||
segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node
|
||||
with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on
|
||||
|
||||
@@ -133,6 +133,12 @@ decision rather than a second implementation.
|
||||
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
||||
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||
decides whether a person can find the one they left standing yesterday.
|
||||
- ~~**Does a scenario snapshot need the machines stopped?**~~ **Answered by the integration
|
||||
test on its first run: no, but they must be flushed.** A snapshot captures disk and not
|
||||
memory, so a write still in the guest's page cache is absent from it — not stale, absent. A
|
||||
file written seconds before a snapshot did not survive the restore. Flushing first buys
|
||||
write-durability; it does not buy application-consistency, and anything mid-transaction is
|
||||
still captured mid-transaction.
|
||||
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||
the instance must not destroy them.
|
||||
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
||||
|
||||
Reference in New Issue
Block a user