From eab45984941dc6f3f5f0eddbc1666248456e1fdb Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 01:28:08 +0200 Subject: [PATCH 1/4] ADR 0033: a router is scenery, not a node MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ADR 0016 makes a lab node a virtual machine, and its reasoning is fidelity: a node boots a stock image and runs the real install, so it has to be a real machine or the thing under test is not the thing that ships. That reasoning does not reach a router. Nothing under test runs on one, it holds no identity, the mesh never installs anything on it, and no assertion is ever made about its internals. It exists so packets behave the way they behave in the world, which is the definition of scenery. So a router is a system container. What it must reproduce is kernel behaviour — translation, connection tracking, filtering, forwarding — and a container has the same kernel. Verified before deciding rather than assumed. In a plain unprivileged container: ip_forward and ipv6 forwarding both settable, nftables masquerade accepted and listed back, and the conntrack timeouts that mapping_ttl depends on both writable. No privileged mode, no nesting, no capability grants. Rejected letting the hypervisor provide NAT, on a stronger ground than speed: it makes the lab provide what the declaration is supposed to own, and it cannot express a mapping that expires, a gateway that refuses to forward, or policy between siblings. The model would shrink to fit the tool. The distinction is now load-bearing and has to stay legible: node means something under test, scenery means something that makes the test real. If the mesh ever installs anything on a router, it has become a node and this record no longer covers it. --- .../0033-a-router-is-scenery-not-a-node.md | 81 +++++++++++++++++++ 03-DESIGN/01-to-be/02-scenario-declaration.md | 5 ++ 2 files changed, 86 insertions(+) create mode 100644 02-DECISIONS/0033-a-router-is-scenery-not-a-node.md diff --git a/02-DECISIONS/0033-a-router-is-scenery-not-a-node.md b/02-DECISIONS/0033-a-router-is-scenery-not-a-node.md new file mode 100644 index 0000000..a98bb85 --- /dev/null +++ b/02-DECISIONS/0033-a-router-is-scenery-not-a-node.md @@ -0,0 +1,81 @@ +--- +status: accepted +date: 2026-08-24 +deciders: jochen +reconstructed: false +extends: 0016-a-lab-node-is-a-virtual-machine.md +--- + +# 33. A router is scenery, not a node — so it is a container + +## Context + +[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual +machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install, +so it has to be a real machine or the thing under test is not the thing that ships. + +A scenario also needs routers. NAT, port forwarding, policy between segments and mapping +expiry are all things a router does, and until one is materialised a multi-segment scenario +raises isolated islands +([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)). +The declaration already implies them: a gateway is *the one implicit machine in an otherwise +explicit declaration*. + +The question is whether ADR 0016 binds those too. + +## Considered options + +1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency + nobody needs. A router boots in roughly ten seconds against a container's one; a + six-segment scenario wanting three routers spends thirty seconds per raise on scenery. +2. **The hypervisor provides NAT** — bridges with translation switched on, and its own + forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide + what the declaration is supposed to own, and it cannot express a mapping that expires, a + gateway that refuses to forward, or policy between siblings. The model would shrink to fit + the tool. +3. **A router is scenery, and scenery is a container.** Chosen. + +## Decision + +**ADR 0016 binds nodes. A router is not a node.** + +Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh +never installs anything on it, and no assertion is ever made about its internals. It exists so +that packets between machines behave the way they behave in the world — which is the definition +of scenery. + +So a router is a **system container**, and the fidelity argument does not reach it: what a +router must reproduce is kernel behaviour — translation, connection tracking, filtering, +forwarding — and a container has the same kernel. + +**Verified before deciding, not assumed.** In a plain unprivileged container: + +| Needed for | Works | +|---|---| +| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` | +| `nat:` | nftables masquerade, rules accepted and listed back | +| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` | + +No privileged mode, no nesting, no capability grants. + +## Consequences + +- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten, + and a scenario's cost tracks the machines actually under test. +- **The distinction is now load-bearing and has to stay legible.** *Node* means something under + test; *scenery* means something that makes the test real. If anything is ever installed on a + router by the mesh, it has become a node and this decision no longer covers it. +- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra + concept — justified by it being true, rather than by the saving. +- A container shares the host kernel, so a scenario cannot reproduce a router running a + *different* kernel from the workstation. Nothing currently wants that; if something does, that + router becomes a virtual machine and this record needs revisiting rather than bending. +- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and + never names the machine that serves it — which is right, because it is not a machine the + scenario has anything to say about. + +## References + +- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged. +- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router, + and why *published but behind NAT* only exists in production today. diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 370be64..4dd19d3 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -256,6 +256,11 @@ reachability, and it does not. The lab materialises a machine to be the gateway. That is the one implicit machine in an otherwise explicit declaration, and it exists because NAT has to run somewhere. +It is a **container, not a virtual machine** — a router is scenery rather than something under +test, so the fidelity argument that makes a node a virtual machine does not reach it +([ADR 0033](../../02-DECISIONS/0033-a-router-is-scenery-not-a-node.md)). What a router must +reproduce is kernel behaviour, and a container has the same kernel. + **`machines[].at`** — segment and addresses, or a **list** of them for a machine on several segments at once. Multi-homing is not exotic: it is what a border machine is, and what any node with both a LAN and a WAN interface is. Each entry carries the addresses that machine holds on -- 2.54.0 From a2495e4d8ee79e7b8b39b5cc9bc685eb236d60e0 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 22:22:19 +0200 Subject: [PATCH 2/4] ADR 0034 (proposed): a test defends a decision MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit how-we-build 5 already says that if a document states a rule about the mesh, it says how the rule is verified — an unenforced rule being indistinguishable from a wrong one, and costing more because people believe it. That has never been applied to decisions, and a decision record states the same kind of claim. The gap was found by review: the lab reached 2,128 lines with 1,072 untested and no stated rule broken, because there is no testing posture in how-we-build at all. Every decision the lab embodies was verified by hand and none of it survives the terminal it ran in — which is 04-ISSUES/005 in miniature, coverage assumed rather than checked. Rejected a coverage percentage: it measures how much code a test touched, not whether anything important is defended, and would have been satisfied by testing the parser harder while the hypervisor integration stayed unasserted. Rejected test-driven development as a hard rule, and not because it is wrong in general. Half this implementation was discovery — that the hypervisor CLI reads a definition from stdin and hangs, that it assigns a MAC without recording it, that a stock image's networking flushes a static address. A test written first against undiscovered behaviour asserts a guess. So: structure and logic tested first, behaviour against a real system tested alongside, mocking the boundary forbidden, and a blocking gate as the definition of done. A test names the decision it defends, which is what makes the pairing checkable — a decision without one can be found rather than noticed. Stated as proposed rather than adopted: 6 requires review by someone who is not the proposer. Records 0001-0033 predate it and are not retroactively invalid, but each should acquire a test or an explicit note that it cannot have one, and until then the rule is aspirational for them — which is the state 5 warns about, recorded rather than hidden. --- 00-META/how-we-build.md | 20 ++++ .../0034-a-test-defends-a-decision.md | 91 +++++++++++++++++++ 2 files changed, 111 insertions(+) create mode 100644 02-DECISIONS/0034-a-test-defends-a-decision.md diff --git a/00-META/how-we-build.md b/00-META/how-we-build.md index 9926fd4..a353a70 100644 --- a/00-META/how-we-build.md +++ b/00-META/how-we-build.md @@ -152,6 +152,26 @@ finds one names the evidence required and returns the work. The reason is the mesh's most consistent failure shape: a green result proves transport, not effect. Absence reads as success unless something looked. +### A test defends a decision + +*Proposed — [ADR 0034](../02-DECISIONS/0034-a-test-defends-a-decision.md), pending review.* + +The rule above applies to prose. It applies to **decisions** too: a decision record states +something that must be true, and a test asserts it. A decision with no test is one that will +quietly stop being true, and nobody will learn that from a document. + +- Structure and logic — what is accepted, what is refused, how a value is derived — is tested + **first**, because the behaviour is knowable before the code. +- Behaviour against a real system is tested **alongside**, because it is discovered rather than + known. +- **Mocking the boundary is forbidden.** A test that fakes the system under integration asserts + that the fake behaves as expected. +- **The gate is blocking.** Green is the definition of done; a change that has not run its tests + is not finished, whatever the diff looks like. + +Not test-driven development as a blanket rule — a test written first against undiscovered +behaviour asserts a guess. The obligation is that every decision has a defender. + ### Search the record before forming a hypothesis The first action on any error message, failing service or unexpected behaviour is to search the diff --git a/02-DECISIONS/0034-a-test-defends-a-decision.md b/02-DECISIONS/0034-a-test-defends-a-decision.md new file mode 100644 index 0000000..93ef48f --- /dev/null +++ b/02-DECISIONS/0034-a-test-defends-a-decision.md @@ -0,0 +1,91 @@ +--- +status: proposed +date: 2026-08-24 +deciders: jochen +reconstructed: false +--- + +# 34. A test defends a decision + +## Context + +[`how-we-build.md`](../00-META/how-we-build.md) §5 already says that **if a document states a +rule about the mesh, it says how the rule is verified**, on the grounds that an unenforced rule +is indistinguishable from a wrong one and costs more, because people believe it. + +That rule is applied to prose and to acceptance criteria. It has never been applied to +**decisions**, and it should be — a decision record states something that must be true, which +is the same kind of claim. + +The gap was found by review. The lab reached 2,128 lines with 1,072 of them untested, and +**no stated rule was broken.** There is no testing posture in `how-we-build.md` at all: no +expectation, no gate, no definition of done. Every decision the lab embodies — the underlay +boundary, the closed address space, routers as scenery, waiting for usable rather than for a +call to return — was verified by hand, by running scenarios and reading output, and none of +that survives the terminal it was run in. + +Which is the fault the mesh already has catalogued at scale: an end-to-end harness that has +not built since 2026-06-04, and nothing said so +([`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)). Coverage +assumed rather than checked. + +## Considered options + +1. **A coverage percentage.** Rejected. It measures how much code a test touched, not whether + anything important is defended, and it is satisfied by tests that assert nothing. A number + would have been met by testing the parser harder while the hypervisor integration stayed + unasserted. +2. **Test-driven development as a hard rule.** Rejected, and not because it is wrong in general. + Half the lab's implementation was discovery: that the hypervisor CLI reads a definition from + stdin and hangs, that it assigns a MAC without recording it, that a stock image's boot-time + networking flushes a static address. A test written first against undiscovered behaviour + asserts a guess. +3. **A test defends a decision.** Chosen. + +## Decision + +**Every decision record states something that must be true. A test asserts it.** + +A decision with no test is a decision that will quietly stop being true, and nobody will find +out from a document. Concretely: + +- Where a decision is about **structure or logic** — what a declaration may say, what is + refused, how a name is derived — the test is a unit test, and it is **written first**, because + the behaviour is knowable before the code. +- Where a decision is about **behaviour against a real system** — a hypervisor, a broker, a + daemon — the test runs against the real thing, and is written **alongside**, because the + behaviour is discovered rather than known. +- **Mocking the boundary is forbidden.** A test that fakes a hypervisor asserts that the fake + behaves as expected, which is the shape of test this whole effort exists to stop shipping. +- **The gate is blocking, and green is the definition of done.** A change that has not run its + tests is not finished, whatever its diff looks like. + +A test names the decision it defends. Not as ceremony: it is what makes the pairing checkable, +so a decision without one can be *found* rather than noticed. + +## Consequences + +- The question *"which tests matter"* has an answer that is not a number. The decisions are the + list, and they are already written down. +- **A new decision costs a test.** That is the intended friction — a decision nobody will assert + is one worth reconsidering. +- Integration tests need real infrastructure and are slow. That cost is accepted: a fast test + suite that mocks the boundary would tell us nothing about the boundary, which is where every + interesting fault in this session actually was. +- Some decisions are not mechanically assertable — *the mesh brokers capabilities; nodes host; + agents think* is a shape, not a predicate. Those should say so in the record rather than being + quietly exempt, so the exemption is visible. +- Records 0001–0033 were made before this rule. They are not retroactively invalid, but each + should acquire a test or an explicit note that it cannot have one, and until then this rule + is aspirational for them — which is exactly the state §5 warns about, recorded rather than + hidden. + +## References + +- [`how-we-build.md`](../00-META/how-we-build.md) §5 — the rule this extends from prose to + decisions. +- [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md) — coverage + assumed rather than checked, for two and a half months. +- The sibling HQ repository for the PAPA platform states the same boundary rule — the contract + is tested against the real system, mocking the client is forbidden, and a blocking gate is the + definition of done. This record adopts that posture and adds the decision pairing. -- 2.54.0 From 8efa063f210cee292793e7eebcfb71c476599e34 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 22:26:47 +0200 Subject: [PATCH 3/4] The snapshot question is answered by a test MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The lifecycle design asked whether a scenario snapshot needs the machines stopped. The integration test answered it on its first run: no, but they must be flushed. A snapshot captures disk and not memory, so a write still in the guest's page cache is absent from it — not stale, absent. A file written seconds before a snapshot did not survive the restore. Flushing first buys write-durability. It does not buy application-consistency: anything mid-transaction is still captured mid-transaction, and that limit is now stated rather than left implied. --- 03-DESIGN/01-to-be/03-scenario-lifecycle.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md index 8a06df0..e0df128 100644 --- a/03-DESIGN/01-to-be/03-scenario-lifecycle.md +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -133,6 +133,12 @@ decision rather than a second implementation. [research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md). - **Instance naming.** A declaration is a kind and instances are many; how they are named decides whether a person can find the one they left standing yesterday. +- ~~**Does a scenario snapshot need the machines stopped?**~~ **Answered by the integration + test on its first run: no, but they must be flushed.** A snapshot captures disk and not + memory, so a write still in the guest's page cache is absent from it — not stale, absent. A + file written seconds before a snapshot did not survive the restore. Flushing first buys + write-durability; it does not buy application-consistency, and anything mid-transaction is + still captured mid-transaction. - **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying the instance must not destroy them. - **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and -- 2.54.0 From b225b0762541ec11f7f500e825445e86712aecde Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 24 Aug 2026 22:54:30 +0200 Subject: [PATCH 4/4] ADR 0035 (proposed): a picture is read from what runs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Drawing a scenario forced a choice that looks cosmetic and is not. A diagram built from the declaration and captioned "as raised" answers "is what is running what I asked for?" with the request, which always agrees with itself. So: a picture captioned as raised reads only the running system, and where the hypervisor does not hold a fact the picture needs, the raise records it on the resource. With the rule that makes the recording worth anything — a behavioural tag is written after the behaviour works, never at creation, because a failed raise leaves wreckage standing and a picture of wreckage must not badge what the wreckage was supposed to be. It earned itself on the first comparison: every VM showed no addresses, because a container's interface carries the device's name and a VM names its own. The two pictures disagreed, so a whole class of machine silently losing its addresses was visible in seconds. §5 of how-we-build gains the general form, marked proposed. The constitution sync is deliberately NOT done — a rule the mesh enforces before a second person agreed to it is what §6 exists to prevent. --- 00-META/how-we-build.md | 26 +++++ .../0035-a-picture-is-read-from-what-runs.md | 100 ++++++++++++++++++ 2 files changed, 126 insertions(+) create mode 100644 02-DECISIONS/0035-a-picture-is-read-from-what-runs.md diff --git a/00-META/how-we-build.md b/00-META/how-we-build.md index a353a70..696a941 100644 --- a/00-META/how-we-build.md +++ b/00-META/how-we-build.md @@ -172,6 +172,32 @@ quietly stop being true, and nobody will learn that from a document. Not test-driven development as a blanket rule — a test written first against undiscovered behaviour asserts a guess. The obligation is that every decision has a defender. +### A report is read from the system, never from what asked for it + +*Proposed — [ADR 0035](../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md), pending +review.* + +The same rule as the two above, pointed at reporting rather than at verification. **Anything +that describes the state of the mesh — a status view, an inventory, a diagram, a health +check — is assembled from the running system.** Assembling it from the intended state produces +a report that always agrees with itself and can never disagree with reality, which is not a +report. + +Where the system does not natively hold a fact the report needs, **the thing that applied the +fact records it** — and: + +> **A record of behaviour is written after the behaviour works, never when the resource is +> created.** + +Written up front it restates the request in a new place and inherits none of the authority of +having happened. A failed run leaves its wreckage standing, and a report of that wreckage must +not describe what the wreckage was supposed to be. + +This is the production form of the mesh's most expensive fault: a firewall key declared in five +manifests and read by no code +([04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)). The +declaration was never wrong. Nothing ever asked the system. + ### Search the record before forming a hypothesis The first action on any error message, failing service or unexpected behaviour is to search the diff --git a/02-DECISIONS/0035-a-picture-is-read-from-what-runs.md b/02-DECISIONS/0035-a-picture-is-read-from-what-runs.md new file mode 100644 index 0000000..91ce780 --- /dev/null +++ b/02-DECISIONS/0035-a-picture-is-read-from-what-runs.md @@ -0,0 +1,100 @@ +--- +status: proposed +date: 2026-08-24 +deciders: jochen +reconstructed: false +extends: 0034-a-test-defends-a-decision.md +--- + +# 35. A picture of a system is read from the system, never from what asked for it + +## Context + +A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets. +The two are supposed to correspond, and the entire value of the lab rests on noticing when +they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a +claim that will quietly stop being true. + +Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not. +A diagram of a running system can be produced two ways: parse the declaration and lay it out, +or interrogate the system and lay *that* out. The first is far easier — the declaration is +already parsed, already validated, already in memory. + +It is also worthless for the only question worth asking of such a picture: *is what is running +what I asked for?* A drawing built from the request and captioned **as raised** answers that +question with the request, which always agrees with itself. + +This is the same fault as +[04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — a firewall key declared in +five manifests and read by no code, so a manifest appears to restrict a port and restricts +nothing. The declaration was never wrong. Nothing ever asked the system. + +## Considered options + +1. **Draw the declaration, and label it honestly.** Cheap, and useful for review before + anything is raised. Insufficient alone: it can never disagree with itself. +2. **Draw the system, inferring the rest from the declaration where the system is silent.** + The tempting middle. Rejected — a picture where some facts are observed and some are + assumed has no honest caption, and the assumed ones are exactly the interesting ones. +3. **Two pictures, one layout, neither borrowing from the other.** Chosen. + +## Decision + +**A picture captioned *as raised* reads only the running system.** It never opens the +declaration, not even for a fact the system happens not to record. + +Where the hypervisor does not natively hold a fact the picture needs — whether a segment is +public, what a gateway translates, whether a machine refuses inbound — **the raise records it +on the resource** as metadata, and the picture reads it back from there. + +That recording carries its own rule, which is the substance of this decision rather than an +implementation note: + +> **A behavioural tag is written after the behaviour works, never when the resource is +> created.** + +Written at creation, a tag restates the request in a new location and inherits none of the +authority of having happened. A failed raise deliberately leaves its wreckage standing, so a +tag written up front would let a picture of that wreckage badge translation the gateway was +never configured to do — reproducing, inside the tool built to catch the fault, exactly the +fault. + +So: the gateway is tagged after its ruleset applies; the machine after the read-back proves +its firewall loaded. + +Both pictures render through **one layout**, so they can be put side by side and the +difference read off directly. + +## Consequences + +- **It earned itself on the first comparison.** Drawn side by side, every virtual machine in + the live picture held no addresses at all. A container's interface carries the name of the + device it was configured as; a virtual machine names its own — so joining addresses to + devices by name attached every address to a container and none to a VM. Nothing failed; + a whole class of machine silently lost its addresses. The two pictures disagreed, so it was + visible in seconds. It is now joined on MAC. +- Raise does more work, and writes metadata it does not itself consume. Accepted: the cost is + a few config keys, and it is what makes a raised instance self-describing. +- A resource raised before a tag existed is missing it. The reader says so rather than filling + the gap from the declaration — an untagged link draws as unknown, not as what the file said + it should be. +- **The rule generalises past diagrams.** Anything reporting on the mesh — a status view, an + inventory, a health check — is subject to it. A report assembled from the intended state is + not a report. +- The declared picture stays, and stays useful: it is review before raising, and it is one half + of the comparison. It carries no runtime status, because it cannot know any. + +- **The constitution sync is not done and must not be.** `how-we-build.md` §5 carries this rule + marked *proposed*, and playbook + [05](../00-META/process/05-constitution-sync.md) publishes the derived page only for rules the + mesh should enforce now. A rule the mesh enforces before a second person has agreed to it is + the failure mode §6 exists to prevent. The sync happens when this record is accepted, and this + line is what makes the gap visible rather than silent. + +## References + +- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true. +- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the + mesh is responsible for; the same instinct, applied to facts rather than to configuration. +- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production + form. -- 2.54.0