From a253afe020b9f9265f99db917fe7928b981e4f33 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 23 Aug 2026 23:53:00 +0200 Subject: [PATCH] Scenario lifecycle, and how two scenarios coexist MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ADR 0032: a scenario is a closed address space. Every segment materialises as its own isolated link belonging to one instance, so two scenarios raised from the same declaration hold the same addresses and never meet. The declaration keeps its literal addresses and they mean what they say — allocating from a pool would have made them a fiction, so a scenario reproducing a specific topology would stop reproducing it. The constraint that follows shapes everything: the lab never reaches into a scenario over IP. It talks to machines through the virtualisation layer's own channel. If it reached them by address, the workstation would need a route into each scenario, and two carrying the same prefix would give it two routes to one destination — failing not with an error but by one scenario's traffic arriving in another. That also makes reachability an honest question. Can this machine reach that one is asked from INSIDE, by executing on the first, rather than probed from a workstation that is not on the network and whose opinion would be a different question with a misleadingly similar answer. The lifecycle itself: six verbs, of which raise and destroy are enough to be useful and the rest are what make repetition cheap. Raising is convergent rather than incremental, because a lab behaving differently from the thing it tests teaches the wrong habit. A failed raise leaves the wreckage standing. Tearing down on failure destroys the only evidence, which is backwards — a scenario that failed to raise is more interesting than one that succeeded. Snapshots are whole-scenario. Per-machine would be cheaper and wrong: the mesh keeps state spanning nodes, so restoring one machine while its peers move on produces a mesh that has never existed, and faults found there would be artefacts of the lab. Closes the declaration's open question about running several scenarios at once. --- ...a-scenario-is-an-isolated-address-space.md | 71 ++++++++++ 03-DESIGN/01-to-be/02-scenario-declaration.md | 4 - 03-DESIGN/01-to-be/03-scenario-lifecycle.md | 134 ++++++++++++++++++ 03-DESIGN/01-to-be/README.md | 1 + 4 files changed, 206 insertions(+), 4 deletions(-) create mode 100644 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md create mode 100644 03-DESIGN/01-to-be/03-scenario-lifecycle.md diff --git a/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md new file mode 100644 index 0000000..09dac33 --- /dev/null +++ b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md @@ -0,0 +1,71 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +--- + +# 32. A scenario is an isolated address space, and the lab never reaches into it over IP + +## Context + +A scenario declares literal addresses — +[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has +to be, because reproducing *published but behind NAT* means saying which address the world sees. + +That raises a question the declaration left open: **two scenarios at once.** Several agents +working means several scenarios, and the lab design already calls that a requirement. But two +scenarios built from the same declaration want the same addresses, and there are only three +documentation ranges in existence. + +## Considered options + +1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals. + Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a + specific topology no longer reproduces it; it breaks the RFC-range validation, since + allocated addresses would have to come from somewhere real; and the numbers a person reads + in the file stop being the numbers they will see in a capture. +2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate + an agent has to queue for is a gate that gets bypassed. +3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen. + +## Decision + +**A scenario is a closed address space.** Every segment materialises as its own isolated link, +belonging to one scenario instance. Two scenarios raised from the same declaration hold the same +addresses and never meet, because nothing joins their links. + +The declaration therefore keeps its literal addresses, and they mean exactly what they say. + +**The consequence that constrains everything else: the lab never reaches into a scenario over +IP.** It talks to a machine through the virtualisation layer's own channel — the same way one +executes a command in a container without the container being routable. + +That is not a preference. If the lab reached machines by address, the workstation running it +would need a route into each scenario, and two scenarios carrying the same prefix would give it +two routes to the same destination. Concurrency would be impossible, and it would fail in the +worst available way: not with an error, but by one scenario's traffic arriving in another. + +## Consequences + +- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit + beyond the machine's capacity. +- The three documentation ranges stop being a scarce resource. Every scenario may use all of + them, because no two scenarios share a link. +- **The lab cannot use IP to check anything**, which is more of a constraint than it first + appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by + executing on a machine, rather than probed from outside. That is the honest way to ask it + anyway — reachability from the workstation is not the question. +- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing + outside holds a reference into it. +- The lab needs a scenario **instance** identity distinct from the scenario name in the + declaration: the declaration is a kind, and several instances of one kind may exist. +- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a + developer wants to reach — a web interface, a database — needs an explicit, deliberate + forward out of the scenario, which is a feature rather than a gap: nothing leaks by default. + +## References + +- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses + this preserves. +- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes. diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 695268d..370be64 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -634,10 +634,6 @@ is the one real absence, and it is exactly the double-NAT case. - **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from outside; afterwards from the mesh itself. The declaration should not have to care, which suggests a named source rather than a path. -- **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above - writes addresses absolutely. Whether a scenario carries literal addresses or a template the - lab allocates from decides whether two can run side by side — and there are only three - documentation ranges to go round. - **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be published through both. - **An address changing in place**, as a DHCP lease expiring under a machine that has not moved. diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md new file mode 100644 index 0000000..3ca2623 --- /dev/null +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -0,0 +1,134 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-23 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0031-the-lab-provides-the-underlay.md + - 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md +--- + +# Scenario lifecycle + +The first thing the lab must do, and the only thing it must do before anything else can be +written: **materialise a mesh, return it to a known state, and destroy it** +([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens +to one. + +## The verbs + +| Verb | Does | +|---|---| +| `raise` | materialise a declaration into a running scenario instance | +| `snapshot` | name the current state of the whole scenario | +| `restore` | return the whole scenario to a named state | +| `move` | change a machine's position while the scenario runs | +| `exec` | run something on a machine, and get its output | +| `destroy` | tear the instance down | + +Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition +cheap, and cheap repetition is what turns the bootstrap path into an inner development loop +rather than a ceremony. + +## Raising, in order + +The order is not arbitrary — each step needs the one before it to exist: + +1. **Segments.** Isolated links, one per declared segment, belonging to this instance and + joined to nothing outside it + ([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). +2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each + distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the + translation, forwarding and mapping-expiry the declaration asked for. +3. **Machines.** Each on its segments, holding its addresses. +4. **Policy.** Rules between segments, applied on the gateways that route between them. +5. **Placement.** Artifacts onto machines. +6. **Snapshot**, if the declaration named one. + +Raising is **convergent, not incremental**: raising an instance that already exists brings it to +the declared state rather than failing or duplicating. That is the same model the mesh itself +uses, and a lab that behaved differently from the thing it tests would be teaching the wrong +habit. + +## A failed raise leaves the wreckage + +A step that fails stops the raise +([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear +down**. + +Tearing down on failure destroys the only evidence of what went wrong, which is precisely +backwards: a scenario that failed to raise is more interesting than one that succeeded. The +instance stays, marked failed, with the step that failed named. + +The caller decides what happens next, and the two callers want different things +— the coordinator captures and destroys, a person opens a shell. That is the same +one-runner-two-callers split the lab design already makes, applied to failure. + +## Snapshots are whole-scenario + +A snapshot captures **every machine and the state of the network between them**, as one thing. +Restoring returns all of it. + +Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes — +what is assigned where, which grants exist, what has been delivered — so restoring one machine +to an earlier moment while its peers move on produces a mesh that has never existed and could +not. The faults found there would be artefacts of the lab. + +This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh +that has never seen a change, another of a mesh running the previous version, and a restore +between runs. + +## Moving a machine + +`move` changes a machine's position while the scenario runs: to another segment, with different +addresses, or to `detached`. + +It is the roaming case, and it is a **lifecycle** operation rather than a declaration because +the interesting part is the transition, not the destination. A mesh that forms correctly with a +node at home and correctly with it away may still fail to notice it moved. + +Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change +made after it, and returning undoes it like any other change. + +## Reaching in + +Everything the lab does to a machine goes through the virtualisation layer, never over IP +([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a +command on a machine and returns its output. + +This has one consequence worth stating plainly: **a reachability question is asked from inside**. +*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from +the workstation. The workstation is not on the scenario's network and its opinion of +reachability would be a different question with a misleadingly similar answer. + +Anything a person wants to open in a browser needs a deliberate forward out of the instance. +Nothing leaks by default. + +## The two callers + +The lab design already establishes that the runner serves the coordinator and a person, and +that anything only one of them can do will drift. Applied here: + +| | the coordinator | someone working on the mesh | +|---|---|---| +| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest | +| on failure | capture, then destroy | leave it, open a shell | + +Both use the same verbs. The difference is what happens after the verdict, which is a caller's +decision rather than a second implementation. + +## Open + +- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the + inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop + is slow, and everything above is theory. +- **Instance naming.** A declaration is a kind and instances are many; how they are named + decides whether a person can find the one they left standing yesterday. +- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying + the instance must not destroy them. +- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and + before the mesh builds itself that somewhere is outside it — the open question from + [research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md). diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index e46578d..70e82bf 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -12,6 +12,7 @@ document is written and this one's status becomes `implemented`. | [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) | | [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | | [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | +| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | ## Not yet written