diff --git a/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md new file mode 100644 index 0000000..09dac33 --- /dev/null +++ b/02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md @@ -0,0 +1,71 @@ +--- +status: accepted +date: 2026-08-23 +deciders: jochen +reconstructed: false +--- + +# 32. A scenario is an isolated address space, and the lab never reaches into it over IP + +## Context + +A scenario declares literal addresses — +[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has +to be, because reproducing *published but behind NAT* means saying which address the world sees. + +That raises a question the declaration left open: **two scenarios at once.** Several agents +working means several scenarios, and the lab design already calls that a requirement. But two +scenarios built from the same declaration want the same addresses, and there are only three +documentation ranges in existence. + +## Considered options + +1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals. + Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a + specific topology no longer reproduces it; it breaks the RFC-range validation, since + allocated addresses would have to come from somewhere real; and the numbers a person reads + in the file stop being the numbers they will see in a capture. +2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate + an agent has to queue for is a gate that gets bypassed. +3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen. + +## Decision + +**A scenario is a closed address space.** Every segment materialises as its own isolated link, +belonging to one scenario instance. Two scenarios raised from the same declaration hold the same +addresses and never meet, because nothing joins their links. + +The declaration therefore keeps its literal addresses, and they mean exactly what they say. + +**The consequence that constrains everything else: the lab never reaches into a scenario over +IP.** It talks to a machine through the virtualisation layer's own channel — the same way one +executes a command in a container without the container being routable. + +That is not a preference. If the lab reached machines by address, the workstation running it +would need a route into each scenario, and two scenarios carrying the same prefix would give it +two routes to the same destination. Concurrency would be impossible, and it would fail in the +worst available way: not with an error, but by one scenario's traffic arriving in another. + +## Consequences + +- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit + beyond the machine's capacity. +- The three documentation ranges stop being a scarce resource. Every scenario may use all of + them, because no two scenarios share a link. +- **The lab cannot use IP to check anything**, which is more of a constraint than it first + appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by + executing on a machine, rather than probed from outside. That is the honest way to ask it + anyway — reachability from the workstation is not the question. +- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing + outside holds a reference into it. +- The lab needs a scenario **instance** identity distinct from the scenario name in the + declaration: the declaration is a kind, and several instances of one kind may exist. +- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a + developer wants to reach — a web interface, a database — needs an explicit, deliberate + forward out of the scenario, which is a feature rather than a gap: nothing leaks by default. + +## References + +- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses + this preserves. +- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes. diff --git a/03-DESIGN/01-to-be/02-scenario-declaration.md b/03-DESIGN/01-to-be/02-scenario-declaration.md index 695268d..370be64 100644 --- a/03-DESIGN/01-to-be/02-scenario-declaration.md +++ b/03-DESIGN/01-to-be/02-scenario-declaration.md @@ -634,10 +634,6 @@ is the one real absence, and it is exactly the double-NAT case. - **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from outside; afterwards from the mesh itself. The declaration should not have to care, which suggests a named source rather than a path. -- **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above - writes addresses absolutely. Whether a scenario carries literal addresses or a template the - lab allocates from decides whether two can run side by side — and there are only three - documentation ranges to go round. - **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be published through both. - **An address changing in place**, as a DHCP lease expiring under a machine that has not moved. diff --git a/03-DESIGN/01-to-be/03-scenario-lifecycle.md b/03-DESIGN/01-to-be/03-scenario-lifecycle.md new file mode 100644 index 0000000..3ca2623 --- /dev/null +++ b/03-DESIGN/01-to-be/03-scenario-lifecycle.md @@ -0,0 +1,134 @@ +--- +layer: to-be +status: designed +code: [mesh-lab] +updated: 2026-08-23 +decisions: + - 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md + - 02-DECISIONS/0031-the-lab-provides-the-underlay.md + - 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md +--- + +# Scenario lifecycle + +The first thing the lab must do, and the only thing it must do before anything else can be +written: **materialise a mesh, return it to a known state, and destroy it** +([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)). + +A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens +to one. + +## The verbs + +| Verb | Does | +|---|---| +| `raise` | materialise a declaration into a running scenario instance | +| `snapshot` | name the current state of the whole scenario | +| `restore` | return the whole scenario to a named state | +| `move` | change a machine's position while the scenario runs | +| `exec` | run something on a machine, and get its output | +| `destroy` | tear the instance down | + +Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition +cheap, and cheap repetition is what turns the bootstrap path into an inner development loop +rather than a ceremony. + +## Raising, in order + +The order is not arbitrary — each step needs the one before it to exist: + +1. **Segments.** Isolated links, one per declared segment, belonging to this instance and + joined to nothing outside it + ([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). +2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each + distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the + translation, forwarding and mapping-expiry the declaration asked for. +3. **Machines.** Each on its segments, holding its addresses. +4. **Policy.** Rules between segments, applied on the gateways that route between them. +5. **Placement.** Artifacts onto machines. +6. **Snapshot**, if the declaration named one. + +Raising is **convergent, not incremental**: raising an instance that already exists brings it to +the declared state rather than failing or duplicating. That is the same model the mesh itself +uses, and a lab that behaved differently from the thing it tests would be teaching the wrong +habit. + +## A failed raise leaves the wreckage + +A step that fails stops the raise +([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear +down**. + +Tearing down on failure destroys the only evidence of what went wrong, which is precisely +backwards: a scenario that failed to raise is more interesting than one that succeeded. The +instance stays, marked failed, with the step that failed named. + +The caller decides what happens next, and the two callers want different things +— the coordinator captures and destroys, a person opens a shell. That is the same +one-runner-two-callers split the lab design already makes, applied to failure. + +## Snapshots are whole-scenario + +A snapshot captures **every machine and the state of the network between them**, as one thing. +Restoring returns all of it. + +Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes — +what is assigned where, which grants exist, what has been delivered — so restoring one machine +to an earlier moment while its peers move on produces a mesh that has never existed and could +not. The faults found there would be artefacts of the lab. + +This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh +that has never seen a change, another of a mesh running the previous version, and a restore +between runs. + +## Moving a machine + +`move` changes a machine's position while the scenario runs: to another segment, with different +addresses, or to `detached`. + +It is the roaming case, and it is a **lifecycle** operation rather than a declaration because +the interesting part is the transition, not the destination. A mesh that forms correctly with a +node at home and correctly with it away may still fail to notice it moved. + +Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change +made after it, and returning undoes it like any other change. + +## Reaching in + +Everything the lab does to a machine goes through the virtualisation layer, never over IP +([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a +command on a machine and returns its output. + +This has one consequence worth stating plainly: **a reachability question is asked from inside**. +*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from +the workstation. The workstation is not on the scenario's network and its opinion of +reachability would be a different question with a misleadingly similar answer. + +Anything a person wants to open in a browser needs a deliberate forward out of the instance. +Nothing leaks by default. + +## The two callers + +The lab design already establishes that the runner serves the coordinator and a person, and +that anything only one of them can do will drift. Applied here: + +| | the coordinator | someone working on the mesh | +|---|---|---| +| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest | +| on failure | capture, then destroy | leave it, open a shell | + +Both use the same verbs. The difference is what happens after the verdict, which is a caller's +decision rather than a second implementation. + +## Open + +- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the + inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop + is slow, and everything above is theory. +- **Instance naming.** A declaration is a kind and instances are many; how they are named + decides whether a person can find the one they left standing yesterday. +- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying + the instance must not destroy them. +- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and + before the mesh builds itself that somewhere is outside it — the open question from + [research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md). diff --git a/03-DESIGN/01-to-be/README.md b/03-DESIGN/01-to-be/README.md index e46578d..70e82bf 100644 --- a/03-DESIGN/01-to-be/README.md +++ b/03-DESIGN/01-to-be/README.md @@ -12,6 +12,7 @@ document is written and this one's status becomes `implemented`. | [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) | | [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) | | [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) | +| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) | ## Not yet written