The scenario model, the lifecycle, and what the lab actually costs #6
@@ -0,0 +1,71 @@
|
|||||||
|
---
|
||||||
|
status: accepted
|
||||||
|
date: 2026-08-23
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
---
|
||||||
|
|
||||||
|
# 32. A scenario is an isolated address space, and the lab never reaches into it over IP
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
A scenario declares literal addresses —
|
||||||
|
[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has
|
||||||
|
to be, because reproducing *published but behind NAT* means saying which address the world sees.
|
||||||
|
|
||||||
|
That raises a question the declaration left open: **two scenarios at once.** Several agents
|
||||||
|
working means several scenarios, and the lab design already calls that a requirement. But two
|
||||||
|
scenarios built from the same declaration want the same addresses, and there are only three
|
||||||
|
documentation ranges in existence.
|
||||||
|
|
||||||
|
## Considered options
|
||||||
|
|
||||||
|
1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals.
|
||||||
|
Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a
|
||||||
|
specific topology no longer reproduces it; it breaks the RFC-range validation, since
|
||||||
|
allocated addresses would have to come from somewhere real; and the numbers a person reads
|
||||||
|
in the file stop being the numbers they will see in a capture.
|
||||||
|
2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate
|
||||||
|
an agent has to queue for is a gate that gets bypassed.
|
||||||
|
3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**A scenario is a closed address space.** Every segment materialises as its own isolated link,
|
||||||
|
belonging to one scenario instance. Two scenarios raised from the same declaration hold the same
|
||||||
|
addresses and never meet, because nothing joins their links.
|
||||||
|
|
||||||
|
The declaration therefore keeps its literal addresses, and they mean exactly what they say.
|
||||||
|
|
||||||
|
**The consequence that constrains everything else: the lab never reaches into a scenario over
|
||||||
|
IP.** It talks to a machine through the virtualisation layer's own channel — the same way one
|
||||||
|
executes a command in a container without the container being routable.
|
||||||
|
|
||||||
|
That is not a preference. If the lab reached machines by address, the workstation running it
|
||||||
|
would need a route into each scenario, and two scenarios carrying the same prefix would give it
|
||||||
|
two routes to the same destination. Concurrency would be impossible, and it would fail in the
|
||||||
|
worst available way: not with an error, but by one scenario's traffic arriving in another.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit
|
||||||
|
beyond the machine's capacity.
|
||||||
|
- The three documentation ranges stop being a scarce resource. Every scenario may use all of
|
||||||
|
them, because no two scenarios share a link.
|
||||||
|
- **The lab cannot use IP to check anything**, which is more of a constraint than it first
|
||||||
|
appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by
|
||||||
|
executing on a machine, rather than probed from outside. That is the honest way to ask it
|
||||||
|
anyway — reachability from the workstation is not the question.
|
||||||
|
- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing
|
||||||
|
outside holds a reference into it.
|
||||||
|
- The lab needs a scenario **instance** identity distinct from the scenario name in the
|
||||||
|
declaration: the declaration is a kind, and several instances of one kind may exist.
|
||||||
|
- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a
|
||||||
|
developer wants to reach — a web interface, a database — needs an explicit, deliberate
|
||||||
|
forward out of the scenario, which is a feature rather than a gap: nothing leaks by default.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses
|
||||||
|
this preserves.
|
||||||
|
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes.
|
||||||
@@ -634,10 +634,6 @@ is the one real absence, and it is exactly the double-NAT case.
|
|||||||
- **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from
|
- **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from
|
||||||
outside; afterwards from the mesh itself. The declaration should not have to care, which
|
outside; afterwards from the mesh itself. The declaration should not have to care, which
|
||||||
suggests a named source rather than a path.
|
suggests a named source rather than a path.
|
||||||
- **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above
|
|
||||||
writes addresses absolutely. Whether a scenario carries literal addresses or a template the
|
|
||||||
lab allocates from decides whether two can run side by side — and there are only three
|
|
||||||
documentation ranges to go round.
|
|
||||||
- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be
|
- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be
|
||||||
published through both.
|
published through both.
|
||||||
- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved.
|
- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved.
|
||||||
|
|||||||
@@ -0,0 +1,134 @@
|
|||||||
|
---
|
||||||
|
layer: to-be
|
||||||
|
status: designed
|
||||||
|
code: [mesh-lab]
|
||||||
|
updated: 2026-08-23
|
||||||
|
decisions:
|
||||||
|
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||||
|
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
|
||||||
|
- 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# Scenario lifecycle
|
||||||
|
|
||||||
|
The first thing the lab must do, and the only thing it must do before anything else can be
|
||||||
|
written: **materialise a mesh, return it to a known state, and destroy it**
|
||||||
|
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
||||||
|
|
||||||
|
A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens
|
||||||
|
to one.
|
||||||
|
|
||||||
|
## The verbs
|
||||||
|
|
||||||
|
| Verb | Does |
|
||||||
|
|---|---|
|
||||||
|
| `raise` | materialise a declaration into a running scenario instance |
|
||||||
|
| `snapshot` | name the current state of the whole scenario |
|
||||||
|
| `restore` | return the whole scenario to a named state |
|
||||||
|
| `move` | change a machine's position while the scenario runs |
|
||||||
|
| `exec` | run something on a machine, and get its output |
|
||||||
|
| `destroy` | tear the instance down |
|
||||||
|
|
||||||
|
Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition
|
||||||
|
cheap, and cheap repetition is what turns the bootstrap path into an inner development loop
|
||||||
|
rather than a ceremony.
|
||||||
|
|
||||||
|
## Raising, in order
|
||||||
|
|
||||||
|
The order is not arbitrary — each step needs the one before it to exist:
|
||||||
|
|
||||||
|
1. **Segments.** Isolated links, one per declared segment, belonging to this instance and
|
||||||
|
joined to nothing outside it
|
||||||
|
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)).
|
||||||
|
2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each
|
||||||
|
distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the
|
||||||
|
translation, forwarding and mapping-expiry the declaration asked for.
|
||||||
|
3. **Machines.** Each on its segments, holding its addresses.
|
||||||
|
4. **Policy.** Rules between segments, applied on the gateways that route between them.
|
||||||
|
5. **Placement.** Artifacts onto machines.
|
||||||
|
6. **Snapshot**, if the declaration named one.
|
||||||
|
|
||||||
|
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
||||||
|
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
||||||
|
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
||||||
|
habit.
|
||||||
|
|
||||||
|
## A failed raise leaves the wreckage
|
||||||
|
|
||||||
|
A step that fails stops the raise
|
||||||
|
([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear
|
||||||
|
down**.
|
||||||
|
|
||||||
|
Tearing down on failure destroys the only evidence of what went wrong, which is precisely
|
||||||
|
backwards: a scenario that failed to raise is more interesting than one that succeeded. The
|
||||||
|
instance stays, marked failed, with the step that failed named.
|
||||||
|
|
||||||
|
The caller decides what happens next, and the two callers want different things
|
||||||
|
— the coordinator captures and destroys, a person opens a shell. That is the same
|
||||||
|
one-runner-two-callers split the lab design already makes, applied to failure.
|
||||||
|
|
||||||
|
## Snapshots are whole-scenario
|
||||||
|
|
||||||
|
A snapshot captures **every machine and the state of the network between them**, as one thing.
|
||||||
|
Restoring returns all of it.
|
||||||
|
|
||||||
|
Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes —
|
||||||
|
what is assigned where, which grants exist, what has been delivered — so restoring one machine
|
||||||
|
to an earlier moment while its peers move on produces a mesh that has never existed and could
|
||||||
|
not. The faults found there would be artefacts of the lab.
|
||||||
|
|
||||||
|
This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh
|
||||||
|
that has never seen a change, another of a mesh running the previous version, and a restore
|
||||||
|
between runs.
|
||||||
|
|
||||||
|
## Moving a machine
|
||||||
|
|
||||||
|
`move` changes a machine's position while the scenario runs: to another segment, with different
|
||||||
|
addresses, or to `detached`.
|
||||||
|
|
||||||
|
It is the roaming case, and it is a **lifecycle** operation rather than a declaration because
|
||||||
|
the interesting part is the transition, not the destination. A mesh that forms correctly with a
|
||||||
|
node at home and correctly with it away may still fail to notice it moved.
|
||||||
|
|
||||||
|
Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change
|
||||||
|
made after it, and returning undoes it like any other change.
|
||||||
|
|
||||||
|
## Reaching in
|
||||||
|
|
||||||
|
Everything the lab does to a machine goes through the virtualisation layer, never over IP
|
||||||
|
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a
|
||||||
|
command on a machine and returns its output.
|
||||||
|
|
||||||
|
This has one consequence worth stating plainly: **a reachability question is asked from inside**.
|
||||||
|
*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from
|
||||||
|
the workstation. The workstation is not on the scenario's network and its opinion of
|
||||||
|
reachability would be a different question with a misleadingly similar answer.
|
||||||
|
|
||||||
|
Anything a person wants to open in a browser needs a deliberate forward out of the instance.
|
||||||
|
Nothing leaks by default.
|
||||||
|
|
||||||
|
## The two callers
|
||||||
|
|
||||||
|
The lab design already establishes that the runner serves the coordinator and a person, and
|
||||||
|
that anything only one of them can do will drift. Applied here:
|
||||||
|
|
||||||
|
| | the coordinator | someone working on the mesh |
|
||||||
|
|---|---|---|
|
||||||
|
| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest |
|
||||||
|
| on failure | capture, then destroy | leave it, open a shell |
|
||||||
|
|
||||||
|
Both use the same verbs. The difference is what happens after the verdict, which is a caller's
|
||||||
|
decision rather than a second implementation.
|
||||||
|
|
||||||
|
## Open
|
||||||
|
|
||||||
|
- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
||||||
|
inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
||||||
|
is slow, and everything above is theory.
|
||||||
|
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||||
|
decides whether a person can find the one they left standing yesterday.
|
||||||
|
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||||
|
the instance must not destroy them.
|
||||||
|
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
||||||
|
before the mesh builds itself that somewhere is outside it — the open question from
|
||||||
|
[research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md).
|
||||||
@@ -12,6 +12,7 @@ document is written and this one's status becomes `implemented`.
|
|||||||
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
|
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
|
||||||
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
|
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
|
||||||
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
|
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
|
||||||
|
| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) |
|
||||||
|
|
||||||
## Not yet written
|
## Not yet written
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user