The scenario model, the lifecycle, and what the lab actually costs #6
@@ -0,0 +1,71 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-23
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
---
|
||||
|
||||
# 32. A scenario is an isolated address space, and the lab never reaches into it over IP
|
||||
|
||||
## Context
|
||||
|
||||
A scenario declares literal addresses —
|
||||
[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has
|
||||
to be, because reproducing *published but behind NAT* means saying which address the world sees.
|
||||
|
||||
That raises a question the declaration left open: **two scenarios at once.** Several agents
|
||||
working means several scenarios, and the lab design already calls that a requirement. But two
|
||||
scenarios built from the same declaration want the same addresses, and there are only three
|
||||
documentation ranges in existence.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals.
|
||||
Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a
|
||||
specific topology no longer reproduces it; it breaks the RFC-range validation, since
|
||||
allocated addresses would have to come from somewhere real; and the numbers a person reads
|
||||
in the file stop being the numbers they will see in a capture.
|
||||
2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate
|
||||
an agent has to queue for is a gate that gets bypassed.
|
||||
3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**A scenario is a closed address space.** Every segment materialises as its own isolated link,
|
||||
belonging to one scenario instance. Two scenarios raised from the same declaration hold the same
|
||||
addresses and never meet, because nothing joins their links.
|
||||
|
||||
The declaration therefore keeps its literal addresses, and they mean exactly what they say.
|
||||
|
||||
**The consequence that constrains everything else: the lab never reaches into a scenario over
|
||||
IP.** It talks to a machine through the virtualisation layer's own channel — the same way one
|
||||
executes a command in a container without the container being routable.
|
||||
|
||||
That is not a preference. If the lab reached machines by address, the workstation running it
|
||||
would need a route into each scenario, and two scenarios carrying the same prefix would give it
|
||||
two routes to the same destination. Concurrency would be impossible, and it would fail in the
|
||||
worst available way: not with an error, but by one scenario's traffic arriving in another.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit
|
||||
beyond the machine's capacity.
|
||||
- The three documentation ranges stop being a scarce resource. Every scenario may use all of
|
||||
them, because no two scenarios share a link.
|
||||
- **The lab cannot use IP to check anything**, which is more of a constraint than it first
|
||||
appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by
|
||||
executing on a machine, rather than probed from outside. That is the honest way to ask it
|
||||
anyway — reachability from the workstation is not the question.
|
||||
- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing
|
||||
outside holds a reference into it.
|
||||
- The lab needs a scenario **instance** identity distinct from the scenario name in the
|
||||
declaration: the declaration is a kind, and several instances of one kind may exist.
|
||||
- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a
|
||||
developer wants to reach — a web interface, a database — needs an explicit, deliberate
|
||||
forward out of the scenario, which is a feature rather than a gap: nothing leaks by default.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses
|
||||
this preserves.
|
||||
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes.
|
||||
@@ -634,10 +634,6 @@ is the one real absence, and it is exactly the double-NAT case.
|
||||
- **Where `place:` gets its artifacts from.** Before the mesh is self-hosting these come from
|
||||
outside; afterwards from the mesh itself. The declaration should not have to care, which
|
||||
suggests a named source rather than a path.
|
||||
- **Multiple scenarios at once.** Each needs its own segments and addresses, and the shape above
|
||||
writes addresses absolutely. Whether a scenario carries literal addresses or a template the
|
||||
lab allocates from decides whether two can run side by side — and there are only three
|
||||
documentation ranges to go round.
|
||||
- **Nested forwarding** — `published:` names one gateway, so a machine behind two cannot be
|
||||
published through both.
|
||||
- **An address changing in place**, as a DHCP lease expiring under a machine that has not moved.
|
||||
|
||||
@@ -0,0 +1,134 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: [mesh-lab]
|
||||
updated: 2026-08-23
|
||||
decisions:
|
||||
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
||||
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
|
||||
- 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
|
||||
---
|
||||
|
||||
# Scenario lifecycle
|
||||
|
||||
The first thing the lab must do, and the only thing it must do before anything else can be
|
||||
written: **materialise a mesh, return it to a known state, and destroy it**
|
||||
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
||||
|
||||
A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens
|
||||
to one.
|
||||
|
||||
## The verbs
|
||||
|
||||
| Verb | Does |
|
||||
|---|---|
|
||||
| `raise` | materialise a declaration into a running scenario instance |
|
||||
| `snapshot` | name the current state of the whole scenario |
|
||||
| `restore` | return the whole scenario to a named state |
|
||||
| `move` | change a machine's position while the scenario runs |
|
||||
| `exec` | run something on a machine, and get its output |
|
||||
| `destroy` | tear the instance down |
|
||||
|
||||
Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition
|
||||
cheap, and cheap repetition is what turns the bootstrap path into an inner development loop
|
||||
rather than a ceremony.
|
||||
|
||||
## Raising, in order
|
||||
|
||||
The order is not arbitrary — each step needs the one before it to exist:
|
||||
|
||||
1. **Segments.** Isolated links, one per declared segment, belonging to this instance and
|
||||
joined to nothing outside it
|
||||
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)).
|
||||
2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each
|
||||
distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the
|
||||
translation, forwarding and mapping-expiry the declaration asked for.
|
||||
3. **Machines.** Each on its segments, holding its addresses.
|
||||
4. **Policy.** Rules between segments, applied on the gateways that route between them.
|
||||
5. **Placement.** Artifacts onto machines.
|
||||
6. **Snapshot**, if the declaration named one.
|
||||
|
||||
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
||||
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
||||
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
||||
habit.
|
||||
|
||||
## A failed raise leaves the wreckage
|
||||
|
||||
A step that fails stops the raise
|
||||
([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear
|
||||
down**.
|
||||
|
||||
Tearing down on failure destroys the only evidence of what went wrong, which is precisely
|
||||
backwards: a scenario that failed to raise is more interesting than one that succeeded. The
|
||||
instance stays, marked failed, with the step that failed named.
|
||||
|
||||
The caller decides what happens next, and the two callers want different things
|
||||
— the coordinator captures and destroys, a person opens a shell. That is the same
|
||||
one-runner-two-callers split the lab design already makes, applied to failure.
|
||||
|
||||
## Snapshots are whole-scenario
|
||||
|
||||
A snapshot captures **every machine and the state of the network between them**, as one thing.
|
||||
Restoring returns all of it.
|
||||
|
||||
Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes —
|
||||
what is assigned where, which grants exist, what has been delivered — so restoring one machine
|
||||
to an earlier moment while its peers move on produces a mesh that has never existed and could
|
||||
not. The faults found there would be artefacts of the lab.
|
||||
|
||||
This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh
|
||||
that has never seen a change, another of a mesh running the previous version, and a restore
|
||||
between runs.
|
||||
|
||||
## Moving a machine
|
||||
|
||||
`move` changes a machine's position while the scenario runs: to another segment, with different
|
||||
addresses, or to `detached`.
|
||||
|
||||
It is the roaming case, and it is a **lifecycle** operation rather than a declaration because
|
||||
the interesting part is the transition, not the destination. A mesh that forms correctly with a
|
||||
node at home and correctly with it away may still fail to notice it moved.
|
||||
|
||||
Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change
|
||||
made after it, and returning undoes it like any other change.
|
||||
|
||||
## Reaching in
|
||||
|
||||
Everything the lab does to a machine goes through the virtualisation layer, never over IP
|
||||
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a
|
||||
command on a machine and returns its output.
|
||||
|
||||
This has one consequence worth stating plainly: **a reachability question is asked from inside**.
|
||||
*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from
|
||||
the workstation. The workstation is not on the scenario's network and its opinion of
|
||||
reachability would be a different question with a misleadingly similar answer.
|
||||
|
||||
Anything a person wants to open in a browser needs a deliberate forward out of the instance.
|
||||
Nothing leaks by default.
|
||||
|
||||
## The two callers
|
||||
|
||||
The lab design already establishes that the runner serves the coordinator and a person, and
|
||||
that anything only one of them can do will drift. Applied here:
|
||||
|
||||
| | the coordinator | someone working on the mesh |
|
||||
|---|---|---|
|
||||
| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest |
|
||||
| on failure | capture, then destroy | leave it, open a shell |
|
||||
|
||||
Both use the same verbs. The difference is what happens after the verdict, which is a caller's
|
||||
decision rather than a second implementation.
|
||||
|
||||
## Open
|
||||
|
||||
- **What a snapshot costs.** Whole-scenario snapshots of several machines are the operation the
|
||||
inner loop repeats most, so their cost sets the loop's speed. If restoring is slow, the loop
|
||||
is slow, and everything above is theory.
|
||||
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
||||
decides whether a person can find the one they left standing yesterday.
|
||||
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
||||
the instance must not destroy them.
|
||||
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
||||
before the mesh builds itself that somewhere is outside it — the open question from
|
||||
[research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md).
|
||||
@@ -12,6 +12,7 @@ document is written and this one's status becomes `implemented`.
|
||||
| [`00-work-breakdown.md`](00-work-breakdown.md) | How the decomposition gets built, in what order, and where a human must look | [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) |
|
||||
| [`01-end-to-end-testing.md`](01-end-to-end-testing.md) | The lab: a real mesh a change can be run against before it reaches nodes | [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md), [0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) |
|
||||
| [`02-scenario-declaration.md`](02-scenario-declaration.md) | What a scenario declares — the underlay, and what to place on it | [ADR 0031](../../02-DECISIONS/0031-the-lab-provides-the-underlay.md) |
|
||||
| [`03-scenario-lifecycle.md`](03-scenario-lifecycle.md) | What happens to a scenario — raise, snapshot, restore, move, destroy | [ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md) |
|
||||
|
||||
## Not yet written
|
||||
|
||||
|
||||
Reference in New Issue
Block a user