Playbook 02 and 04 were followed for the substance — decisions before design, design before build — and skipped for the bookkeeping. This closes that. 004 graduates. Its one open item was "not yet stood up"; the lab is stood up, and the substitution the effort turned on is now enforced by the validator before anything is raised rather than left as a thing to remember. Its certificate conclusion has a home in 01-end-to-end-testing and is designed but not built — implementation is a third axis, and an effort graduates on its conclusions. One item leaves 004 without a home and is recorded rather than lost: the reverse proxy does not set caServer, so it defaults to the production endpoint. The two lab designs read `designed` while running in production of a sort, so they become `in-progress`. And the lab gets an as-is document, which it did not have. It records what runs including the parts nobody would choose again: that `place:` is refused and the lab therefore raises EMPTY MACHINES, that the drawing shipped with no design document behind it, that a router is tagged as a machine for a reason found by a bug, and that the integration suite raises two of five scenarios while both faults found so far lived in the three it does not. 006 stays active, deliberately. Two of its open questions ARE the tier 0 design — whether absorbing six concerns makes the host too large, and whether an unprivileged node earns a place in the inventory. Playbook 04 is explicit that an open question is a reason to research, not to build around.
147 lines
7.2 KiB
Markdown
147 lines
7.2 KiB
Markdown
---
|
|
layer: to-be
|
|
status: in-progress
|
|
code: [mesh-lab]
|
|
updated: 2026-08-25
|
|
decisions:
|
|
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
|
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
|
|
- 02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md
|
|
---
|
|
|
|
# Scenario lifecycle
|
|
|
|
The first thing the lab must do, and the only thing it must do before anything else can be
|
|
written: **materialise a mesh, return it to a known state, and destroy it**
|
|
([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
|
|
|
A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens
|
|
to one.
|
|
|
|
## The verbs
|
|
|
|
| Verb | Does |
|
|
|---|---|
|
|
| `raise` | materialise a declaration into a running scenario instance |
|
|
| `snapshot` | name the current state of the whole scenario |
|
|
| `restore` | return the whole scenario to a named state |
|
|
| `move` | change a machine's position while the scenario runs |
|
|
| `exec` | run something on a machine, and get its output |
|
|
| `destroy` | tear the instance down |
|
|
|
|
Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition
|
|
cheap, and cheap repetition is what turns the bootstrap path into an inner development loop
|
|
rather than a ceremony.
|
|
|
|
## Raising, in order
|
|
|
|
The order is not arbitrary — each step needs the one before it to exist:
|
|
|
|
1. **Segments.** Isolated links, one per declared segment, belonging to this instance and
|
|
joined to nothing outside it
|
|
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)).
|
|
2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each
|
|
distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the
|
|
translation, forwarding and mapping-expiry the declaration asked for.
|
|
3. **Machines.** Each on its segments, holding its addresses.
|
|
4. **Policy.** Rules between segments, applied on the gateways that route between them.
|
|
5. **Placement.** Artifacts onto machines.
|
|
6. **Snapshot**, if the declaration named one.
|
|
|
|
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
|
|
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
|
|
own recurring failure, transport reported as effect.
|
|
|
|
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
|
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
|
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
|
habit.
|
|
|
|
## A failed raise leaves the wreckage
|
|
|
|
A step that fails stops the raise
|
|
([ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md)) — and **does not tear
|
|
down**.
|
|
|
|
Tearing down on failure destroys the only evidence of what went wrong, which is precisely
|
|
backwards: a scenario that failed to raise is more interesting than one that succeeded. The
|
|
instance stays, marked failed, with the step that failed named.
|
|
|
|
The caller decides what happens next, and the two callers want different things
|
|
— the coordinator captures and destroys, a person opens a shell. That is the same
|
|
one-runner-two-callers split the lab design already makes, applied to failure.
|
|
|
|
## Snapshots are whole-scenario
|
|
|
|
A snapshot captures **every machine and the state of the network between them**, as one thing.
|
|
Restoring returns all of it.
|
|
|
|
Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes —
|
|
what is assigned where, which grants exist, what has been delivered — so restoring one machine
|
|
to an earlier moment while its peers move on produces a mesh that has never existed and could
|
|
not. The faults found there would be artefacts of the lab.
|
|
|
|
This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh
|
|
that has never seen a change, another of a mesh running the previous version, and a restore
|
|
between runs.
|
|
|
|
## Moving a machine
|
|
|
|
`move` changes a machine's position while the scenario runs: to another segment, with different
|
|
addresses, or to `detached`.
|
|
|
|
It is the roaming case, and it is a **lifecycle** operation rather than a declaration because
|
|
the interesting part is the transition, not the destination. A mesh that forms correctly with a
|
|
node at home and correctly with it away may still fail to notice it moved.
|
|
|
|
Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change
|
|
made after it, and returning undoes it like any other change.
|
|
|
|
## Reaching in
|
|
|
|
Everything the lab does to a machine goes through the virtualisation layer, never over IP
|
|
([ADR 0032](../../02-DECISIONS/0032-a-scenario-is-an-isolated-address-space.md)). `exec` runs a
|
|
command on a machine and returns its output.
|
|
|
|
This has one consequence worth stating plainly: **a reachability question is asked from inside**.
|
|
*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from
|
|
the workstation. The workstation is not on the scenario's network and its opinion of
|
|
reachability would be a different question with a misleadingly similar answer.
|
|
|
|
Anything a person wants to open in a browser needs a deliberate forward out of the instance.
|
|
Nothing leaks by default.
|
|
|
|
## The two callers
|
|
|
|
The lab design already establishes that the runner serves the coordinator and a person, and
|
|
that anything only one of them can do will drift. Applied here:
|
|
|
|
| | the coordinator | someone working on the mesh |
|
|
|---|---|---|
|
|
| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest |
|
|
| on failure | capture, then destroy | leave it, open a shell |
|
|
|
|
Both use the same verbs. The difference is what happens after the verdict, which is a caller's
|
|
decision rather than a second implementation.
|
|
|
|
## Open
|
|
|
|
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
|
|
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
|
|
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
|
|
cycle. The cause is the storage driver, not virtual machines. See
|
|
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
|
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
|
decides whether a person can find the one they left standing yesterday.
|
|
- ~~**Does a scenario snapshot need the machines stopped?**~~ **Answered by the integration
|
|
test on its first run: no, but they must be flushed.** A snapshot captures disk and not
|
|
memory, so a write still in the guest's page cache is absent from it — not stale, absent. A
|
|
file written seconds before a snapshot did not survive the restore. Flushing first buys
|
|
write-durability; it does not buy application-consistency, and anything mid-transaction is
|
|
still captured mid-transaction.
|
|
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
|
the instance must not destroy them.
|
|
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
|
before the mesh builds itself that somewhere is outside it — the open question from
|
|
[research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md).
|