The connectivity design still said a hub cannot be filtered — a gap recorded in the morning and closed in the afternoon, left standing as though it were current. Worse than a stale date: it would send somebody away from something that works. `restart-on` was described nowhere, including the part added today that lets a service reflect a file another module put on the machine. A rule the host enforces and no document mentions is a rule nobody can rely on. And nine of fifteen design documents claimed an `updated:` older than their last change, some by a week. That field is what cross-cutting views are generated from, so it is not decoration.
147 lines
7.0 KiB
Markdown
147 lines
7.0 KiB
Markdown
---
|
|
layer: to-be
|
|
status: in-progress
|
|
code: [mesh-lab]
|
|
updated: 2026-08-28
|
|
decisions:
|
|
- 02-DECISIONS/0016-the-lab.md
|
|
- 02-DECISIONS/0016-the-lab.md
|
|
- 02-DECISIONS/0016-the-lab.md
|
|
---
|
|
|
|
# Scenario lifecycle
|
|
|
|
The first thing the lab must do, and the only thing it must do before anything else can be
|
|
written: **materialise a mesh, return it to a known state, and destroy it**
|
|
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
|
|
|
|
A [declaration](02-scenario-declaration.md) describes a scenario. This describes what happens
|
|
to one.
|
|
|
|
## The verbs
|
|
|
|
| Verb | Does |
|
|
|---|---|
|
|
| `raise` | materialise a declaration into a running scenario instance |
|
|
| `snapshot` | name the current state of the whole scenario |
|
|
| `restore` | return the whole scenario to a named state |
|
|
| `move` | change a machine's position while the scenario runs |
|
|
| `exec` | run something on a machine, and get its output |
|
|
| `destroy` | tear the instance down |
|
|
|
|
Six verbs, and `raise` plus `destroy` are enough to be useful. The rest are what make repetition
|
|
cheap, and cheap repetition is what turns the bootstrap path into an inner development loop
|
|
rather than a ceremony.
|
|
|
|
## Raising, in order
|
|
|
|
The order is not arbitrary — each step needs the one before it to exist:
|
|
|
|
1. **Segments.** Isolated links, one per declared segment, belonging to this instance and
|
|
joined to nothing outside it
|
|
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
|
|
2. **Gateways.** Derived, never declared as machines: a gateway is materialised for each
|
|
distinct `gateway:` declaration, sitting on both its segment and its parent, carrying the
|
|
translation, forwarding and mapping-expiry the declaration asked for.
|
|
3. **Machines.** Each on its segments, holding its addresses.
|
|
4. **Policy.** Rules between segments, applied on the gateways that route between them.
|
|
5. **Placement.** Artifacts onto machines.
|
|
6. **Snapshot**, if the declaration named one.
|
|
|
|
`raise` must wait for the machines to be **usable**, not for the call to return. Measured, those
|
|
are 14 seconds apart for a single machine — and reporting the earlier number would be the mesh's
|
|
own recurring failure, transport reported as effect.
|
|
|
|
Raising is **convergent, not incremental**: raising an instance that already exists brings it to
|
|
the declared state rather than failing or duplicating. That is the same model the mesh itself
|
|
uses, and a lab that behaved differently from the thing it tests would be teaching the wrong
|
|
habit.
|
|
|
|
## A failed raise leaves the wreckage
|
|
|
|
A step that fails stops the raise
|
|
([ADR 0010](../../02-DECISIONS/0010-delivery.md)) — and **does not tear
|
|
down**.
|
|
|
|
Tearing down on failure destroys the only evidence of what went wrong, which is precisely
|
|
backwards: a scenario that failed to raise is more interesting than one that succeeded. The
|
|
instance stays, marked failed, with the step that failed named.
|
|
|
|
The caller decides what happens next, and the two callers want different things
|
|
— the coordinator captures and destroys, a person opens a shell. That is the same
|
|
one-runner-two-callers split the lab design already makes, applied to failure.
|
|
|
|
## Snapshots are whole-scenario
|
|
|
|
A snapshot captures **every machine and the state of the network between them**, as one thing.
|
|
Restoring returns all of it.
|
|
|
|
Per-machine snapshots would be cheaper and are wrong. The mesh keeps state that spans nodes —
|
|
what is assigned where, which grants exist, what has been delivered — so restoring one machine
|
|
to an earlier moment while its peers move on produces a mesh that has never existed and could
|
|
not. The faults found there would be artefacts of the lab.
|
|
|
|
This is what makes *fresh* and *upgrade* both cheap and both default: one snapshot of a mesh
|
|
that has never seen a change, another of a mesh running the previous version, and a restore
|
|
between runs.
|
|
|
|
## Moving a machine
|
|
|
|
`move` changes a machine's position while the scenario runs: to another segment, with different
|
|
addresses, or to `detached`.
|
|
|
|
It is the roaming case, and it is a **lifecycle** operation rather than a declaration because
|
|
the interesting part is the transition, not the destination. A mesh that forms correctly with a
|
|
node at home and correctly with it away may still fail to notice it moved.
|
|
|
|
Moving does not invalidate a snapshot. A snapshot is a state to return to; a move is a change
|
|
made after it, and returning undoes it like any other change.
|
|
|
|
## Reaching in
|
|
|
|
Everything the lab does to a machine goes through the virtualisation layer, never over IP
|
|
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)). `exec` runs a
|
|
command on a machine and returns its output.
|
|
|
|
This has one consequence worth stating plainly: **a reachability question is asked from inside**.
|
|
*Can this machine reach that one* is `exec` on the first, testing the second — not a probe from
|
|
the workstation. The workstation is not on the scenario's network and its opinion of
|
|
reachability would be a different question with a misleadingly similar answer.
|
|
|
|
Anything a person wants to open in a browser needs a deliberate forward out of the instance.
|
|
Nothing leaks by default.
|
|
|
|
## The two callers
|
|
|
|
The lab design already establishes that the runner serves the coordinator and a person, and
|
|
that anything only one of them can do will drift. Applied here:
|
|
|
|
| | the coordinator | someone working on the mesh |
|
|
|---|---|---|
|
|
| wants | structured results, clean teardown, no prompts | readable output, the scenario **left standing**, one thing re-run without repeating the rest |
|
|
| on failure | capture, then destroy | leave it, open a shell |
|
|
|
|
Both use the same verbs. The difference is what happens after the verdict, which is a caller's
|
|
decision rather than a second implementation.
|
|
|
|
## Open
|
|
|
|
- **What a snapshot costs — measured, and the answer is a blocker.** On the current host a
|
|
snapshot is a full copy: 1.6 GB and ten seconds for one small machine, over two minutes when
|
|
observed again, projecting to roughly a minute and a half for a four-machine reset-and-rerun
|
|
cycle. The cause is the storage driver, not virtual machines. See
|
|
[research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md).
|
|
- **Instance naming.** A declaration is a kind and instances are many; how they are named
|
|
decides whether a person can find the one they left standing yesterday.
|
|
- ~~**Does a scenario snapshot need the machines stopped?**~~ **Answered by the integration
|
|
test on its first run: no, but they must be flushed.** A snapshot captures disk and not
|
|
memory, so a write still in the guest's page cache is absent from it — not stale, absent. A
|
|
file written seconds before a snapshot did not survive the restore. Flushing first buys
|
|
write-durability; it does not buy application-consistency, and anything mid-transaction is
|
|
still captured mid-transaction.
|
|
- **What survives `destroy`.** Logs and captures are the output of a failed run, so destroying
|
|
the instance must not destroy them.
|
|
- **Placement before the mesh is self-hosting.** `place:` needs artifacts from somewhere, and
|
|
before the mesh builds itself that somewhere is outside it — the open question from
|
|
[research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md).
|