133 lines
7.4 KiB
Markdown
133 lines
7.4 KiB
Markdown
---
|
|
layer: to-be
|
|
status: in-progress
|
|
code: [mesh-lab]
|
|
updated: 2026-09-11
|
|
decisions:
|
|
- 02-DECISIONS/0016-the-lab.md
|
|
- 02-DECISIONS/0010-delivery.md
|
|
---
|
|
|
|
# Installing the lab on a clean machine
|
|
|
|
The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity
|
|
permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is
|
|
built ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
|
|
|
|
So the lab needs an install path of its own. This describes it, and the shape it has to take is
|
|
determined by two failures observed while measuring
|
|
([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)).
|
|
|
|
## The two failures that shape this
|
|
|
|
**One: installed is not available.** The virtualisation package was present and explicitly
|
|
installed. Both its units were disabled, the operator was in no group, and the client reported
|
|
the server unreachable. Nothing had failed — the declaration was satisfied exactly as written
|
|
([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)).
|
|
|
|
**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon
|
|
offered one driver, everything worked, and snapshots took **seventy-six times longer** than they
|
|
needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too
|
|
slow to use — and the inner loop it exists to provide quietly does not exist.
|
|
|
|
The second is the more dangerous shape and it is a variant the mesh has not catalogued before.
|
|
Its usual failure is *reported success and did nothing*. This is **reported success and did it
|
|
seventy-six times slower**, which no error surface catches because nothing is wrong.
|
|
|
|
## What follows: the lab verifies capability, never installation
|
|
|
|
The install path may differ. **The verification does not.**
|
|
|
|
Before the lab will raise anything, it asserts the outcomes it needs — not that packages are
|
|
present, but that the machine can actually do the work:
|
|
|
|
| Assertion | Failing means |
|
|
|---|---|
|
|
| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect |
|
|
| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies |
|
|
| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom |
|
|
| hardware virtualisation is present | machines will be emulated and unusably slow |
|
|
| an image can be fetched or is cached | the first raise will fail late instead of early |
|
|
| **a machine on an uplink reaches something real**, by fetching it — not by reading a route or a policy | forwarding is being dropped by something else on the workstation; names still resolve, and every pull hangs |
|
|
|
|
Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir`
|
|
driver"* means nothing to someone who does not already know it means seventy-six times slower
|
|
and unbounded at worst.
|
|
|
|
**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner
|
|
loop is read once and ignored forever, and the loop stays slow. This is
|
|
[ADR 0010](../../02-DECISIONS/0010-delivery.md) applied where the failure is
|
|
performance rather than an error.
|
|
|
|
## Two ways the prerequisites arrive
|
|
|
|
**On a machine the mesh manages** — a module declares them, and a hook turns them into
|
|
capabilities: the units enabled, the group granted, the pool created on the right driver. This
|
|
already works; it is what was done for the virtualisation daemon itself.
|
|
|
|
**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a
|
|
clean machine, that installs what is missing and configures it.
|
|
|
|
The second path is not a convenience. It is **required**, because the lab must work before the
|
|
mesh does, and a lab that could only be installed by a mesh would be unable to host the
|
|
development of the mesh that installs it.
|
|
|
|
Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks
|
|
them itself — because the lesson of the first failure is precisely that *something else said it
|
|
was done* is not evidence.
|
|
|
|
## The lab is the second thing installed by hand
|
|
|
|
Worth stating, because it looks like an exception and is not.
|
|
|
|
The node host is the one thing installed by hand on a machine
|
|
([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else
|
|
arrives through it. The lab is the same shape on a workstation — installed once, by hand, and
|
|
then everything about the mesh is developed inside it.
|
|
|
|
Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending
|
|
otherwise produces a circularity that gets papered over with a script nobody exercises.**
|
|
|
|
## What a clean install actually needs
|
|
|
|
In order, on a machine with nothing:
|
|
|
|
1. **A virtualisation daemon**, running, with its socket enabled.
|
|
2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the
|
|
userspace half that is missing and that decides whether the driver is offered at all.
|
|
3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which
|
|
matters: requiring a dedicated filesystem would make the lab uninstallable on a machine
|
|
already in use.
|
|
4. **Group membership** for the operator — which does **not** apply to sessions that already
|
|
existed. Observed directly: a shell whose process tree predated the grant could not reach
|
|
the daemon while a fresh lookup showed the membership present. The bootstrap has to say so,
|
|
or the first thing a person meets is a permission error that looks like a broken install.
|
|
5. **A path out to the internet for the machines that need one.** A scenario's own segments are
|
|
isolated on purpose, but a machine that fetches what it starts from is attached to an uplink
|
|
the daemon translates. That is the daemon's business and it does it — and then a **container
|
|
runtime on the same workstation sets the kernel's forwarding policy to drop**, which the
|
|
daemon's own accept rules do not override, because both are consulted and a drop anywhere is
|
|
the answer.
|
|
|
|
The result is the sharpest instance of *available is not adequate* yet: the machines get
|
|
addresses, they resolve names — the daemon's resolver is on the bridge, so that half works —
|
|
and every packet to anything real is discarded. Nothing is misconfigured, nothing logs, and
|
|
the failure presents as *every image pull hangs*. A workstation that runs containers is the
|
|
ordinary case, so this is a prerequisite rather than a quirk: forwarding must be permitted for
|
|
the lab's own bridges, and it must be **verified by reaching something**, never by reading a
|
|
setting.
|
|
|
|
6. **Verification**, as above, before anything is raised.
|
|
|
|
## Open
|
|
|
|
- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules
|
|
forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the
|
|
sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is —
|
|
but that is an argument to record, not to assume.
|
|
- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time
|
|
threshold would catch a copy-on-write pool that is slow for some other reason, and would be a
|
|
real assertion rather than a proxy.
|
|
- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the
|
|
driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool.
|