A container runtime on the same workstation sets the kernel's forwarding policy to drop, and the virtualisation daemon's own accept rules do not override it — both are consulted and a drop anywhere is the answer. The machines then get addresses and resolve names, because the daemon's resolver is on the bridge, and discard every packet to anything real. The sharpest 'available is not adequate' yet: nothing is misconfigured, nothing logs, and it presents as every image pull hanging. A workstation that runs containers is the ordinary case, so it is a prerequisite — verified by reaching something, never by reading a setting. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
132 lines
7.3 KiB
Markdown
132 lines
7.3 KiB
Markdown
---
|
|
layer: to-be
|
|
status: designed
|
|
code: [mesh-lab]
|
|
updated: 2026-09-11
|
|
decisions:
|
|
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
|
|
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
|
|
---
|
|
|
|
# Installing the lab on a clean machine
|
|
|
|
The lab has prerequisites — a virtualisation daemon, copy-on-write storage, a pool, an identity
|
|
permitted to talk to it — and it cannot get them from the mesh, because it is where the mesh is
|
|
built ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)).
|
|
|
|
So the lab needs an install path of its own. This describes it, and the shape it has to take is
|
|
determined by two failures observed while measuring
|
|
([research 010](../../01-RESEARCH/010-lab-inner-loop-cost/measurements.md)).
|
|
|
|
## The two failures that shape this
|
|
|
|
**One: installed is not available.** The virtualisation package was present and explicitly
|
|
installed. Both its units were disabled, the operator was in no group, and the client reported
|
|
the server unreachable. Nothing had failed — the declaration was satisfied exactly as written
|
|
([`04-ISSUES/007`](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md)).
|
|
|
|
**Two, and worse: available is not adequate.** With the storage tooling absent, the daemon
|
|
offered one driver, everything worked, and snapshots took **seventy-six times longer** than they
|
|
needed to. Nothing failed. Nothing warned. A lab in that state runs correctly and is simply too
|
|
slow to use — and the inner loop it exists to provide quietly does not exist.
|
|
|
|
The second is the more dangerous shape and it is a variant the mesh has not catalogued before.
|
|
Its usual failure is *reported success and did nothing*. This is **reported success and did it
|
|
seventy-six times slower**, which no error surface catches because nothing is wrong.
|
|
|
|
## What follows: the lab verifies capability, never installation
|
|
|
|
The install path may differ. **The verification does not.**
|
|
|
|
Before the lab will raise anything, it asserts the outcomes it needs — not that packages are
|
|
present, but that the machine can actually do the work:
|
|
|
|
| Assertion | Failing means |
|
|
|---|---|
|
|
| the daemon answers **as the invoking user**, not as root | a group membership that was granted but never took effect |
|
|
| a copy-on-write storage driver is offered | the userspace tooling is missing; snapshots will be full copies |
|
|
| **the pool the lab will use is on that driver** | a pool exists, and is the slow kind — the failure that has no symptom |
|
|
| hardware virtualisation is present | machines will be emulated and unusably slow |
|
|
| an image can be fetched or is cached | the first raise will fail late instead of early |
|
|
|
|
Each check states **why it matters**, in the terms of what it costs. *"The pool uses the `dir`
|
|
driver"* means nothing to someone who does not already know it means seventy-six times slower
|
|
and unbounded at worst.
|
|
|
|
**The lab refuses to run degraded.** It does not warn and continue: a warning about a slow inner
|
|
loop is read once and ignored forever, and the loop stays slow. This is
|
|
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied where the failure is
|
|
performance rather than an error.
|
|
|
|
## Two ways the prerequisites arrive
|
|
|
|
**On a machine the mesh manages** — a module declares them, and a hook turns them into
|
|
capabilities: the units enabled, the group granted, the pool created on the right driver. This
|
|
already works; it is what was done for the virtualisation daemon itself.
|
|
|
|
**On a machine the mesh does not manage** — the lab's own bootstrap does it. One command, on a
|
|
clean machine, that installs what is missing and configures it.
|
|
|
|
The second path is not a convenience. It is **required**, because the lab must work before the
|
|
mesh does, and a lab that could only be installed by a mesh would be unable to host the
|
|
development of the mesh that installs it.
|
|
|
|
Both paths end at the same verification. Whoever satisfied the prerequisites, the lab checks
|
|
them itself — because the lesson of the first failure is precisely that *something else said it
|
|
was done* is not evidence.
|
|
|
|
## The lab is the second thing installed by hand
|
|
|
|
Worth stating, because it looks like an exception and is not.
|
|
|
|
The node host is the one thing installed by hand on a machine
|
|
([research 006](../../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)): everything else
|
|
arrives through it. The lab is the same shape on a workstation — installed once, by hand, and
|
|
then everything about the mesh is developed inside it.
|
|
|
|
Two bootstraps, at two levels, for the same reason: **something has to be first, and pretending
|
|
otherwise produces a circularity that gets papered over with a script nobody exercises.**
|
|
|
|
## What a clean install actually needs
|
|
|
|
In order, on a machine with nothing:
|
|
|
|
1. **A virtualisation daemon**, running, with its socket enabled.
|
|
2. **Copy-on-write storage tooling** — the kernel side is usually already present; it is the
|
|
userspace half that is missing and that decides whether the driver is offered at all.
|
|
3. **A pool on that driver.** A loop-backed file is sufficient and needs no partitioning, which
|
|
matters: requiring a dedicated filesystem would make the lab uninstallable on a machine
|
|
already in use.
|
|
4. **Group membership** for the operator — which does **not** apply to sessions that already
|
|
existed. Observed directly: a shell whose process tree predated the grant could not reach
|
|
the daemon while a fresh lookup showed the membership present. The bootstrap has to say so,
|
|
or the first thing a person meets is a permission error that looks like a broken install.
|
|
5. **A path out to the internet for the machines that need one.** A scenario's own segments are
|
|
isolated on purpose, but a machine that fetches what it starts from is attached to an uplink
|
|
the daemon translates. That is the daemon's business and it does it — and then a **container
|
|
runtime on the same workstation sets the kernel's forwarding policy to drop**, which the
|
|
daemon's own accept rules do not override, because both are consulted and a drop anywhere is
|
|
the answer.
|
|
|
|
The result is the sharpest instance of *available is not adequate* yet: the machines get
|
|
addresses, they resolve names — the daemon's resolver is on the bridge, so that half works —
|
|
and every packet to anything real is discarded. Nothing is misconfigured, nothing logs, and
|
|
the failure presents as *every image pull hangs*. A workstation that runs containers is the
|
|
ordinary case, so this is a prerequisite rather than a quirk: forwarding must be permitted for
|
|
the lab's own bridges, and it must be **verified by reaching something**, never by reading a
|
|
setting.
|
|
|
|
6. **Verification**, as above, before anything is raised.
|
|
|
|
## Open
|
|
|
|
- **Whether the lab's bootstrap may install packages at all**, given that the mesh's rules
|
|
forbid installing by hand. The resolution is probably that the lab's bootstrap *is* the
|
|
sanctioned mechanism on an unmanaged machine, in the way the mesh's own first-node script is —
|
|
but that is an argument to record, not to assume.
|
|
- **What "adequate" means numerically.** The checks above are qualitative. A snapshot-time
|
|
threshold would catch a copy-on-write pool that is slow for some other reason, and would be a
|
|
real assertion rather than a proxy.
|
|
- **Whether the lab should own its pool** rather than using an existing one. Owning it makes the
|
|
driver guaranteed; sharing it avoids duplicating storage on a machine that already has a pool.
|