Files
hq/02-DECISIONS/0068-the-lab-takes-requests.md
jschoubben 9da45c68d6 Propose that the lab takes requests, one at a time, from its own copy
Raising a scenario occupies the machine and the person who started it, and
running in the background against a working copy is worse than waiting: a run
reads that copy as it goes, so editing while it runs yields a result describing a
state that never existed.

Proposed rather than accepted. The load-bearing part is the restriction — a
request names a bed and a commit and nothing else — because a request that could
say what to install and where would make the lab a second way of installing a
mesh, which is the arrangement that just cost a year of late-found faults.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-12 16:51:37 +02:00

113 lines
6.5 KiB
Markdown

---
topic: building it
status: proposed
date: 2026-09-12
deciders: jochen
reconstructed: false
extends: 0016-the-lab.md
---
# 68. The lab takes requests, one at a time, and runs each from its own copy
## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
addresses and names on the workstation for as long as it stands, and only one scenario can stand
at a time. So a run is not merely slow — it occupies the machine and the person who started it,
who then waits rather than works.
**Running it in the background against the working copy is worse than waiting.** The obvious fix
is to start a run and carry on editing. But a run reads the working copy as it goes: binaries are
rebuilt from it, manifests are read out of it, and the bed's own code is loaded from it. Edit
while it runs and the result describes a state that never existed — a mixture of what was there
when each file happened to be read. A green result obtained that way is not evidence, and a red
one costs a day to disbelieve.
**Nothing today records what was asked for.** A run is a command line in somebody's terminal. What
commit it exercised, what it was trying to find out, and what it answered all live in scrollback,
which is why the same question gets re-run rather than looked up.
**Most of the parts already exist.** The lab writes a receipt of its last run. The mesh already
carries messages between nodes and can notify a person. The machine already runs work on a
schedule. What is missing is the thing in the middle.
## Considered Options
**1. Leave it as it is — a person drives the lab and waits.** Rejected. It is the loop
[ADR 0010](0010-delivery.md) removed everywhere else, kept here by habit rather than by argument,
and the cost compounds: because a run is expensive to start and blocks the person, fewer are run,
so faults are found later and in larger batches.
**2. Run in the background against the working copy.** Rejected on the reasoning above. The
failure is silent, which is the kind this repository exists to refuse.
**3. Put the lab behind the ordinary build pipeline.** Rejected for now. The pipeline builds
artifacts and does not own a machine that can raise virtual machines; giving it one makes the
pipeline's slowest job the lab's, and couples every push to hardware only one machine has. This
may become right later; it is not the smallest thing that works.
**4. A queue in front of the lab, and an isolated copy behind it.** Chosen.
## Decision
**The lab accepts requests rather than commands.** A request is recorded, queued, and answered.
The person who made it is told when it is answered and does not wait.
**A request names a bed and a commit, and nothing else.** This is the load-bearing restriction. A
request may say *run this bed, at this version of these repositories*. It may not say what to
install, on which machine, or with which settings — because a request that could say those things
would be a second way of installing a mesh, and the whole reason the installer exists is that the
lab already was one ([ADR 0067](0067-genesis-is-a-pivot.md)). The bed decides what is installed;
the request only decides which bed and which version.
**Requests are released one at a time.** The hardware admits one standing scenario, so the queue
enforces what the hardware already requires, rather than leaving it to whoever remembers.
**Every run happens in a copy the lab owns.** The lab checks the requested commit out into its own
path and builds and runs from there. A working copy is never read by a run. This is what makes the
queue safe to use while work continues, and without it the rest of this record is not worth
having.
**The lab is reached through tools, not only a command line.** A command line is available only
to whoever is sitting at the machine, which is the constraint this record exists to remove. The
lab answers three questions to anything that can reach the mesh — *what is standing now*, *what is
queued or running*, and *what did this request answer* — and accepts a request and a cancellation.
An agent can therefore start a run, stop attending to it, and come back; and somebody who did not
start a run can still see it, which is the difference between a shared lab and a private one.
**The restriction holds at every door.** A tool submits a bed and a commit, exactly as a command
line does. A tool that could name a module, a node or a setting would reintroduce the second
installer through a different entrance, and the entrance is not what made it dangerous.
**Every run leaves a record that outlives the terminal**: what was asked, which commit, when it
ran, what it answered, and where its output went. A question already answered is looked up rather
than re-run.
## Consequences
Work continues while the lab runs, which is the point. A second session may edit freely, because
nothing it edits is what the lab is reading.
A request is reproducible by construction: it names a commit, so the same request can be asked
again and compared. Today two runs of "the same thing" are only as alike as the tree happened to be.
The lab gains a second copy of every repository it exercises, costing disk and needing to be kept
from drifting into a place people edit by hand.
Anything that can reach the mesh can now see what the lab is doing, including an agent working on
something else. That is the intended gain and also the obvious hazard: a thing that is easy to ask
is easy to ask too often, and the hardware still admits one scenario at a time.
The queue becomes a thing that can fail — stuck, backed up, or lost — and a queue nobody watches
is worse than no queue, because it absorbs requests silently.
## How this is checked
| Rule | Checked by |
|---|---|
| A run never reads a working copy | The runner is given a path it owns and no other; a run started while a working copy is deliberately dirtied produces a result matching the commit, not the edits. |
| One scenario stands at a time | A second request submitted while one runs is observed to wait, not to raise. |
| A request cannot say what to install | The request format admits a bed and a commit only. A request naming a module, a node or a setting is refused, and the refusal is exercised. |
| A request is answered | Every queued request reaches a terminal state with a record. A request that vanishes is a failure of the queue, not a quiet nothing. |
| The lab can be asked from elsewhere | What is standing is asked from a session that did not raise it, and the answer matches the machine. A lab that only answers its own caller has not left the terminal. |