Merge main: 0115 was taken by another record

This commit is contained in:
2026-09-26 18:54:59 +02:00
4 changed files with 181 additions and 0 deletions
@@ -0,0 +1,49 @@
---
topic: what runs on it
status: accepted
date: 2026-09-26
deciders: jochen
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 115. One assignment of a module per node: the module's name is the assignment's identity
## Context
Everything an assignment owns is named after its module: the database user
(`mesh_<node>_<module>`), the broker account (`<node>-<module>`), the containers, the sealed
secrets, and — since ADR 0112 — the placed directory (`<root>/<module>`). A second assignment
of the same module on the same node would collide on every one of those names at once, which
is why the mesh has never allowed it.
[Issue 119](../04-ISSUES/119-a-module-definition-decides-where-its-files-live/00-report.md)
recorded this as a kept limitation, and 0112 deliberately did not fix it — giving assignments
identities of their own would have touched every naming recipe in one already-large change.
The question stayed open: is multi-assignment a requirement deferred, or a requirement at all?
The original motivation was real — one module serving two tenants on one machine, a second
photo site, a second mail domain. The operator held that requirement once and has now weighed
it against what it costs.
## Decision
**Dropped.** One assignment of a module per node is the rule, not a limitation. The module's
name IS the assignment's identity on a node, permanently, and every naming recipe may rely on
it.
Wanting the same software twice on one node has a spelling the mesh already supports: **two
modules.** A module definition is cheap — two photo sites are two modules sharing artifacts
(the build's images are content-addressed; nothing is built twice), each with its own name,
its own directory, its own grants and its own routes. The tenant boundary lands where every
other boundary already is: the module name.
## Consequences
- The naming recipes stay as simple as they are. No instance suffixes, no assignment ids
threaded through six systems, no migration of every existing name.
- `<root>/<module>` is the assignment's directory with nothing left open (0112's placement
language stands unchanged).
- The controller may refuse a second assignment *plainly* — "novox already runs mailu, and one
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the
control plane.
@@ -365,3 +365,29 @@ Each phase ends at a check that holds, so none of them leaves a mechanism half-r
- The layout a node's default root uses beneath it, beyond one directory per assignment.
- Whether a module provider's answer can change without the provider being asked, for example a
provider moving. The rule so far is that it cannot, and moving is re-resolving.
## The three gaps, answered (2026-09-26, operator)
**Where the root comes from.** A node setting, fixed at installation; a node that states none
gets the default under `/var/lib`. (The adopted-machine placement stays as proposed: a
per-assignment placement for data that must sit where it already is — mssql is novox's live
case.)
**What sits beneath it — dissolved, not decided.** The question assumed the mesh's own
writes (`/var/lib/mesh/<module>`: sealed credentials, composed bindings) need a
module-visible reservation. They do not: a module *requires* a `host-path` and receives a
location; what the mesh writes for the module is the mesh's plumbing, placed where the mesh
chooses and mounted in — never part of the module's contract. One reservation per
requirement, `<root>/<module>/<name>`.
**Resolution happens in the controller, at declaration composition.** The node receives
concrete paths exactly as today — the wire format and the host's apply do not change for
this. What changes is that no *manifest* carries a path; the controller resolves
requirement → location against the node's root setting. (The host still needs issue 126's
fix — volumes in the spec comparison — or a resolved path change cannot reach a running
container.)
**Retirement order** stays per-module data migration with verification, smallest and
empties first, the mail spool and the store last — the novox session's six-module window
(mesh-catalog #97, hq 126) is the worked example, hazard included.
@@ -0,0 +1,59 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-host internal/apply, mesh-controller]
---
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
## What was observed
During the route-proxy edge cutover on novox (2026-09-26): the module was assigned, the
push reported success, `status` said the node was doing everything it was told — and the
module's three containers did not exist. The operator stopped the predecessor's proxy on
the strength of those reports, and every public name on the node went dark until rollback.
The cause was correct behaviour, invisibly reported. The first (rolled-back) route-proxy
attempt had left `/var/lib/route-proxy/*` on disk; on re-assign, the adopted node *found*
those directories, held them (ADR 0100, exactly as designed), and held every container
that mounts them — `"would mount /var/lib/route-proxy/ca, found on this adopted node;
not run until route-proxy is taken"`. All of that lived only in `state.json`. What the
operator saw:
- the push: `sent novox 346 resource(s)` — the controller's count of what it sent;
- the node's journal: `applied 330 resource(s)` — sixteen fewer, with no line saying
which sixteen or why;
- `status`: green — a held resource is not "wrong", so nothing was flagged;
- `node show novox`: the holds list did NOT include route-proxy's (it showed only holds
the *controller* knew about from take-time listings, not what the node decided at
apply-time).
Four surfaces, none carrying the one sentence that mattered: *route-proxy is assigned
but not taken, and its containers will not run until it is.*
## Why this is a real fault and not operator error alone
The operator error (an edge-flip runbook that omitted `take`) was only possible because
every surface reported success. A system whose correct refusals are indistinguishable
from completed work will keep converting small procedural gaps into outages. The
`sent 346 / applied 330` discrepancy was the single visible symptom, and interpreting it
required reading `state.json` by hand.
## What would have prevented it
Any one of:
1. **The apply report says what it held.** `applied 330 resource(s), 16 held for
untaken modules (route-proxy: 13, …)` — one line in the journal.
2. **`status` counts holds against untaken-but-assigned modules.** A module assigned,
pushed, and running zero of its containers is at minimum worth a "waiting on take"
line — it is never converged in any useful sense.
3. **`node show <node>` shows the node's own held list**, not only what take-time
computed — the node already records it in `state.json` with reasons.
## Precedent
The photos cutover hit the same semantics benignly the same week (assign → held
containers in `Created` state → take), and the mailu cutover documented "take is the
verb, and ADR 0100 meant it". The semantics are consistent and right; the reporting is
what let them be forgotten at the worst moment.
@@ -0,0 +1,47 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-host internal/apply]
---
# 126 — a volume path is not in the spec comparison, and a roll-out raced a data move
## What was observed
Landing the "module data lives in /var/lib" change on novox (mesh-catalog #97), two
distinct faults surfaced in one hour:
1. **Building a module with a roll-out upgrade policy IS deploying it.** gitea's policy
was roll-out; the `build` that registered its repathed manifest sent it to the node
immediately, which recreated the container mounting the *not-yet-renamed* (empty)
`/var/lib/gitea/data`. The forge came back as its own install page, fresh host keys
and all, and every subsequent pipeline build died on `repository not found` — which
also blocked the fix, since re-registering the other modules needed the forge. The
operator narrative "build, then move data, then push" is only safe under the record
policy; nothing warned that one module in the batch would skip the pause.
2. **Changing a container's volume paths does not recreate the container.** After the
final push, five of the six repathed modules kept their old containers running
("Up 13–26 hours") — the new declaration's volume paths differ from the running
containers' mounts, and the apply judged them current. Same class as mesh-host #27
(`dns`/`ip` absent from the comparison): a field the comparison does not read is a
field that can never change a running container. Benign here only because a rename
on one filesystem preserves the mounted inode — the running containers keep serving
the same bytes the new path names, and the next natural recreation converges. A
cross-filesystem move, or a path change to *different* data, would have silently
split the module between two worlds.
## What would have prevented it
- `build` printing the module's upgrade policy when that policy will act on the result
("gitea rolls out on build — the node will receive this immediately"), or a
`--register-only` flag for exactly this choreography.
- Volumes (and every other container field) in the spec comparison, or the honest
refusal: "this field changed and I cannot apply it without recreation."
## Recovery that worked
Instant renames both ways broke the circular dependency (forge needed for builds,
builds needed for the push, push needed for the forge): data back to the old path,
old-spec forge started, artifacts rebuilt, data renamed forward, push. Nothing lost;
the install-page junk was discarded twice.