Compare commits

..
Author SHA1 Message Date
jschoubben 743051efe7 Issue 163 (was 161): another record took 161 on main first 2026-09-30 13:28:22 +02:00
jschoubben e8470057aa Merge remote-tracking branch 'origin/main' into issue/161-an-assignment-does-not-record-its-provider 2026-09-30 13:28:22 +02:00
jschoubben 39340fcd76 Issue 161: an assignment does not record which provider answers it
ADR 0110 decided each assignment records where its requirements are answered
from; the control plane keeps only a per-machine pin (none recorded) and
resolves every requirement implicitly. Harmless with one provider; a second
one silently moves consumers' data.
2026-09-30 11:54:31 +02:00
3 changed files with 63 additions and 184 deletions
@@ -1,56 +0,0 @@
---
status: resolved
opened: 2026-09-30
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
amended-design:
---
# 163 — A delivered host stood aside on every push, and reported nothing
## What was observed
*2026-09-30, rolling the mesh-built host onto the last two machines.*
Every push to a machine running a delivered host produced, in order:
```
host 093231796eb0 is delivered; standing aside so the launcher runs it
applied 333 resource(s)
applied, and could not tell the mesh: reporting: context canceled
nox-mesh-host-launch: the host exited cleanly; starting it again
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
```
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
never received a single report from it: `node show` kept the version from before the crossover, and
the operator's push waited its full three minutes for an answer that was never coming.
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
## Why
After an apply the host asks whether a newer host has been delivered than the one running, and the
question was asked with the **link-time version stamp**. Since
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
host's version comes from where it sits and its stamp is `development build` — so the comparison never
matched the newest delivered version, and "a newer host is waiting" was always true.
Standing aside cancels the context the report is published with, so the report was lost on every one
of those applies. Two faults from one wrong argument.
The change that moved the version to the path was applied to the report and to the known-good record,
and not here. Half a change, and the half left behind was the one that decides whether to exit.
## Why the three-minute wait made it invisible
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
answered.
## How it is checked
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
version; it stands aside once, and the next push it does not.
@@ -0,0 +1,63 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller cmd/mesh-controller/modules.go (assign takes no provider; pin is a separate, per-machine command)
- mesh-controller internal/inventory (provision_pin keyed by (node, name))
fixed-by:
amended-design:
---
# 163 — An assignment does not record which provider answers it
## What was observed
Planning ace's modules that need a database (baserow, letta, n8n, and the apps using ace's
predecessor postgres). The operator's model — and ADR 0110's — is that **an assignment states where
each of its requirements is answered from**: gitea's assignment on novox says its `postgres-database`
comes from novox; an app assigned to ace says whether its database comes from ace or from novox.
The mesh holds no such statement for any assignment. Read on novox (2026-09-30):
```
select … from provision_pin; -- 0 rows
```
Every requirement in the mesh resolves implicitly, each time, by ADR 0084's order (a pin, then the
provider on the consumer's own node, then the only provider).
## What was decided, and what exists
[ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md):
> Where several remain and none is local, **a person chooses when the module is assigned**.
> Assignment lists the candidates, with the holder of a seat that delivers the provision suggested
> first, and records the answer on the assignment as its pin. Without an answer the module is not
> assigned.
What the control plane implements:
| decided | implemented |
|---|---|
| the answer is recorded **on the assignment** | `provision_pin` is keyed `(node, name)` — one answer per machine per provision, shared by every module on it |
| chosen **at assignment** | `assign <node> <module>` takes no provider; `pin <node> <provision> <from-node>` is a separate command |
| an assignment may be answered from its own machine (gitea ← novox) | `pin` refuses a machine pinning to itself ("does not need saying") |
| every assignment has an answer | none recorded; resolution guesses the same answer every time |
## Consequence
Nothing is wrong *today* — with one postgres provider, every guess is the intended answer. But the
answer is not a fact anyone stated, so:
- **it changes silently** the day a second provider appears (e.g. a postgres assigned on ace): every
unpinned consumer re-resolves — a consumer on ace moves from novox's database to an empty one on ace
at the next push, which is data a module stops seeing without anything saying so;
- two modules on one machine cannot take one provision from different providers;
- a person reading an assignment cannot see where its data lives.
## What would be right
ADR 0110 as written: `assign` records, per requirement, the node that answers it (its own node
included), offering the candidates and refusing an assignment without an answer where several exist;
the per-machine `provision_pin` becomes a per-assignment record, with existing assignments backfilled
from what they resolve to now so nothing moves.
@@ -1,128 +0,0 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-catalog (no module shares a path over the network)
- hq 02-DECISIONS (a file-share seat, per ADR 0126, is a module's own to define)
fixed-by:
amended-design:
---
# 169 — A machine shares its files, and the mesh does not know
## What was observed
ace serves the operator's media library to the home network with two host services no module
declares and HAL never managed either:
```
/etc/exports: /storage/media 192.168.1.0/24(rw,sync,root_squash,…) nfs-server active, :2049
/etc/samba/smb.conf: [media] path = /storage/media/ valid users = media smb active, :139/:445
```
Two LAN clients were connected at survey (2026-09-30). The library itself is operator data
(ADR 0051: ~40 TB on ZFS, the mesh owns nothing about it — [issue 153](../153-an-adopted-machines-data-cannot-be-placed-where-it-is/00-report.md)
is about modules reaching it in place).
Under the mesh as it stands, this arrangement has no expression and one failure mode:
- **Nothing declares the listens.** At `converge ace` the filter is the sum of what modules listen
on (ADR 0045); 2049 and 445 are nobody's, so the shares close — silently, for the two clients
that mount them.
- **Nothing owns the configuration.** `/etc/exports` and `smb.conf` are hand-written files on one
machine; a second machine sharing a directory would be written by hand again.
- **Nothing can consume it.** A module on another node that wanted the library (a player, an
indexer, a backup) has no `requires` to state and no binding to read; it would mount by a
hand-typed host and path.
- The clients are LAN devices, so this also meets [issue 154](../154-a-machines-own-network-is-not-a-reach/00-report.md)
(no reach for the machine's own network).
## The proposal (the operator's, 2026-09-30, settled after two rounds)
**Two module-defined seats, one per protocol, because NFS and SMB share an intent and not a
contract.** A seat in the mesh's sense is a contract — what it accepts, emits and serves, and the
tools its holder must answer (ADR 0126, 0132) — and lined up, the two share almost none of it:
| | `nfs-share` | `smb-share` |
|---|---|---|
| serves | export path(s); the client ranges allowed (`sec=sys` authorises by address) | share name(s), path |
| pair credential | none | a user and password per consumer |
| consumer's mount | `at:/path` | `//at/share` with credentials |
| holder's tools | export / unexport a path for a range | add / remove a share, create a user |
One `file-share` seat would be the union with every field optional — a consumer could bind it and
still not know how to mount what it got (the emptiness ADR 0129 warns against). "Export a path to
the network" is a category, and the mesh needs no seat category: a consumer requires the one it
can mount. If "give me the library, however" is ever needed, it is a provision an umbrella module
serves, not a seat.
Both are node-scoped, one holder per node (ADR 0110), so ace holds both. `nfs` and `samba` are the
first implementations; a second (Ganesha for `nfs-share`, ksmbd for `smb-share`) is what proves
0126's promise that "replacing the implementation changes nothing for any caller".
The holder module:
- declares the exported paths as `accesses` (ADR 0051: it owns nothing about them — never creates,
chowns or removes), and *which* paths as the assignment's settings (ADR 0046/0112);
- writes the share configuration (`/etc/exports`, `smb.conf`) as mesh-managed files and drives the
units, like `dnsmasq`/`sshd` do for theirs;
- declares its endpoints (`nfs` 2049/tcp; `smb` 445/tcp, …) so the reach — internal, or the LAN
once 154 has an answer — is the assignment's, and converge keeps them open;
- **provides** the seat's provision, so a consumer on another node `requires nfs-share` (or
`smb-share`) and reads `${bound:nfs-share:at}` and the path from its binding instead of a
hand-typed mount.
## The design gap this exposes
**A seat definition has no home outside the module that first declared it.** Today a seat is
declared inside a manifest (`showcase` declares `the-showcase`, `ca-trust` its own). If `nfs`
declared `nfs-share`, Ganesha could hold it only by depending on nfs's manifest — the coupling
0126 removed for callers, reintroduced for implementations. The protocol needs a neutral place in
the catalogue beside the modules (a seat definition registered like a manifest), with a module
saying which seats it implements. This is the first role with an obvious second implementation,
which is what makes it the exemplar for that mechanism.
## The consumer's half: a module mounts it (2026-09-30, third and fourth round)
A binding tells a consumer *where* the share is; it does not put the files on its machine. Mounting
is something done on a machine, and something done on a machine is a module's work — not the host's
(the vocabulary stays closed; no `mount` resource kind).
**A consumer-side module, `network-share` — the module responsible for setting up the network
shares a node uses** (the operator's framing). A node role, like `node-uplink` or
`node-dns-resolver`: each machine has it at most once, which is a reason for it to hold a
node-scoped seat, so two modules can never both be writing mount units on one machine. Assigned on
the node that wants the files:
- `requires nfs-share` (or `smb-share`); several shares on one node are several local names of
the requirement (ADR 0094);
- its manifest is a `package` (nfs-utils), a `file` writing a systemd `.mount` unit filled from
the binding — `What=${bound:nfs-share:at}:${bound:nfs-share:path}` — and a `service` enabling it
after the overlay is up: the same shape as `resolv-conf` or `sshd`, files and a unit;
- *where* it mounts is the assignment's setting (`/srv/media` on one machine, elsewhere on
another); which machine mounts what is an operator decision made at assignment, exactly as which
paths a machine shares is.
**The modules that use the files never learn about NFS.** A player, an indexer, a backup declares
the mounted path as an `access` — an operator-chosen, pre-existing path the mesh never owns
(ADR 0051), exactly as `/storage/media` is on ace. The same app manifest then runs on ace against
the local library and on another node against the mounted one, with only its assignment differing.
**The one check to add, because it is the data-loss case.** An `access` is confirmed today by the
path being present. For a mountpoint that is not enough: a writer whose container starts before the
mount is up writes into the empty directory underneath it, and the files vanish when the mount
lands. The access check must confirm the path is *a mountpoint* when the module says so (or the
module's unit is ordered before the consumer's container — which crosses modules and is exactly
what the mesh does not order). Which of the two is the decision's.
**Identity crosses the wire.** `sec=sys` NFS trusts the client's uid, so a consumer must run as the
library's owner on the server (ace: `media`, 1001:2000) — hq 153's `${access:<id>:uid}`, read from
the mounted tree, answers it on the consumer's side too.
## Open questions for the decision
- Whether an NFS export over the overlay is an `internal` reach of the same endpoint or a second
export line — NFS authorises by client address, so the mesh range and the LAN range are two
entries in one file.
- How a consumer's binding expresses a *path* to mount (today bindings carry `at`, `port`, `as` and
whatever the provider `serves`), and whether one share can serve several paths.