Compare commits

..
Author SHA1 Message Date
jschoubben 49b1136ded Issues 164-168: found provisioning every dependency on ace
164 a credential that must be accepted is minted anyway
165 one accepted value must be accepted once per consumer
166 a requirement cannot be optional
167 code several modules share has no home
168 a setting reaches every file and every contribution
2026-09-30 13:28:53 +02:00
7 changed files with 135 additions and 184 deletions
@@ -1,56 +0,0 @@
---
status: resolved
opened: 2026-09-30
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
amended-design:
---
# 163 — A delivered host stood aside on every push, and reported nothing
## What was observed
*2026-09-30, rolling the mesh-built host onto the last two machines.*
Every push to a machine running a delivered host produced, in order:
```
host 093231796eb0 is delivered; standing aside so the launcher runs it
applied 333 resource(s)
applied, and could not tell the mesh: reporting: context canceled
nox-mesh-host-launch: the host exited cleanly; starting it again
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
```
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
never received a single report from it: `node show` kept the version from before the crossover, and
the operator's push waited its full three minutes for an answer that was never coming.
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
## Why
After an apply the host asks whether a newer host has been delivered than the one running, and the
question was asked with the **link-time version stamp**. Since
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
host's version comes from where it sits and its stamp is `development build` — so the comparison never
matched the newest delivered version, and "a newer host is waiting" was always true.
Standing aside cancels the context the report is published with, so the report was lost on every one
of those applies. Two faults from one wrong argument.
The change that moved the version to the path was applied to the report and to the known-good record,
and not here. Half a change, and the half left behind was the one that decides whether to exit.
## Why the three-minute wait made it invisible
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
answered.
## How it is checked
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
version; it stands aside once, and the next push it does not.
@@ -0,0 +1,29 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller internal/inventory/secrets.go (SecretFor mints a pair credential nobody accepted)
fixed-by:
amended-design:
---
# 164 — A credential that must be accepted is minted anyway
## What was observed
Provisioning ace's modules. Several providers hold exactly one credential they did not get from the
mesh and cannot take one from it: a Servarr app's API key (sonarr, radarr, lidarr), jackett's API key,
plex's X-Plex-Token, nzbget's ControlPassword, qBittorrent's WebUI password. Their consumers' pair
credential must be **accepted** by the operator (ADR 0092). Until it is, `SecretFor` mints a random
value, seals it to both ends, and reports nothing: the value can never work.
Every consumer therefore had to learn to detect it — try the credential against the provider first,
refuse a value the provider rejects, print the `secret accept` command — six write-in steps, one probe
each (ombi, home-assistant, and the four download-stack consumers). qBittorrent bans an address after
five failed logins, so a consumer retrying a minted value locks itself out.
## What would be right
A provision (or a provider's `serves`) can declare its pair credential **accepted-only**. The plan then
refuses the pair — naming the accept command — instead of minting, and a consumer is never handed a
value the mesh knows cannot work.
@@ -0,0 +1,24 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller internal/inventory/secrets.go (AcceptSecretForPair is per consumer)
fixed-by:
amended-design:
---
# 165 — One accepted value must be accepted once per consumer
## What was observed
On ace, jackett's API key is the pair credential for sonarr, radarr, lidarr and bookshelf; sonarr's is
the credential for ombi, bazarr and home-assistant. It is **one value**, owned by the provider — yet
`secret accept` is per pair, so ace's download stack alone needs 12 accepts of 3 values, and rotating
a provider's key means finding and re-accepting every pair. Missing one leaves that consumer on a
stale (or minted, 164) value.
## What would be right
A provider-level accept: "this provider's credential for `<provision>` is X" — delivered to every
consumer pair, current and future, and rotated in one place. Pairs whose credential is genuinely per
consumer (postgres, keycloak, mosquitto, influxdb — minted and created by a provisioner) are unaffected.
@@ -0,0 +1,24 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue (requires is a list of hard requirements)
fixed-by:
amended-design:
---
# 166 — A requirement cannot be optional
## What was observed
Making every dependency on ace a provision turned soft dependencies into hard ones. grafana now
requires `influxdb-api` (a data source), ombi requires `sonarr-api`, `radarr-api` and `lidarr-api`,
home-assistant requires the Servarr APIs and `mqtt-topic`. Each is optional to the software — grafana
runs without a data source, ombi without lidarr — but a mesh without influxdb cannot assign grafana at
all, and a mesh without lidarr cannot run ombi.
## What would be right
A requirement a module can run without: resolved and bound when a provider exists, absent (with its
`${bound:…}` placeholders refused or defaulted explicitly, never rendered empty) when none does — so
the module description stays true on every mesh.
@@ -0,0 +1,26 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-catalog (each module builds from its own directory, ADR 0069)
- mesh-sdk
fixed-by:
amended-design:
---
# 167 — Code several modules share has no home
## What was observed
The download-stack write-in step (register download clients and torznab indexers through the Servarr
API) is identical for sonarr, radarr, lidarr and bookshelf. Because a module builds from its own
directory, it now exists as four byte-identical copies under `modules/<m>/downloads/`, kept honest by a
test that fails when one differs. The same shape repeats: an MQTT probe copied into two modules, and a
"write the provider into the app through its API, idempotently, refuse a minted value" step in ombi,
home-assistant, nodered, tautulli and the four downloaders.
## What would be right
A home for shared module code the builder can use — an sdk helper (a write-in step harness: read
bindings and pair credentials, probe the provider, diff, write, report) or a shared package the
catalogue builds once — so a fix lands in one place.
@@ -0,0 +1,32 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/settings.go (settle: every key but `ports` merges into every mergeable file and every contribution)
- mesh-controller internal/catalogue/declaration.go (a provider's settings are laid over what it serves)
fixed-by:
amended-design:
---
# 168 — A setting reaches every file and every contribution
## What was observed
Settings merge key by key into **every** `"merge": "json"` file of a module **and** every contribution
it makes; a provider's settings are also laid over what it serves. Seen on ace:
- searxng's `endpoints` and a route `label` land in searxng's own `settings.yml`; nodered's
`timeZone` and `mqtt` keys land in mosquitto's grants file; keycloak's `issuer` lands in its
`postgres-database` and `route` contributions.
- every consumer's `plex-api` binding carries plex's `endpoints` and `expose` settings — and a provider
setting named `port` would silently redirect every consumer.
- a module cannot have two configurable files: searxng's sidecar config had to stop being mergeable
so searxng's keys would not reach it.
Harmless today only because every receiver happens to ignore unknown keys.
## What would be right
A setting is aimed: at a file (by resource id), at a contribution (by requirement), or at what the
module serves — declared settable by the module (ADR 0046 already says settings drive "the fields the
manifest marks") — and an unaimed key is refused like any unknown setting.
@@ -1,128 +0,0 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-catalog (no module shares a path over the network)
- hq 02-DECISIONS (a file-share seat, per ADR 0126, is a module's own to define)
fixed-by:
amended-design:
---
# 169 — A machine shares its files, and the mesh does not know
## What was observed
ace serves the operator's media library to the home network with two host services no module
declares and HAL never managed either:
```
/etc/exports: /storage/media 192.168.1.0/24(rw,sync,root_squash,…) nfs-server active, :2049
/etc/samba/smb.conf: [media] path = /storage/media/ valid users = media smb active, :139/:445
```
Two LAN clients were connected at survey (2026-09-30). The library itself is operator data
(ADR 0051: ~40 TB on ZFS, the mesh owns nothing about it — [issue 153](../153-an-adopted-machines-data-cannot-be-placed-where-it-is/00-report.md)
is about modules reaching it in place).
Under the mesh as it stands, this arrangement has no expression and one failure mode:
- **Nothing declares the listens.** At `converge ace` the filter is the sum of what modules listen
on (ADR 0045); 2049 and 445 are nobody's, so the shares close — silently, for the two clients
that mount them.
- **Nothing owns the configuration.** `/etc/exports` and `smb.conf` are hand-written files on one
machine; a second machine sharing a directory would be written by hand again.
- **Nothing can consume it.** A module on another node that wanted the library (a player, an
indexer, a backup) has no `requires` to state and no binding to read; it would mount by a
hand-typed host and path.
- The clients are LAN devices, so this also meets [issue 154](../154-a-machines-own-network-is-not-a-reach/00-report.md)
(no reach for the machine's own network).
## The proposal (the operator's, 2026-09-30, settled after two rounds)
**Two module-defined seats, one per protocol, because NFS and SMB share an intent and not a
contract.** A seat in the mesh's sense is a contract — what it accepts, emits and serves, and the
tools its holder must answer (ADR 0126, 0132) — and lined up, the two share almost none of it:
| | `nfs-share` | `smb-share` |
|---|---|---|
| serves | export path(s); the client ranges allowed (`sec=sys` authorises by address) | share name(s), path |
| pair credential | none | a user and password per consumer |
| consumer's mount | `at:/path` | `//at/share` with credentials |
| holder's tools | export / unexport a path for a range | add / remove a share, create a user |
One `file-share` seat would be the union with every field optional — a consumer could bind it and
still not know how to mount what it got (the emptiness ADR 0129 warns against). "Export a path to
the network" is a category, and the mesh needs no seat category: a consumer requires the one it
can mount. If "give me the library, however" is ever needed, it is a provision an umbrella module
serves, not a seat.
Both are node-scoped, one holder per node (ADR 0110), so ace holds both. `nfs` and `samba` are the
first implementations; a second (Ganesha for `nfs-share`, ksmbd for `smb-share`) is what proves
0126's promise that "replacing the implementation changes nothing for any caller".
The holder module:
- declares the exported paths as `accesses` (ADR 0051: it owns nothing about them — never creates,
chowns or removes), and *which* paths as the assignment's settings (ADR 0046/0112);
- writes the share configuration (`/etc/exports`, `smb.conf`) as mesh-managed files and drives the
units, like `dnsmasq`/`sshd` do for theirs;
- declares its endpoints (`nfs` 2049/tcp; `smb` 445/tcp, …) so the reach — internal, or the LAN
once 154 has an answer — is the assignment's, and converge keeps them open;
- **provides** the seat's provision, so a consumer on another node `requires nfs-share` (or
`smb-share`) and reads `${bound:nfs-share:at}` and the path from its binding instead of a
hand-typed mount.
## The design gap this exposes
**A seat definition has no home outside the module that first declared it.** Today a seat is
declared inside a manifest (`showcase` declares `the-showcase`, `ca-trust` its own). If `nfs`
declared `nfs-share`, Ganesha could hold it only by depending on nfs's manifest — the coupling
0126 removed for callers, reintroduced for implementations. The protocol needs a neutral place in
the catalogue beside the modules (a seat definition registered like a manifest), with a module
saying which seats it implements. This is the first role with an obvious second implementation,
which is what makes it the exemplar for that mechanism.
## The consumer's half: a module mounts it (2026-09-30, third and fourth round)
A binding tells a consumer *where* the share is; it does not put the files on its machine. Mounting
is something done on a machine, and something done on a machine is a module's work — not the host's
(the vocabulary stays closed; no `mount` resource kind).
**A consumer-side module, `network-share` — the module responsible for setting up the network
shares a node uses** (the operator's framing). A node role, like `node-uplink` or
`node-dns-resolver`: each machine has it at most once, which is a reason for it to hold a
node-scoped seat, so two modules can never both be writing mount units on one machine. Assigned on
the node that wants the files:
- `requires nfs-share` (or `smb-share`); several shares on one node are several local names of
the requirement (ADR 0094);
- its manifest is a `package` (nfs-utils), a `file` writing a systemd `.mount` unit filled from
the binding — `What=${bound:nfs-share:at}:${bound:nfs-share:path}` — and a `service` enabling it
after the overlay is up: the same shape as `resolv-conf` or `sshd`, files and a unit;
- *where* it mounts is the assignment's setting (`/srv/media` on one machine, elsewhere on
another); which machine mounts what is an operator decision made at assignment, exactly as which
paths a machine shares is.
**The modules that use the files never learn about NFS.** A player, an indexer, a backup declares
the mounted path as an `access` — an operator-chosen, pre-existing path the mesh never owns
(ADR 0051), exactly as `/storage/media` is on ace. The same app manifest then runs on ace against
the local library and on another node against the mounted one, with only its assignment differing.
**The one check to add, because it is the data-loss case.** An `access` is confirmed today by the
path being present. For a mountpoint that is not enough: a writer whose container starts before the
mount is up writes into the empty directory underneath it, and the files vanish when the mount
lands. The access check must confirm the path is *a mountpoint* when the module says so (or the
module's unit is ordered before the consumer's container — which crosses modules and is exactly
what the mesh does not order). Which of the two is the decision's.
**Identity crosses the wire.** `sec=sys` NFS trusts the client's uid, so a consumer must run as the
library's owner on the server (ace: `media`, 1001:2000) — hq 153's `${access:<id>:uid}`, read from
the mounted tree, answers it on the consumer's side too.
## Open questions for the decision
- Whether an NFS export over the overlay is an `internal` reach of the same endpoint or a second
export line — NFS authorises by client address, so the mesh range and the LAN range are two
entries in one file.
- How a consumer's binding expresses a *path* to mount (today bindings carry `at`, `port`, `as` and
whatever the provider `serves`), and whether one share can serve several paths.