1 Commits
Author SHA1 Message Date
jschoubben d222091fe9 Issues 245, 246; 240 resolved, 242 located
245: a media server's previews were reached through a link its container never mounted — fixed by
mounting them. 246: settings set replaces the whole layer with nothing to read it first and no
history. 240 is fixed by mesh-controller#47; 242 has backups running and records what the rollout
taught.
2026-10-05 14:54:38 +02:00
4 changed files with 122 additions and 6 deletions
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-10-04
located-in: []
fixed-by:
located-in: [mesh-controller cmd/mesh-controller/main.go, mesh-controller cmd/mesh-builder, mesh-controller internal/link/build.go]
fixed-by: mesh-controller#47
amended-design:
---
@@ -35,3 +35,11 @@ acts on it, review becomes a formality: the unreviewed definition reaches a mach
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it?
## Resolution (2026-10-05)
A dry run is marked on the request (`DryRun`), the builder echoes the mark on its outcome, and the
controller's daemon sets a marked outcome aside: no record, no registration, no plan, nothing a push
could send (mesh-controller#47, with a test that the daemon takes a dry run in with no store at all).
Left open as a follow-up: the catalogue module also hears build outcomes and records their edges; it
should skip a dry run too.
@@ -1,8 +1,8 @@
---
status: open
status: located
opened: 2026-10-05
located-in: []
fixed-by:
located-in: [mesh-controller internal/catalogue/seats.go, mesh-catalog modules/restic]
fixed-by: mesh-controller#49, mesh-catalog#49, mesh-catalog#54, mesh-catalog#55, mesh-catalog#56, mesh-media-catalog#1
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
---
@@ -47,3 +47,27 @@ research 030; proposed as ADR 0214.
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.
## Where it stands (2026-10-05)
Backups run (ADR 0214, to-be 43): the control node keeps nightly restore points of its eight stores
and services, the home server of its databases and its media apps' libraries and cover art; each on
its own machine, on its larger filesystem. The first runs were tried one module, then two, then a
whole node, each proven by a restore beside the live data. Not yet built, which is why this stays
open: the weekly test restore into a throwaway instance, the 48-hour status line, and failures
reaching the operator rather than the holder's log.
What the rollout taught, for the next module that takes contributions:
- **A contribution makes the contributing module depend on the seat.** Merging the stores'
contributions before a holder was assigned left every node's plan unresolvable — the whole mesh,
not only the machines running a store — until the holder was assigned. Assign the holder in the
same step as the merge.
- **A brand-new module is not built by the push that adds it**; its first build is asked for by hand.
- **The holder must look as the account that can see.** Its first run called a store's dumps missing
because it checked as the runtime's account, which cannot see inside the store's own directory.
- **A kept single file is not a directory.** restic restores a snapshot's subfolder, not a file;
the first restore of a module keeping its settings file refused it.
- **An update kills a running backup** — the holder's restic is its child. A hand-off to the new
version, the run living on as its own unit under the machine's service manager, is the idea to
take forward.
@@ -0,0 +1,39 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-media-catalog modules/plex]
fixed-by: mesh-media-catalog#2
amended-design:
---
# 245 — A media server's previews were reached through a link its container never mounted
## What was observed
On the home server the media server's scrub previews — 423 GB, about 57 000 preview files, one per
video — had not grown in seven months: the newest was from the day its preview folder was moved off
the server's own disk onto the large storage pool. The server's log repeated, live, that it could not
create a directory under its preview folder.
## Why
The move left the preview folder as a **symbolic link** inside the server's configuration directory,
pointing at the pool. The container mounts the configuration directory and the media libraries, and
not the pool's path, so inside the container the link pointed at nothing: the server could neither
show the previews it had nor make new ones, and said so only in its own log. Nothing the mesh reports
showed it — the container ran, and answered.
It is the failure the mesh's rule against symbolic links exists for: a link resolves differently in
every place that reads it, and a container is such a place.
## Resolution
The previews are a directory of the module, mounted at the server's preview path; where that
directory lives on a machine is the machine's placement setting, here the pool path the files were
already in. The link was removed at the cutover, the container recreated, the preview folders visible
inside it again and the log's errors gone. The previews are not backed up, by the operator's choice.
## How it is checked
The catalogue check that every mount is declared passes over the module. On the machine: the preview
folder seen from inside the container lists the same folders as the pool path.
@@ -0,0 +1,45 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller (settings), mesh-controller internal/inventory (SetSettings)]
fixed-by:
amended-design:
---
# 246 — Setting a module's settings replaces the whole layer, and nothing shows it first
## What was observed
An operator's agent set one placement — where a media server's preview folder lives on one machine —
with `settings set <module> {"places": {…}} --node <machine>`. The command answered that the setting
was recorded. The machine's plan then mounted the server's configuration from an empty default
directory: the node's layer had held the placements of three other directories, eight media
accesses, a public exposure, four endpoints and the account's ids, and every one of them was gone.
Caught before any push, by reading the plan. The previous layer was read back from the controller
database's nightly dump — the backups that issue 242 asked for, a few hours old.
## Why
The layer is a statement of the whole, by design (the inventory's `SetSettings`: "replacing rather
than merging … removing a key is done by leaving it out"). That design is sound; what is missing
around it is everything that makes it safe to use:
- **There is no way to read a layer.** `settings` has `set` and `clear`, no `show`; the console verb
likewise. To change one key, a person must already know every other key in the layer.
- **There is no history.** The row is updated in place; the previous values exist nowhere but a
database backup.
- **The answer does not say what was dropped.** "places" was reported as set; the six keys removed
were not mentioned.
## What would fix it
1. A way to read a layer — `settings show <module> [--node]` and the same on the verb.
2. `set` answers with what changed: keys added, changed and **removed**, so dropping one is never
silent. A removal could even require saying so.
3. The previous value kept: a settings history row per change, so an undo needs no backup.
## Status
Open. Until it is fixed: read the layer (from the store, read-only) before setting it, and compare
the node's plan before and after.