10 Commits
Author SHA1 Message Date
mesh-admin ae58570209 Merge pull request 'Issue 247: a module cannot put the operator's account in a group' (#99) from issues/247-a-module-cannot-put-the-operator-in-a-group into main 2026-10-05 13:45:52 +00:00
mesh-admin 855c3f9b33 Merge pull request 'Issue 246: the console says a module runs nowhere when a runtime answers late' (#98) from issues/246-the-console-loses-the-largest-runtime into main 2026-10-05 13:45:49 +00:00
jochen f93b02900c Issue 247: a module cannot put the operator's account in a group 2026-10-05 15:45:24 +02:00
jochen 9306192f93 Issue 246: the console says a module runs nowhere when a runtime answers late 2026-10-05 15:40:14 +02:00
mesh-admin 8bd56a2820 Merge pull request 'Issue 243: resolved' (#87) from issue/243-resolved into main 2026-10-05 13:17:44 +00:00
mesh-admin 7497149eed Merge pull request 'Issue 245: behind is what a plan decided and has not yet done' (#96) from issues/245-the-plan-is-the-truth into main 2026-10-05 13:17:04 +00:00
jochen 1b9196b807 Issue 245: behind is what a plan decided and has not yet done (the operator's direction) 2026-10-05 15:16:51 +02:00
mesh-admin b3e7191ddc Merge pull request 'Issue 245: status calls a module behind when only its repository moved' (#95) from issues/245-behind-means-its-files-changed into main 2026-10-05 13:15:57 +00:00
jochen 87f80573b8 Issue 245: status calls a module behind when only its repository moved 2026-10-05 15:15:44 +02:00
jochen 9197b99153 Issue 243: resolved by nodes applying any differing binding and the manager never repeating a generation 2026-10-05 10:03:49 +02:00
9 changed files with 181 additions and 124 deletions
@@ -1,8 +1,8 @@
---
status: resolved
status: open
opened: 2026-10-04
located-in: [mesh-controller cmd/mesh-controller/main.go, mesh-controller cmd/mesh-builder, mesh-controller internal/link/build.go]
fixed-by: mesh-controller#47
located-in: []
fixed-by:
amended-design:
---
@@ -35,11 +35,3 @@ acts on it, review becomes a formality: the unreviewed definition reaches a mach
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it?
## Resolution (2026-10-05)
A dry run is marked on the request (`DryRun`), the builder echoes the mark on its outcome, and the
controller's daemon sets a marked outcome aside: no record, no registration, no plan, nothing a push
could send (mesh-controller#47, with a test that the daemon takes a dry run in with no store at all).
Left open as a follow-up: the catalogue module also hears build outcomes and records their edges; it
should skip a dry run too.
@@ -1,8 +1,8 @@
---
status: located
status: open
opened: 2026-10-05
located-in: [mesh-controller internal/catalogue/seats.go, mesh-catalog modules/restic]
fixed-by: mesh-controller#49, mesh-catalog#49, mesh-catalog#54, mesh-catalog#55, mesh-catalog#56, mesh-media-catalog#1
located-in: []
fixed-by:
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
---
@@ -47,27 +47,3 @@ research 030; proposed as ADR 0214.
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.
## Where it stands (2026-10-05)
Backups run (ADR 0214, to-be 43): the control node keeps nightly restore points of its eight stores
and services, the home server of its databases and its media apps' libraries and cover art; each on
its own machine, on its larger filesystem. The first runs were tried one module, then two, then a
whole node, each proven by a restore beside the live data. Not yet built, which is why this stays
open: the weekly test restore into a throwaway instance, the 48-hour status line, and failures
reaching the operator rather than the holder's log.
What the rollout taught, for the next module that takes contributions:
- **A contribution makes the contributing module depend on the seat.** Merging the stores'
contributions before a holder was assigned left every node's plan unresolvable — the whole mesh,
not only the machines running a store — until the holder was assigned. Assign the holder in the
same step as the merge.
- **A brand-new module is not built by the push that adds it**; its first build is asked for by hand.
- **The holder must look as the account that can see.** Its first run called a store's dumps missing
because it checked as the runtime's account, which cannot see inside the store's own directory.
- **A kept single file is not a directory.** restic restores a snapshot's subfolder, not a file;
the first restore of a module keeping its settings file refused it.
- **An update kills a running backup** — the holder's restic is its child. A hand-off to the new
version, the run living on as its own unit under the machine's service manager, is the idea to
take forward.
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
fixed-by:
fixed-by: mesh-catalog PR #45
amended-design:
---
@@ -31,3 +31,9 @@
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
same second it was published. The licences themselves were sound. The remaining licence refreshed
on every attempt, and the second account's login was adopted from its first report.
## Resolved — 2026-10-05
Live on all four machines and the manager the same morning. After the restart every machine reported
the generation it held, and the manager's bindings matched them. The manager logged no failure while
moving its sequence past them.
@@ -1,39 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-media-catalog modules/plex]
fixed-by: mesh-media-catalog#2
amended-design:
---
# 245 — A media server's previews were reached through a link its container never mounted
## What was observed
On the home server the media server's scrub previews — 423 GB, about 57 000 preview files, one per
video — had not grown in seven months: the newest was from the day its preview folder was moved off
the server's own disk onto the large storage pool. The server's log repeated, live, that it could not
create a directory under its preview folder.
## Why
The move left the preview folder as a **symbolic link** inside the server's configuration directory,
pointing at the pool. The container mounts the configuration directory and the media libraries, and
not the pool's path, so inside the container the link pointed at nothing: the server could neither
show the previews it had nor make new ones, and said so only in its own log. Nothing the mesh reports
showed it — the container ran, and answered.
It is the failure the mesh's rule against symbolic links exists for: a link resolves differently in
every place that reads it, and a container is such a place.
## Resolution
The previews are a directory of the module, mounted at the server's preview path; where that
directory lives on a machine is the machine's placement setting, here the pool path the files were
already in. The link was removed at the cutover, the container recreated, the preview folders visible
inside it again and the log's errors gone. The previews are not backed up, by the operator's choice.
## How it is checked
The catalogue check that every mount is declared passes over the module. On the machine: the preview
folder seen from inside the container lists the same folders as the pool path.
@@ -0,0 +1,50 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 245 — `status` calls a module behind when only its repository moved
## What was observed
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
"`build --behind` builds them; `push --behind` sends them on".
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
not change; only the repository's commit did.
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
control node. Every node's runtime lost the bus for about a minute.
## Why it matters beyond this instance
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
is held to a commit of its repository, so after any merge almost every module of that repository
reads behind. The list then says nothing about what needs building: it hides the few modules that
really are behind among the many that are not, and it invites a rebuild of everything, which is not
a harmless act (above).
## The operator's direction (2026-10-05)
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
second answer to the same question is how the two came to disagree.
## Open questions
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
tiers?
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
move forward, or does the record stop carrying a commit that only says when it was last built?
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
never act on a module no plan named?
@@ -1,45 +0,0 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller (settings), mesh-controller internal/inventory (SetSettings)]
fixed-by:
amended-design:
---
# 246 — Setting a module's settings replaces the whole layer, and nothing shows it first
## What was observed
An operator's agent set one placement — where a media server's preview folder lives on one machine —
with `settings set <module> {"places": {…}} --node <machine>`. The command answered that the setting
was recorded. The machine's plan then mounted the server's configuration from an empty default
directory: the node's layer had held the placements of three other directories, eight media
accesses, a public exposure, four endpoints and the account's ids, and every one of them was gone.
Caught before any push, by reading the plan. The previous layer was read back from the controller
database's nightly dump — the backups that issue 242 asked for, a few hours old.
## Why
The layer is a statement of the whole, by design (the inventory's `SetSettings`: "replacing rather
than merging … removing a key is done by leaving it out"). That design is sound; what is missing
around it is everything that makes it safe to use:
- **There is no way to read a layer.** `settings` has `set` and `clear`, no `show`; the console verb
likewise. To change one key, a person must already know every other key in the layer.
- **There is no history.** The row is updated in place; the previous values exist nowhere but a
database backup.
- **The answer does not say what was dropped.** "places" was reported as set; the six keys removed
were not mentioned.
## What would fix it
1. A way to read a layer — `settings show <module> [--node]` and the same on the verb.
2. `set` answers with what changed: keys added, changed and **removed**, so dropping one is never
silent. A removal could even require saying so.
3. The previous value kept: a settings history row per change, so an undo needs no backup.
## Status
Open. Until it is fixed: read the layer (from the store, read-only) before setting it, and compare
the node's plan before and after.
@@ -0,0 +1,65 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-tools]
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
amended-design:
---
# 246 — The console says a module runs nowhere when a runtime answers late
## What was observed
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
answer and bypasses everything the console stands for: one way in, one account, one record of what
was called. The operator asked for a tool instead.
## What was measured
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
against the live bus, whose round trip from the laptop was about 40 ms:
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
is far below the bus's message limit, and it was never shortened.
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
answered within 180 to 275 ms, and the controller within 50 ms.
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
while a runtime re-serves after a restart.
## Root cause
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
answers dropping a machine.
## Resolution
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
- A runtime that said it is there and did not say what it serves in time, or that answered recently
and not now, is named. While one is unheard, the console never says an address is missing or runs
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
and which runtimes or machines were not heard.
- An announcement still too large after its descriptions are cut to their first line now leaves the
descriptions out, and says so.
## How it is checked
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
`mesh_runtimes` shows every machine's answer and its time.
@@ -0,0 +1,52 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 247 — A module cannot put the operator's account in a group
## What was observed
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
devices.
The module's own check names the fix, which is to add the account to the group and log in again. No
module can declare that fix:
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
- A second module that declares the same account, only to add one group, is refused as a duplicate
name.
- There is no resource for one membership on its own. A whole-account declaration that lists groups
would also take from the account every group it does not list, including the operator's own.
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
the group.
## Why it matters beyond this instance
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
reinstall only by memory. A membership added by hand is also never taken away when the module that
needed it is unassigned.
## Open questions
1. Is a membership its own resource (account, group), held by the module that needs it and given
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
module contribute to a seat?
2. A membership takes effect at the next login. How does the module say so: a finding, or a
moment the power or session seat already knows?
3. What does undeclare do with a membership the account already had before any module declared it?
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
## How it is checked
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
account in the group, says that a new login is needed, and leaves the account's other groups as they
were. Unassigning it removes only a membership the module added.