Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ae58570209 | ||
|
|
855c3f9b33 | ||
|
|
f93b02900c | ||
|
|
9306192f93 | ||
|
|
8bd56a2820 | ||
|
|
7497149eed | ||
|
|
1b9196b807 | ||
|
|
b3e7191ddc | ||
|
|
87f80573b8 | ||
|
|
9197b99153 |
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: resolved
|
||||
status: open
|
||||
opened: 2026-10-04
|
||||
located-in: [mesh-controller cmd/mesh-controller/main.go, mesh-controller cmd/mesh-builder, mesh-controller internal/link/build.go]
|
||||
fixed-by: mesh-controller#47
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -35,11 +35,3 @@ acts on it, review becomes a formality: the unreviewed definition reaches a mach
|
||||
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
|
||||
any other build's?
|
||||
2. Did the controller roll the dry run out, or did a later push compose from it?
|
||||
|
||||
## Resolution (2026-10-05)
|
||||
|
||||
A dry run is marked on the request (`DryRun`), the builder echoes the mark on its outcome, and the
|
||||
controller's daemon sets a marked outcome aside: no record, no registration, no plan, nothing a push
|
||||
could send (mesh-controller#47, with a test that the daemon takes a dry run in with no store at all).
|
||||
Left open as a follow-up: the catalogue module also hears build outcomes and records their edges; it
|
||||
should skip a dry run too.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller internal/catalogue/seats.go, mesh-catalog modules/restic]
|
||||
fixed-by: mesh-controller#49, mesh-catalog#49, mesh-catalog#54, mesh-catalog#55, mesh-catalog#56, mesh-media-catalog#1
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
|
||||
---
|
||||
|
||||
@@ -47,27 +47,3 @@ research 030; proposed as ADR 0214.
|
||||
|
||||
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
|
||||
held on another machine, it would have been a ten-minute restore of yesterday.
|
||||
|
||||
## Where it stands (2026-10-05)
|
||||
|
||||
Backups run (ADR 0214, to-be 43): the control node keeps nightly restore points of its eight stores
|
||||
and services, the home server of its databases and its media apps' libraries and cover art; each on
|
||||
its own machine, on its larger filesystem. The first runs were tried one module, then two, then a
|
||||
whole node, each proven by a restore beside the live data. Not yet built, which is why this stays
|
||||
open: the weekly test restore into a throwaway instance, the 48-hour status line, and failures
|
||||
reaching the operator rather than the holder's log.
|
||||
|
||||
What the rollout taught, for the next module that takes contributions:
|
||||
|
||||
- **A contribution makes the contributing module depend on the seat.** Merging the stores'
|
||||
contributions before a holder was assigned left every node's plan unresolvable — the whole mesh,
|
||||
not only the machines running a store — until the holder was assigned. Assign the holder in the
|
||||
same step as the merge.
|
||||
- **A brand-new module is not built by the push that adds it**; its first build is asked for by hand.
|
||||
- **The holder must look as the account that can see.** Its first run called a store's dumps missing
|
||||
because it checked as the runtime's account, which cannot see inside the store's own directory.
|
||||
- **A kept single file is not a directory.** restic restores a snapshot's subfolder, not a file;
|
||||
the first restore of a module keeping its settings file refused it.
|
||||
- **An update kills a running backup** — the holder's restic is its child. A hand-off to the new
|
||||
version, the run living on as its own unit under the machine's service manager, is the idea to
|
||||
take forward.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
|
||||
fixed-by:
|
||||
fixed-by: mesh-catalog PR #45
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -31,3 +31,9 @@
|
||||
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
|
||||
same second it was published. The licences themselves were sound. The remaining licence refreshed
|
||||
on every attempt, and the second account's login was adopted from its first report.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
Live on all four machines and the manager the same morning. After the restart every machine reported
|
||||
the generation it held, and the manager's bindings matched them. The manager logged no failure while
|
||||
moving its sequence past them.
|
||||
|
||||
-39
@@ -1,39 +0,0 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-media-catalog modules/plex]
|
||||
fixed-by: mesh-media-catalog#2
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 245 — A media server's previews were reached through a link its container never mounted
|
||||
|
||||
## What was observed
|
||||
|
||||
On the home server the media server's scrub previews — 423 GB, about 57 000 preview files, one per
|
||||
video — had not grown in seven months: the newest was from the day its preview folder was moved off
|
||||
the server's own disk onto the large storage pool. The server's log repeated, live, that it could not
|
||||
create a directory under its preview folder.
|
||||
|
||||
## Why
|
||||
|
||||
The move left the preview folder as a **symbolic link** inside the server's configuration directory,
|
||||
pointing at the pool. The container mounts the configuration directory and the media libraries, and
|
||||
not the pool's path, so inside the container the link pointed at nothing: the server could neither
|
||||
show the previews it had nor make new ones, and said so only in its own log. Nothing the mesh reports
|
||||
showed it — the container ran, and answered.
|
||||
|
||||
It is the failure the mesh's rule against symbolic links exists for: a link resolves differently in
|
||||
every place that reads it, and a container is such a place.
|
||||
|
||||
## Resolution
|
||||
|
||||
The previews are a directory of the module, mounted at the server's preview path; where that
|
||||
directory lives on a machine is the machine's placement setting, here the pool path the files were
|
||||
already in. The link was removed at the cutover, the container recreated, the preview folders visible
|
||||
inside it again and the log's errors gone. The previews are not backed up, by the operator's choice.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The catalogue check that every mount is declared passes over the module. On the machine: the preview
|
||||
folder seen from inside the container lists the same folders as the pool path.
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 245 — `status` calls a module behind when only its repository moved
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
|
||||
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
|
||||
"`build --behind` builds them; `push --behind` sends them on".
|
||||
|
||||
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
|
||||
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
|
||||
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
|
||||
not change; only the repository's commit did.
|
||||
|
||||
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
|
||||
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
|
||||
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
|
||||
control node. Every node's runtime lost the bus for about a minute.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
|
||||
is held to a commit of its repository, so after any merge almost every module of that repository
|
||||
reads behind. The list then says nothing about what needs building: it hides the few modules that
|
||||
really are behind among the many that are not, and it invites a rebuild of everything, which is not
|
||||
a harmless act (above).
|
||||
|
||||
## The operator's direction (2026-10-05)
|
||||
|
||||
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
|
||||
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
|
||||
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
|
||||
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
|
||||
second answer to the same question is how the two came to disagree.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
|
||||
tiers?
|
||||
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
|
||||
move forward, or does the record stop carrying a commit that only says when it was last built?
|
||||
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
|
||||
never act on a module no plan named?
|
||||
-45
@@ -1,45 +0,0 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller cmd/mesh-controller (settings), mesh-controller internal/inventory (SetSettings)]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 246 — Setting a module's settings replaces the whole layer, and nothing shows it first
|
||||
|
||||
## What was observed
|
||||
|
||||
An operator's agent set one placement — where a media server's preview folder lives on one machine —
|
||||
with `settings set <module> {"places": {…}} --node <machine>`. The command answered that the setting
|
||||
was recorded. The machine's plan then mounted the server's configuration from an empty default
|
||||
directory: the node's layer had held the placements of three other directories, eight media
|
||||
accesses, a public exposure, four endpoints and the account's ids, and every one of them was gone.
|
||||
|
||||
Caught before any push, by reading the plan. The previous layer was read back from the controller
|
||||
database's nightly dump — the backups that issue 242 asked for, a few hours old.
|
||||
|
||||
## Why
|
||||
|
||||
The layer is a statement of the whole, by design (the inventory's `SetSettings`: "replacing rather
|
||||
than merging … removing a key is done by leaving it out"). That design is sound; what is missing
|
||||
around it is everything that makes it safe to use:
|
||||
|
||||
- **There is no way to read a layer.** `settings` has `set` and `clear`, no `show`; the console verb
|
||||
likewise. To change one key, a person must already know every other key in the layer.
|
||||
- **There is no history.** The row is updated in place; the previous values exist nowhere but a
|
||||
database backup.
|
||||
- **The answer does not say what was dropped.** "places" was reported as set; the six keys removed
|
||||
were not mentioned.
|
||||
|
||||
## What would fix it
|
||||
|
||||
1. A way to read a layer — `settings show <module> [--node]` and the same on the verb.
|
||||
2. `set` answers with what changed: keys added, changed and **removed**, so dropping one is never
|
||||
silent. A removal could even require saying so.
|
||||
3. The previous value kept: a settings history row per change, so an undo needs no backup.
|
||||
|
||||
## Status
|
||||
|
||||
Open. Until it is fixed: read the layer (from the store, read-only) before setting it, and compare
|
||||
the node's plan before and after.
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-tools]
|
||||
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 246 — The console says a module runs nowhere when a runtime answers late
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
|
||||
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
|
||||
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
|
||||
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
|
||||
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
|
||||
|
||||
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
|
||||
answer and bypasses everything the console stands for: one way in, one account, one record of what
|
||||
was called. The operator asked for a tool instead.
|
||||
|
||||
## What was measured
|
||||
|
||||
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
|
||||
against the live bus, whose round trip from the laptop was about 40 ms:
|
||||
|
||||
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
|
||||
is far below the bus's message limit, and it was never shortened.
|
||||
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
|
||||
answered within 180 to 275 ms, and the controller within 50 ms.
|
||||
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
|
||||
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
|
||||
while a runtime re-serves after a restart.
|
||||
|
||||
## Root cause
|
||||
|
||||
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
|
||||
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
|
||||
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
|
||||
|
||||
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
|
||||
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
|
||||
answers dropping a machine.
|
||||
|
||||
## Resolution
|
||||
|
||||
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
|
||||
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
|
||||
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
|
||||
- A runtime that said it is there and did not say what it serves in time, or that answered recently
|
||||
and not now, is named. While one is unheard, the console never says an address is missing or runs
|
||||
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
|
||||
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
|
||||
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
|
||||
and which runtimes or machines were not heard.
|
||||
- An announcement still too large after its descriptions are cut to their first line now leaves the
|
||||
descriptions out, and says so.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
|
||||
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
|
||||
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
|
||||
`mesh_runtimes` shows every machine's answer and its time.
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 247 — A module cannot put the operator's account in a group
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
|
||||
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
|
||||
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
|
||||
devices.
|
||||
|
||||
The module's own check names the fix, which is to add the account to the group and log in again. No
|
||||
module can declare that fix:
|
||||
|
||||
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
|
||||
- A second module that declares the same account, only to add one group, is refused as a duplicate
|
||||
name.
|
||||
- There is no resource for one membership on its own. A whole-account declaration that lists groups
|
||||
would also take from the account every group it does not list, including the operator's own.
|
||||
|
||||
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
|
||||
the group.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
|
||||
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
|
||||
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
|
||||
reinstall only by memory. A membership added by hand is also never taken away when the module that
|
||||
needed it is unassigned.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Is a membership its own resource (account, group), held by the module that needs it and given
|
||||
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
|
||||
module contribute to a seat?
|
||||
2. A membership takes effect at the next login. How does the module say so: a finding, or a
|
||||
moment the power or session seat already knows?
|
||||
3. What does undeclare do with a membership the account already had before any module declared it?
|
||||
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
|
||||
|
||||
## How it is checked
|
||||
|
||||
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
|
||||
account in the group, says that a new login is needed, and leaves the account's other groups as they
|
||||
were. Unassigning it removes only a membership the module added.
|
||||
Reference in New Issue
Block a user