Compare commits
10
Commits
issues/245-246
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ae58570209 | ||
|
|
855c3f9b33 | ||
|
|
f93b02900c | ||
|
|
9306192f93 | ||
|
|
8bd56a2820 | ||
|
|
7497149eed | ||
|
|
1b9196b807 | ||
|
|
b3e7191ddc | ||
|
|
87f80573b8 | ||
|
|
9197b99153 |
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
|
||||
fixed-by:
|
||||
fixed-by: mesh-catalog PR #45
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -31,3 +31,9 @@
|
||||
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
|
||||
same second it was published. The licences themselves were sound. The remaining licence refreshed
|
||||
on every attempt, and the second account's login was adopted from its first report.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
Live on all four machines and the manager the same morning. After the restart every machine reported
|
||||
the generation it held, and the manager's bindings matched them. The manager logged no failure while
|
||||
moving its sequence past them.
|
||||
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 245 — `status` calls a module behind when only its repository moved
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
|
||||
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
|
||||
"`build --behind` builds them; `push --behind` sends them on".
|
||||
|
||||
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
|
||||
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
|
||||
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
|
||||
not change; only the repository's commit did.
|
||||
|
||||
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
|
||||
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
|
||||
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
|
||||
control node. Every node's runtime lost the bus for about a minute.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
|
||||
is held to a commit of its repository, so after any merge almost every module of that repository
|
||||
reads behind. The list then says nothing about what needs building: it hides the few modules that
|
||||
really are behind among the many that are not, and it invites a rebuild of everything, which is not
|
||||
a harmless act (above).
|
||||
|
||||
## The operator's direction (2026-10-05)
|
||||
|
||||
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
|
||||
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
|
||||
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
|
||||
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
|
||||
second answer to the same question is how the two came to disagree.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
|
||||
tiers?
|
||||
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
|
||||
move forward, or does the record stop carrying a commit that only says when it was last built?
|
||||
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
|
||||
never act on a module no plan named?
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-tools]
|
||||
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 246 — The console says a module runs nowhere when a runtime answers late
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
|
||||
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
|
||||
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
|
||||
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
|
||||
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
|
||||
|
||||
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
|
||||
answer and bypasses everything the console stands for: one way in, one account, one record of what
|
||||
was called. The operator asked for a tool instead.
|
||||
|
||||
## What was measured
|
||||
|
||||
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
|
||||
against the live bus, whose round trip from the laptop was about 40 ms:
|
||||
|
||||
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
|
||||
is far below the bus's message limit, and it was never shortened.
|
||||
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
|
||||
answered within 180 to 275 ms, and the controller within 50 ms.
|
||||
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
|
||||
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
|
||||
while a runtime re-serves after a restart.
|
||||
|
||||
## Root cause
|
||||
|
||||
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
|
||||
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
|
||||
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
|
||||
|
||||
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
|
||||
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
|
||||
answers dropping a machine.
|
||||
|
||||
## Resolution
|
||||
|
||||
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
|
||||
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
|
||||
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
|
||||
- A runtime that said it is there and did not say what it serves in time, or that answered recently
|
||||
and not now, is named. While one is unheard, the console never says an address is missing or runs
|
||||
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
|
||||
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
|
||||
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
|
||||
and which runtimes or machines were not heard.
|
||||
- An announcement still too large after its descriptions are cut to their first line now leaves the
|
||||
descriptions out, and says so.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
|
||||
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
|
||||
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
|
||||
`mesh_runtimes` shows every machine's answer and its time.
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 247 — A module cannot put the operator's account in a group
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
|
||||
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
|
||||
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
|
||||
devices.
|
||||
|
||||
The module's own check names the fix, which is to add the account to the group and log in again. No
|
||||
module can declare that fix:
|
||||
|
||||
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
|
||||
- A second module that declares the same account, only to add one group, is refused as a duplicate
|
||||
name.
|
||||
- There is no resource for one membership on its own. A whole-account declaration that lists groups
|
||||
would also take from the account every group it does not list, including the operator's own.
|
||||
|
||||
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
|
||||
the group.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
|
||||
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
|
||||
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
|
||||
reinstall only by memory. A membership added by hand is also never taken away when the module that
|
||||
needed it is unassigned.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Is a membership its own resource (account, group), held by the module that needs it and given
|
||||
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
|
||||
module contribute to a seat?
|
||||
2. A membership takes effect at the next login. How does the module say so: a finding, or a
|
||||
moment the power or session seat already knows?
|
||||
3. What does undeclare do with a membership the account already had before any module declared it?
|
||||
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
|
||||
|
||||
## How it is checked
|
||||
|
||||
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
|
||||
account in the group, says that a new login is needed, and leaves the account's other groups as they
|
||||
were. Unassigning it removes only a membership the module added.
|
||||
Reference in New Issue
Block a user