Author SHA1 Message Date
jschoubben d222091fe9 Issues 245, 246; 240 resolved, 242 located
245: a media server's previews were reached through a link its container never mounted — fixed by
mounting them. 246: settings set replaces the whole layer with nothing to read it first and no
history. 240 is fixed by mesh-controller#47; 242 has backups running and records what the rollout
taught.
2026-10-05 14:54:38 +02:00
mesh-admin 8aeef9969d Merge pull request 'ADR 0189 collects again; issue 244: a verb with an empty schema cannot be called' (#93) from decision/0189-collects-again into main 2026-10-05 12:38:40 +00:00
jschoubben 8ade75c550 ADR 0189 collects again; issue 244: a verb with an empty schema cannot be called
The nightly collector is back on the store (mesh-catalog#57). It was out for
a day because while-stopped named the module-local id and the host refuses a
declaration naming a container it does not have, whole; the namespacing was
fixed the same night and the composition was read before anything was sent
this time — plan's distribution.collect shows while-stopped as
distribution.store. novox applied it with no refusal and the registry was
untouched.

The store is 40G at 14:37, the filesystem 56% full; the collector first fires
at 03:30. Deleting a manifest frees no bytes until then, so that is the
number to read against.

244: mesh-controller.plan and .node publish an empty schema and then refuse
with 'needs node', including when node is passed — the console drops what
the schema does not declare. A tool that cannot be called is worse than one
that is absent, because it is listed. Worked around through the login shell,
which is the path the console exists to replace.
2026-10-05 14:38:17 +02:00
jschoubben 30bd65ad7b Merge pull request 'to-be 43: in progress' (#92) from feat/node-backup into main 2026-10-05 10:24:16 +00:00
mesh-admin b59d479992 Merge pull request 'ADR 0216: the agent's configuration is registered through its module, at three scopes, and served as one plugin' (#89) from decision/the-agent-configured-through-its-module into main 2026-10-05 09:50:44 +00:00
jschoubben 4f599f361b to-be 43: in progress — the seat, the restic holder and the stores' contributions 2026-10-05 11:47:59 +02:00
jochen f291d113c8 ADR 0216: the agent's configuration is registered through its module, at three scopes, and served as one plugin
Graduates research 029 and amends design 36 (section 8): skills, subagents,
commands, hooks and output styles in one nox-mesh plugin; servers, settings
and instructions in the managed files; mesh, node and home scopes.
2026-10-05 11:45:58 +02:00
jschoubben daf2f6d2b1 Merge pull request 'to-be 43: backups against mistakes (graduates research 030)' (#90) from design/43-backups-against-mistakes into main 2026-10-05 09:35:22 +00:00
mesh-admin 4924b26b3c Merge pull request 'ADR 0215: the machine's message bus is a node seat, and it is never restarted live' (#91) from decision/0215-the-message-bus-is-a-node-seat into main 2026-10-05 09:33:36 +00:00
jochen d088f8ec2f ADR 0215: the machine's message bus is a node seat, and it is never restarted live 2026-10-05 11:33:20 +02:00
jschoubben 431375d16c to-be 43: backups against mistakes, declared by modules, kept on the machine
Graduates research 030 into a design for ADR 0214, which merged without one and left the cycle
check failing on main.
2026-10-05 11:31:19 +02:00
jschoubben e51d6f1d86 Merge pull request 'Research 030 + proposed ADR 0214: backups against our own mistakes' (#88) from research/030-backups-against-our-own-mistakes into main 2026-10-05 09:30:35 +00:00
jschoubben a906cd8e67 ADR 0214: accepted by the operator 2026-10-05 11:30:33 +02:00
jschoubben 3d0ce3e6dd Research 030 and proposed ADR 0214: backups guard against mistakes and stay on the machine
The operator scoped backups to mistakes, not disasters (issue 242). Issue 238: fix the failed logins
at their source instead of exempting the operator's address.
2026-10-05 10:07:36 +02:00
mesh-admin bc51d8e7aa Merge pull request 'Issue 195: resolved' (#86) from issue/195-resolved into main 2026-10-05 08:01:20 +00:00
mesh-admin 85c6db0600 Merge pull request 'Research 029: the agent configured through its module, as a plugin the mesh serves' (#85) from research/029-the-agent-configured-through-its-module into main 2026-10-05 08:01:16 +00:00
mesh-admin 9d18968323 Merge pull request 'Issue 243: a rebuilt licence store silences every node' (#84) from issue/243-a-rebuilt-licence-store-silences-every-node into main 2026-10-05 08:01:13 +00:00
jochen 127675e2da Issue 243: a rebuilt licence store silences every node
The manager's generation counter restarted with its store after issue 241,
and nodes that wait for a higher number discarded every binding unseen.
2026-10-05 09:59:17 +02:00
jochen 0e87aad949 Research 029: a running session takes the plugin on reload 2026-10-05 09:53:13 +02:00
jschoubben 2a501af0eb Merge pull request 'Issues 239–242: a name taken over, a dry run recorded, every database dropped, no backups' (#83) from issues/239-242 into main 2026-10-05 00:25:50 +00:00
jschoubben bd7fc40099 Issue 239: name no installation 2026-10-05 02:24:27 +02:00
jschoubben 93d4d29b72 Issues 241, 242: one unreadable grants file dropped every database, and the mesh has no backups
241: the SDK harness read a failed read as "no consumer" and withdrew all seven databases on the
control node; withdrawal destroyed data in seven providers. Fixed in mesh-sdk 0.1.10, mesh-catalog#44
and mesh-host#22; the recovery and what it lost are recorded. 242: nothing backs anything up.
2026-10-05 02:24:11 +02:00
jochen 6ff2440e5d Research 029: the plugin route confirmed on one workstation
Six of seven checks confirmed and autoMode from managed settings very likely;
the agent cannot undo its own managed settings, which a design must allow for.
2026-10-04 23:46:32 +02:00
jochen 65c3315a14 Research 029: the agent's configuration is managed at three scopes
The operator's direction: mesh, node and home. A plugin carries the first
two; the home scope places items the module owns path by path; instructions
follow the same layering.
2026-10-04 17:52:48 +02:00
jochen 854341d4f0 Research 029: the agent configured through its module, as a plugin the mesh serves
Skills, subagents and hooks have no machine-wide place but a plugin, and
hand-copied homes have drifted and gone stale on every machine. Records the
evidence, what the vendor allows, and the options a decision must settle.
2026-10-04 17:46:24 +02:00
jochen d444458bf6 Issue 195: resolved by the controller composing users only for modules that can read an account 2026-10-04 17:37:20 +02:00
mesh-admin dd8b15a0df Merge pull request 'Issue 195: located in the controller's user composition' (#370) from issue/195-diagnosis into main 2026-10-04 15:34:11 +00:00
mesh-admin 02b4ca9bec Merge pull request 'ADR 0213: the operator sets the agent's managed settings through the agent module' (#369) from feat/claude-code-agent-settings into main 2026-10-04 15:34:09 +00:00
jochen 7cd5b37afe Issue 195: located in the controller's user composition
Records why a module that cannot read an account is no bus user, what else
read those users (durable consumers), and answers the report's questions.
2026-10-04 17:33:27 +02:00
jschoubben aac0da9c0c Merge pull request 'ADR 0199: a module that answers names declares its zone, and a node's hosts file is one module's' (#337) from decision/0199-zones-and-the-hosts-file into main 2026-10-04 15:33:07 +00:00
jochen 57b818b669 ADR 0213: the operator sets the agent's managed settings through the agent module
Rules for what the agent may do were set by hand per machine, invisible to
the mesh, and a session cannot loosen its own permissions; design 36 now
takes them as a module setting under the mesh's own keys.
2026-10-04 17:32:18 +02:00
jschoubben 33c10aa85a Issues 239, 240, and 238's diagnosis: a module name taken over, a dry run recorded, the operator banned
239: two repositories defined photos; a rebuild to the catalogue's main replaced the app with a stub
and nothing refused it. 240: a dry-run build was recorded and its definition reached the machine.
238: the forge refused three ssh logins as the operator's account in 31 seconds; nothing in the mesh's
ssh configuration names the forge, and no jail ignores the mesh's own public addresses.
2026-10-04 17:31:29 +02:00
mesh-admin e5042e1e7a Merge pull request 'Issues 237, 238: a self-contradicting assign answer, and the operator's address banned' (#367) from issues/237-238-a-stale-answer-and-a-banned-operator into main 2026-10-04 15:06:51 +00:00
jochen 2580fb6839 Issues 237, 238: a self-contradicting assign answer, and the operator's address banned 2026-10-04 17:06:38 +02:00
mesh-admin 4a8a378ef9 Merge pull request 'ADR 0212: a seat says what it receives, and the machine's hotkeys are a seat' (#366) from decision/0212-contributions-to-a-seat-and-hotkeys into main 2026-10-04 14:42:31 +00:00
jochen 3f48e685cc ADR 0212: a seat says what it receives, and the machine's hotkeys are a seat 2026-10-04 16:42:21 +02:00
mesh-admin 1e0b9ee11f Merge pull request 'As-is: a module's state, and the agent with its licences; designs 36, 39, 40 implemented' (#365) from design/state-and-licences-as-is into main 2026-10-04 14:33:50 +00:00
jochen 1faa63b2d5 As-is: a module's state and the agent with its licences; designs 36, 39 and 40 implemented
What runs since 2026-10-04 and what its first live use showed: refused
requests as timeouts, a late machine reading the whole set, the partial
secrets guard; the agent module and the licence manager on every machine.
2026-10-04 16:32:29 +02:00
mesh-admin 3a359bf99e Merge pull request 'ADR 0211: a machine's power is a node seat, its moments take contributions, and its states are events' (#364) from decision/0211-power-is-a-node-seat into main 2026-10-04 13:58:25 +00:00
jochen ec6d106af8 ADR 0211: a machine's power is a node seat, its moments take contributions, and its states are events 2026-10-04 15:58:16 +02:00
jochen e2b4f3a5c4 Regenerate the decision index 2026-10-04 15:58:09 +02:00
mesh-admin dfb2817fe7 Merge pull request 'ADR 0210: a tool's configuration is its seat holder's, and every other module extends it through the seat' (#361) from decision/0210-a-contribution-depends-on-the-seat-that-receives-it into main 2026-10-04 13:42:56 +00:00
mesh-admin 12742e3a3f Merge pull request 'Designs 36 and 39: the manager's verb is public-key' (#363) from fix/the-verb-is-public-key into main 2026-10-04 13:41:32 +00:00
jochen 5b00bdc7af Designs 36 and 39: the manager's verb is public-key (a seat's verb takes no underscore) 2026-10-04 15:41:23 +02:00
mesh-admin d5a96e5c63 Merge pull request 'Research 028: the mesh's output channel' (#362) from research/028-the-meshs-output-channel into main 2026-10-04 13:36:17 +00:00
jochen 73e6400802 Research 028: the mesh's output channel — how the mesh tells its operator what it noticed 2026-10-04 15:36:03 +02:00
jochen f144ad7be4 To-be 42: who writes what in phase 2 (ADR 0210) 2026-10-04 15:20:01 +02:00
jochen 8394a7bfb1 ADR 0210: a tool's configuration is its seat holder's, and every other module extends it through the seat 2026-10-04 15:19:44 +02:00
mesh-admin 6ac936f170 Merge pull request 'Issues 235, 236: an assignment and a check that let through what then fails' (#360) from issues/235-236-an-assignment-and-a-check-that-let-through-what-fails into main 2026-10-04 13:19:41 +00:00
jochen 3af4179755 Issues 235, 236: an assignment and a check that let through what then fails 2026-10-04 15:17:46 +02:00
mesh-admin dd04729ec5 Merge pull request 'ADR 0209: a login on a node moves that node to its account; an API key is added from any node, sealed' (#359) from decision/0209-a-login-moves-its-node into main 2026-10-04 13:12:01 +00:00
jochen 2e5f40c9fa ADR 0209: a login on a node moves that node to its account; an API key is added from any node, sealed
Traced live: a login to a second account was adopted and left its node
bound to the first, holding a spent refresh token. Designs 36 and 39
amended.
2026-10-04 15:09:56 +02:00
mesh-admin 2f41b4124d Merge pull request 'Issues 233, 234: a stale declaration removed four modules, and the host could not recover' (#358) from issues/233-234-a-stale-declaration-and-a-host-that-cannot-recover into main 2026-10-04 13:09:47 +00:00
jochen 9cc00d57ea Issues 233, 234: a stale declaration removed four modules, and the host could not recover 2026-10-04 14:54:50 +02:00
mesh-admin 438162b5a5 Merge pull request 'Issues 225, 226 and 227 resolved with their live proofs; 232 opened and resolved' (#357) from issues/228-photos-authenticates-against-admin into main 2026-10-04 10:37:53 +00:00
jschoubben 3ed55a3420 connectivity: name the code that builds the one resolver, zones and the hosts file 2026-10-04 00:23:43 +02:00
jschoubben 8184585213 ADR 0199: a module that answers names declares its zone, and a node's hosts file is one module's
The per-node resolvers 0194 retires held two kinds of names that are neither nodes nor routes: the
lab's scenario machines and an operator's own lines. A zone a module declares is forwarded by the
mesh's resolver to that module; /etc/hosts is held per node through node-hosts-file, the operator's
lines in its kept region. Research 023 parks seats that define what their holder owns.
2026-10-03 23:19:36 +02:00
54 changed files with 3192 additions and 20 deletions
@@ -0,0 +1,37 @@
---
status: active
initiated: 2026-10-03
touches: [the seats, the seat protocol, the controller's ownership check, 03-DESIGN/01-to-be/26-the-seats.md]
---
# 023 — A seat protocol that defines what its holder owns
## What is investigated
**A seat is a definition — a protocol — and a module occupies it by implementing that protocol.**
Today the protocol is what the holder accepts, emits and serves (ADR 0118, 0129, 0132): its verbs, as MCP
tool definitions. This asks whether the protocol should also name the **files and directories the
holder owns**, so that occupying the seat means owning them: `node-resolver-config` owns
`/etc/resolv.conf`, `node-hosts-file` owns `/etc/hosts`, the intrusion prevention owns its jail file.
The direction is the protocol's, not the holder's: the seat states what any holder must own; a module
that wants the seat must declare those paths among its resources, or the controller refuses the claim
as not implementing the seat. Two seats may not name one path.
## Why
Who owns a singular file is today answered by reading every manifest, and enforced only after the fact,
when two modules on one machine both declare the same path. The question *which module owns
`/etc/resolv.conf`?* came up on 2026-10-03 with no place to look it up. A seat that names the path answers
it from the seat table, before any module is written, and makes "implements the seat" checkable.
## What it touches
- The seat definition and its table (ADR 0122) — a new part of the protocol.
- The controller's ownership check (`checkResources`), which already refuses two modules owning one path.
- Every node seat that is really about a file: `node-resolver-config`, `node-hosts-file`
([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)),
`node-intrusion-prevention`, `node-packet-filter`.
Raised by the operator during the resolver work of ADRs 0194–0199 and parked there so that work was not
widened by it.
@@ -0,0 +1,71 @@
---
status: active
initiated: 2026-10-04
touches:
- 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md
- 04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md
- 04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md
- 04-ISSUES/233-a-host-without-its-package-managers-configuration-refuses-the-declaration-that-would-restore-it/00-report.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md
- 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
- 03-DESIGN/01-to-be/32-what-a-module-declares.md
became: []
---
# 028 — The mesh's output channel
## What
How the mesh tells its operator what it noticed. The operator's framing: sending notifications
is **an output channel for the mesh**. The mesh already knows when a machine stops answering, when
a failure repeats and will not fix itself, when a rollout waits for ever. Today it keeps that to
itself until someone asks.
The effort looks at:
- **the seat:** one, held once for the mesh, that every other part uses to say something to the
operator;
- **the channels**, each a module: a desktop notification on the machine the operator is at,
**Telegram**, a phone push service, chat, mail and others (see [02](02-the-channels.md));
- **the routing**, by severity and by where the operator is;
- **the life of a message:** deduplicated while it holds, resolved when it stops, acknowledged or
silenced by the operator;
- **the watcher's watcher:** who tells the operator when the parts that would tell them are the
ones that failed.
## Why
[Issue 187](../../04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md) is the
class: *the mesh tells nobody when it stops working*. [01](01-what-the-mesh-already-knows.md)
counts it.
- 15 of the 236 issue reports say the fault was found because a person happened to look.
- 74 describe something failing silently.
On the day this effort opened, the mesh knew three things and told nobody:
- a workstation had refused every declaration for ninety minutes;
- the same workstation had been out of touch for ten minutes after an upgrade;
- one failure on the laptop had repeated thirteen times.
Every one of them was in `status`, for whoever asked.
The pieces exist. [To-be 32](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) already uses a
`telegram-sender` seat as its worked example of a work queue with retention. ADR 0208 made a
machine's desktop notifier a node seat with a `send` verb. What is missing is a seat that speaks
for the mesh, sources that call it, and channels that deliver.
## What it touches
- **The controller**, which would become the first source of what it already computes for `status`.
- **The node-notifier seat**, which would become one channel among several.
- **Issues 187, 229, 230 and 233**, each of which ends in "and nothing said so".
- **ADR 0210**, because a channel extends the output seat through a contribution, and therefore
depends on it.
## Documents
1. [What the mesh already knows](01-what-the-mesh-already-knows.md): the evidence, and the events
that exist.
2. [The channels](02-the-channels.md): the candidates, Telegram first, weighed on the same questions.
3. [Open questions](03-open-questions.md): the seat, routing, life of a message, the watcher's
watcher, what may leave the mesh.
@@ -0,0 +1,60 @@
# 01 — What the mesh already knows, and who hears it
## The count
Over the 236 issue reports in `04-ISSUES/`, on the day this effort opened:
- **15** say, in some wording, that a person found the fault by looking: "nobody was told",
"nothing logged / said / alerted / emitted", "a person asked", "found by a person". Five of them
are still open.
- **74** describe something that failed silently.
The search was a word match over the reports' text, so it undercounts reports that tell the same
story in other words. It never overcounts by much: each of the fifteen was read.
The fifteen fall into three groups:
- **The mesh knew, and kept it in a query.** The fault was in `status`, `plans` or a node's record,
for whoever asked. Examples: a rollout waiting for ever (230), a machine refusing every declaration
(233), a setting that cannot work stored and stopping the node (096).
- **The fault was in a log nothing reads.** Examples: the bus refusing the controller's publishes
(187), a dropped report (187, 230).
- **The fault was invisible to the mesh itself.** Examples: a resolver outside the mesh closed by its
filter (198), a port narrowed without saying (086).
Only the first group is a matter of telling: the fact exists, and only delivery is missing. The
other two need a source first. This effort is about the first, and about giving the other two a
place to say something once they can.
## What the controller computes and does not say
Read from `status` and `node show` on the day this effort opened. Each line is a fact the controller
already holds:
| fact | where it is today | example that day |
|---|---|---|
| a machine is out of touch | `node show`: "last heard from — out of touch 10m" | a workstation after an upgrade |
| a machine refused its declaration | `status`: "refused" with the reason | the same workstation, for 90 minutes |
| a failure repeats and will not fix itself | `status`: "stuck: the same failure N times since …" | 13 times on the laptop |
| machines run different hosts | `status`: the version table | after a host release |
| something runs that the mesh did not write | `node show`: strays | 16 containers on one machine |
| a filter rule the mesh did not write | `status` | one machine |
| a plan is waiting | `plans` | issue 230: "for 0s", for ever |
| an assignment does not compose | the `assign` answer only | issue 235 |
None of these is published. The bus carries a seat's own events (a build's outcome), a module's
declared events, tool calls and declarations. It carries no event for any line above.
## What exists to deliver with
- **A machine's desktop:** the `node-notifier` seat (ADR 0208), held on the laptop. Its `send` verb
shows a notification, and `history` lists them. It was used through the console the day this effort
opened.
- **Mail:** a mail module provides `smtp` to the mesh.
- **Chat:** a Matrix server runs as a module on the home server.
- **Home automation:** a home-automation module runs there too, and its phone app can receive pushes.
- **A seat shape for exactly this:** to-be 32 §5 uses `telegram-sender` (`accepts: send`,
`retain 7d`, `emits: delivered, failed`, `serves: status`) as its worked example. A seat's stream
exists from registration, so work queues until a holder appears.
No module sends to Telegram, a phone push service or SMS today.
@@ -0,0 +1,116 @@
# 02 — The channels
Each channel is a candidate module that delivers what the output seat hands it. They are weighed on
the same questions:
- **Reach:** does it reach the operator away from the machines (phone), or only at a desk?
- **Off-mesh:** does it still work when the mesh's own parts (the bus, the controller, the control
node's network) are what failed?
- **Two-way:** can the operator answer through it: acknowledge, silence, ask?
- **Where the words go:** does the message leave the operator's own machines, and to whom?
- **What it costs to hold:** a secret, a server, an account, money.
## The candidates
### Telegram (required by the operator)
A bot created with Telegram's bot service sends to one chat: the operator's own, or a group.
- **Reach:** the phone and every desktop, with push.
- **Off-mesh:** sending needs only outbound HTTPS from any machine. No inbound port, no server of the
mesh's own. A second machine can hold the same bot token and send when the first is the one that
failed.
- **Two-way:** yes. Inline buttons on a message (acknowledge, silence for an hour) and commands to
the bot, read by long polling over outbound HTTPS. This makes Telegram the strongest candidate for
answering, and the riskiest (see [03](03-open-questions.md), Q7).
- **Where the words go:** to Telegram's servers. Bot chats are not end-to-end encrypted. What a
message may contain is therefore a rule this effort must set.
- **Cost:** one secret (the bot token) and the chat's id. Free. Rate limits are far above what an
operator should receive.
- **Formatting:** short text with a little markup, buttons and links. Enough for a subject, a
machine role, a severity and one line of why.
### The desktop notifier (exists)
The `node-notifier` seat's `send` verb on the machine the operator is at.
- **Reach:** only at that machine, only while a session is up.
- **Off-mesh:** no. It is reached through the mesh's tools.
- **Two-way:** dunst has actions, which a click can answer, but nothing reads them back yet.
- **Where the words go:** nowhere; it is local.
- **Cost:** none.
- **Its place:** the gentlest channel, for a warning while the operator is at a desk. "At a desk" is
itself a question: an unlocked session on a machine with recent input.
### ntfy (or Gotify): a self-hosted phone push
A small server publishes topics; its phone app subscribes.
- **Reach:** the phone, with push.
- **Off-mesh:** only if the server runs outside what failed. On the control node it fails with it.
- **Two-way:** action buttons can call a URL, which is an inbound path to design.
- **Where the words go:** stays on the operator's machines when self-hosted. ntfy's iOS push passes
through an upstream relay unless configured otherwise.
- **Cost:** a module with a container and a routed name; a token per topic.
### Matrix (a server exists as a module)
A bot account posts to a room the operator is in.
- **Reach:** phone and desktop through any Matrix client.
- **Off-mesh:** no, the server is one of the mesh's modules.
- **Two-way:** yes, by messages to the bot.
- **Where the words go:** stays on the operator's server, end-to-end encrypted if the bot supports it.
- **Cost:** a bot account, a secret.
### Mail (a mail module provides `smtp`)
- **Reach:** everywhere, without urgency.
- **Off-mesh:** no, if the mesh's own mail server sends. Yes, through an outside relay.
- **Two-way:** no, not usefully.
- **Its place:** the record and the digest: a daily summary of what was said and resolved, and the
fallback when nothing else acknowledged.
### The home-automation companion app (a module exists)
Its phone app takes pushes and actionable notifications, and the home has lights and speakers.
- **Reach:** the phone, and the house itself: a light that turns a colour.
- **Off-mesh:** no, the home server is a node.
- **Its place:** a playful critical channel, not a primary one.
### The bar on the desktop
An `i3status-rust` block showing the count of open messages, red while one is critical.
- **Reach:** the desk only, and silent.
- **Its place:** the ambient state. Nothing interrupts the operator, and they always see whether
something is open.
### The console (an agent session)
A message the next agent session opens with ("two things happened while you were away").
- **Its place:** context for the agent working on the mesh rather than an alert. It falls out of the
message store if the store is queryable.
### Others, noted and not pursued now
- **SMS or a voice call** through a paid gateway. It is the only channel that works with no data
connection, and the only one that costs per message.
- **Signal**, through an unofficial client: no bot API, and a registered number.
- **Discord or Slack** webhooks: the words go to a third party, as with Telegram, without its two-way
strength.
- **Pushover:** paid, closed, and a phone push service much like ntfy.
- **An external dead-man service** (a heartbeat URL that alerts when pings stop). It belongs to
[03](03-open-questions.md), Q6, as the watcher's watcher rather than as a channel.
## A first reading
- **Telegram** is the primary phone channel, and the only candidate that is cheap, off-mesh capable
and two-way at once.
- **The desktop notifier** is for the desk.
- **The bar** shows the ambient state.
- **Mail** carries the digest and the record.
- **ntfy and Matrix** are self-hosted alternatives for an operator who keeps words off third parties.
The seat must make that a choice, not a rewrite.
@@ -0,0 +1,127 @@
# 03 — Open questions
Each question names the options seen so far. None is decided here.
## Q1. The seat
**What speaks for the mesh to its operator?**
- **a.** One seat in the mesh's own set, held once for the mesh. Working name: `operator-channel`.
- It **accepts** `notify` (a work queue, as to-be 32 §5 designs `telegram-sender`), so a message
waits until a holder appears.
- It **emits** `delivered`, `acknowledged` and `resolved`.
- It **serves** `open` (what is unresolved now) and `history`.
- **b.** No seat: every source calls every channel. Rejected in advance, because each source would
learn every channel. This is the inversion ADR 0126 exists to prevent.
- **c.** Each channel as its own seat, with routing in the sources. Same objection as b, one level up.
Under a, the holder routes. The channels are modules that **contribute** themselves to the seat
(ADR 0210): a channel extends the seat, and so depends on it. Whether the holder is a module of its
own or part of the controller is open. A module keeps the controller small. The controller already
holds most of the facts.
## Q2. What a message is
The first shape seen: a **subject** (what it is about: a machine's role, a module, a plan), a
**kind** (out of touch, refused, stuck, late, …), a **severity**, a one-line **why**, a link to
the tool that shows more, and a **key** that makes it the same message the next time it is said.
- **Severity:** two levels (needs you now / when you can), or three (critical / warning / info)?
Every extra level is a routing rule somebody must keep right.
- **The key** is what makes deduplication possible. "Machine X out of touch" said every minute is
one message, still open, not sixty.
## Q3. The life of a message
open → (acknowledged) → resolved.
- **Deduplicate** by key while open.
- **Resolve** when the source stops saying it, or says it is over ("back in touch after 14 min"). A
channel that can edit its message (Telegram can) updates it in place rather than sending a second.
- **Acknowledge** from any channel that can answer, which stops escalation and repeats.
- **Repeat or escalate** an unacknowledged critical message after a while, to the next channel.
- **Where the open set lives:** the seat's own state, in a key-value bucket (to-be 32's `state:`), so
`open` answers after a restart.
## Q4. Routing and presence
- **By severity:** critical goes to every channel at once. A warning goes to the desk when the
operator is at one, otherwise to the phone, otherwise to the digest.
- **Presence:** "at a desk" needs a fact the mesh does not hold yet. Candidates: an unlocked
graphical session with recent input, read from the `node-lock-screen` and `node-login-manager`
seats' holders. Nothing more invasive.
- **Quiet hours:** a setting of the seat's holder (ADR 0174). Critical overrides it, or not, as the
operator chooses.
- **Rate:** a cap per hour per channel, with the excess folded into one summary, so a storm (a
network outage where every node is out of touch) arrives as one message naming many.
## Q5. The sources
The first sources are the facts in [01](01-what-the-mesh-already-knows.md), all in the controller
today:
- a machine out of touch;
- a declaration refused;
- a stuck failure;
- a plan late (once issue 230 gives a wait an age);
- an assignment that does not compose (issue 235);
- a host version split.
**How each becomes an event:**
- **a.** The controller emits an event per change of state, and the seat's holder consumes them.
- **b.** The controller calls `notify` itself.
With a, the controller learns nothing about telling: other consumers (a board, a log) get the same
facts, and the holder decides what is worth a message. With b, the controller decides severity.
**Modules as sources:** a module may `use` the seat to tell the operator something of its own
(a backup failed, a certificate is close to expiry), with the same message shape.
## Q6. The watcher's watcher
When the controller, the bus or the control node is what failed, nothing above runs. Options:
- **A dead-man signal:** the seat's holder sends a heartbeat out of the mesh (a ping to an external
heartbeat service), which alerts the operator by its own means when pings stop.
- **A second holder of the Telegram channel on another machine** that sends directly, without the
bus, when it stops hearing the controller for longer than a bound.
- **Each host** sending a last message itself when it loses the mesh for longer than a bound. This
needs the channel's secret on every machine, a cost to weigh.
The first is the cheapest and the only one that also covers "the whole house is offline".
## Q7. Answering back
Telegram, and Matrix, can carry the operator's answers.
- **Acknowledge and silence** are safe: they change only the message's state.
- **Commands** ("push the workstation", "show status") turn a chat account into a door to the
controller. If it ever comes, it needs:
- its own record;
- a narrow verb set;
- a check that the answer came from the operator's own account and chat;
- and probably a confirmation step.
The first version should probably answer with acknowledge and silence only.
## Q8. What may leave the mesh
Telegram, and any third-party channel, carries the words to someone else's servers. A message
names a machine, a module and a reason, which is operational detail.
- **What may a message contain?** Roles rather than addresses; no secrets, tokens or paths; a reason
in words. The rule must be enforced by the seat's holder, not hoped for from each source.
- **Is a self-hosted channel required for anything above a severity?**
- **The bot token and chat id** are secrets of the channel's module, delivered as any module secret is.
## Q9. How it is checked
A rule this effort produces must say how it is verified. Candidates:
- a message said twice with one key is one message;
- a resolved source resolves its message;
- a critical message reaches every channel within a bound;
- a message containing an address or a secret is refused;
- the dead-man signal fires when the holder is stopped.
Each is a test of the holder, or a live drill: stop a machine's host and time the message.
@@ -0,0 +1,72 @@
---
status: graduated
initiated: 2026-10-04
touches:
- 03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md
- 03-DESIGN/00-as-is/15-the-agent-and-its-licences.md
- 02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
- 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
became:
- 02-DECISIONS/0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md
- 03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md
---
# 029 — The agent configured through its module
## What
Everything about the operator's coding agent that can be configured is configured through the agent
module's tools, and reaches the machines as **one plugin the mesh serves**:
- subagents, skills, slash commands, hooks, output styles and tool servers;
- the agent's settings and its instructions.
Each item is registered once, from any machine, and goes to one machine, several, or all of them,
including a machine that joins later. The mechanism is the one the module already uses for tool
servers: the registration is kept in the module's state on the bus, and each machine's instance writes
what applies to it. The operator chose the plugin route on the day this effort opened.
**Three scopes** (the operator's direction, the same day):
- **mesh:** in the mesh's plugin and the managed files, on every machine;
- **node:** the same places, rendered for one machine or a list of them;
- **home:** placed in the operator account's own agent directory on a machine, beside what the
person writes there by hand.
Instructions follow the same scopes: the mesh's piece, the node's piece, then further customisation per
machine. See [03](03-options.md).
## Why
The vendor gives the agent a machine-wide directory for its settings, its tool servers and one
instruction file, and **nothing machine-wide for skills, subagents, commands or hooks**. Those exist
only in a home or a project. So today they are copied into each home by hand, and they drift and go
stale. [01](01-what-is-configured-today.md) measures that on four machines.
A plugin is the vendor's own unit for carrying all of those at once. A machine-wide setting can name
a marketplace and enable a plugin from it. If the module serves the plugin and its own managed settings
enable it, the mesh gets a machine-wide place for everything the vendor left home-only, and the home
stays the person's ([ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md)).
## What it touches
- **The agent module's design** ([to-be 36](../../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md)):
its managed directory, its state, and its tools.
- **Module state** ([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)):
where registrations are kept, and which file content fits in a bucket.
- **Contributions** ([ADR 0210](../../02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)):
a module other than the agent's, say the forge's, wanting the agent to have a skill for it. That is a
contribution to the agent's seat, not a file it writes.
- **The managed settings key the operator sets** ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)),
which a settings tool would write rather than a hand-composed settings layer.
## Documents
- [01 — What is configured today](01-what-is-configured-today.md): the evidence.
- [02 — What the vendor allows](02-what-the-vendor-allows.md): plugins, marketplaces and managed
settings, as documented, with sources.
- [03 — Options](03-options.md): where each kind of item goes, how it is registered and stored, and
the questions a decision has to answer.
- [04 — What was confirmed](04-what-was-confirmed.md): the checks, tried on one workstation.
@@ -0,0 +1,51 @@
# 01 — What is configured today
Measured on 2026-10-04 on four machines that run the agent module: two workstations, the control node
and a home server. The figures count what sits in each operator account's agent directory in its home,
outside the module's managed directory.
## What sits in the homes
| what | workstation A | workstation B | control node | home server |
|---|---|---|---|---|
| rule files (`rules/`) | 4 | 2 | 2 | 1 |
| skills of the person's own (`skills/`, beside the vendor's synced ones) | 6 | 2 | 2 | 2 |
| subagents (`agents/`) | 0 | 1 | 0 | 0 |
| slash commands (`commands/`) | 0 | 0 | 0 | 0 |
| plugin marketplaces known | 1 | 2 | 1 | 1 |
## What that shows
- **Six files the design says the operator removes are still on every machine.** To-be 36 §1 lists the
predecessor's rule files and skills and leaves their removal to the operator, "once, on each
workstation". On all four machines, the two predecessor skills are present, byte-identical to each
other:
- one that switches licences through tools that no longer exist;
- one that names the predecessor's forge.
So a session can still load a skill whose every instruction fails.
- **One instruction, three versions.** The predecessor's node-identity rule file is on three machines,
with three different contents. It was written per machine and then left alone.
- **Two instruction sets that contradict each other, loaded together.** On a workstation, one session
reads two sets of instructions:
- the module's managed instruction file says to search the mesh's records first;
- the predecessor's rule files in the home say to search the predecessor's knowledge base first,
through tools that are no longer served.
Both are loaded, and neither says the other is stale.
- **A subagent exists on one machine only.** A reviewer for module definitions was written on one
workstation. The other three machines cannot use it, and nothing says it exists.
- **Settings are per home, and so per machine.** The agent's auto-mode environment, the rules that
decide which actions the agent may take unasked, is written in one home's settings file. It
describes another organisation's cloud, and it answers for this mesh's forge only through a list of
trusted domains. When the agent refused a merge the operator had approved, the only lawful fix was
a managed key ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)).
The agent could not change its own settings, and nothing else in the mesh could either.
## What already works the way this effort wants
Tool servers. A server registered through the module's register tool is kept in the module's state
on the bus, keyed `all.<server>` or `<node>.<server>`. Every instance watches that state and writes what
applies to it into the managed tool-server file, and a machine that joins later takes it at its first
start. Today one server is registered there, for one workstation. That is the shape this effort extends
to everything else.
@@ -0,0 +1,106 @@
# 02 — What the vendor allows
Read from the vendor's documentation on 2026-10-04; the agent installed on the machines measured in
[01](01-what-is-configured-today.md) was a 2.1 release. Each fact names the page it came from. Where the
documentation is silent, this says so. A fact that a design rests on is to be confirmed on one machine
before it is built on (see [03](03-options.md), *What to confirm first*).
## What a plugin can carry
A plugin is a directory with a manifest in `.claude-plugin/plugin.json` and, beside it, any of:
- skills, slash commands and subagents;
- hooks;
- tool servers (`.mcp.json`) and language servers;
- output styles, workflows, themes and monitors;
- a `bin/` directory;
- a `settings.json`.
Its components are namespaced by the plugin's name, so a subagent `reviewer` in a plugin `mesh` is
`mesh:reviewer`, and it never collides with a person's own of the same name.
— *plugins/manifest-reference, plugins/loading (name conflicts)*
**What a plugin cannot carry:**
- **Settings.** Only two keys of a plugin's `settings.json` take effect: the default agent and the
subagent status line. The rest are dropped. — *plugins/manifest-reference, settings*
- **Permission rules.** Not documented as a plugin capability.
- **Instructions.** A `CLAUDE.md` at a plugin's root is not loaded, and the validator warns about it.
Instructions reach a session through skills only. — *plugins/manifest-reference, standard layout*
## Marketplaces, and a marketplace on the machine's own disk
A marketplace is a `marketplace.json` listing plugins and where each comes from. Its sources include:
- a relative path inside the marketplace;
- a forge repository, a git URL or a subdirectory of one;
- a package from a registry;
- an archive over HTTPS;
- the output of a command.
**A marketplace can be a directory on the machine.** Its plugins with relative paths are **loaded in
place**, not copied into the cache. An edit takes effect at the next session start, or at
`/reload-plugins` in a running session, and the plugin's version need not change.
— *plugins/marketplace-reference (marketplace sources), plugins/loading (in-place and copied plugins)*
A plugin from any other source is copied into a cache in the home, under
`plugins/cache/<marketplace>/<plugin>/<version>/`. — *plugins/loading*
## What managed settings do with plugins
These keys work in the machine-wide managed settings file — *plugins/org*:
| key | what it does |
|---|---|
| `extraKnownMarketplaces` | registers a marketplace on every session of the machine |
| `enabledPlugins` | `true` installs and enables a plugin; `false` blocks and hides it at every scope. The managed value outranks every other scope |
| `strictKnownMarketplaces`, `blockedMarketplaces` | allow-list or block-list of marketplace sources |
| `strictPluginOnlyCustomization` | refuses skills, subagents, hooks and tool servers that come from neither a plugin nor managed settings |
| `allowManagedHooksOnly` | runs only the hooks from managed sources |
| `disableSideloadFlags` | blocks loading a plugin from the command line |
| `syncClaudeAiPlugins` | stops plugins synced from the vendor's web account |
**Installed without anyone being asked.** Once the settings reach a machine, the marketplace is
registered and the plugins installed at the next session start. A non-interactive run installs them in
the background. Managed plugins do not wait for the workspace trust prompt. — *plugins/org*
## What the managed settings file honours besides
`permissions` (with its default mode and the switch that disables bypassing it), `autoMode`, `hooks`,
`env`, `model`, `statusLine`, `outputStyle`, `apiKeyHelper`, and the managed-only switches for permission
rules, hooks and tool servers. — *managed-settings*
That `autoMode` is honoured from the managed file is documented. That it changes what the agent
refuses on these machines is still to be seen ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)
left that open).
## Tool servers: the exclusive file wins over a plugin's
When the managed tool-server file is present, as the module writes it, it is **exclusive**: only its
servers load. The vendor's web connectors load too when a managed key allows them. **A plugin's
`.mcp.json` servers are blocked.** — *managed-mcp (exclusive control)*
So a tool server registered through the module stays in the managed tool-server file. Putting it in
the plugin would silently stop it loading.
## Variables inside a plugin
- `${CLAUDE_PLUGIN_ROOT}`: the plugin's directory.
- `${CLAUDE_PLUGIN_DATA}`: a directory that survives updates.
- `${CLAUDE_PROJECT_DIR}`: the project's root.
These resolve in hook commands, tool and language server configuration, and the content of skills,
subagents and commands. A plugin's declared options (`userConfig`) can be marked sensitive; the agent
asks the person for them and stores them itself. — *plugins/manifest-reference (environment variables)*
## Reload
A running session does not see a changed plugin until `/reload-plugins` or a new session.
`/reload-plugins` reloads skills, subagents, hooks and servers. It does not restart monitors.
— *plugins/loading*
## Not documented
- a machine-wide directory for bare skills, subagents or commands. Only a plugin enabled by managed
settings puts them machine-wide;
- permission rules or instructions carried by a plugin.
@@ -0,0 +1,159 @@
# 03 — Options
The route is chosen: a plugin the mesh serves. What is left open is where each kind of item goes, how it
is registered and kept, and what the module does about what it finds in the homes.
## Scopes (the operator's direction, 2026-10-04)
The plugin is not the only place the module manages. **The agent's configuration is managed at three
scopes, and each item is registered at one of them:**
| scope | where it lands | reaches |
|---|---|---|
| **mesh** | the mesh's plugin, and the mesh's part of the managed files | every machine running the agent, including one that joins later |
| **node** | the same plugin and managed files, as rendered on that machine | one machine, or a list of them |
| **home** | the operator account's own agent directory on a machine (`~/.claude`) | that account on that machine |
Each machine renders its own plugin from the registrations that apply to it, so a node-scoped skill sits
in the same `mesh` plugin as a mesh-scoped one, on that machine only. The home scope places an item
where the person's own items live, without the plugin's prefix, as if written there by hand. The
difference is that the mesh knows it placed the item and can change or remove it.
**Instructions follow the same scopes.** The agent reads the managed instruction file first, then the
home's instruction file and its rule files, then the project's. These are concatenated, not overridden:
a later file does not cancel an earlier one, which is how the contradiction measured in
[01](01-what-is-configured-today.md) came about.
- **The mesh's piece** sits in the managed instruction file and is the same on every machine: how a
session on this mesh works, and the conventions.
- **The node's piece** sits in the same file, rendered per machine: its role, and instruction sections
registered for it.
- **Further customisation per machine** sits in the home: a rule file the module places, at the home
scope, beside whatever the person writes there by hand.
**What the home scope needs from [ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md).**
That ADR already lets the mesh own what it places in a home and hold the rest as found. So the module
owns each home item it placed, path by path, recorded in its state. It never writes, renames or
removes an item it did not place. A home item with the same name as one the person made is refused
at registration, never overwritten.
## Where each kind of item goes
[02](02-what-the-vendor-allows.md) puts a hard limit on the plugin: it carries skills, subagents, commands,
hooks, output styles and language servers, but no settings, no permission rules, no instructions, and
no tool server the exclusive managed file does not list. So there are four places, not one:
| kind | goes to | why there |
|---|---|---|
| skills, subagents, slash commands, output styles | **the mesh's plugin** | the only machine-wide place the vendor has for them |
| hooks | **the mesh's plugin**, with the scripts beside them | a hook's script can live in the plugin and be named through `${CLAUDE_PLUGIN_ROOT}`. A hook in the managed settings would need its script placed somewhere else |
| tool servers | **the managed tool-server file**, as today | the exclusive file blocks a plugin's servers |
| settings and permission rules | **the managed settings file**, beside the mesh's keys ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)) | a plugin's settings are dropped |
| instructions | **the managed instruction file**, in sections | a plugin's instruction file is not loaded |
The plugin is reached through two managed keys the module already owns the file for:
`extraKnownMarketplaces`, naming a marketplace directory the module writes, and `enabledPlugins`, set to
`true` for the mesh's plugin. Neither is the operator's to set. Like the attribution key, they are the
mesh's keys and outrank whatever the operator sets.
### Option A — one plugin
Everything the mesh serves is in one plugin, `mesh`, so every invocation reads `mesh:<name>`. That is
simple, and the name says where an item came from.
### Option B — a plugin per source
One plugin for what the operator registers, and one for what other modules contribute (below). An item
then says in its name whether a person or a module definition put it there. But the operator has two
prefixes to remember, and an item has two owners to ask about.
*Leaning:* A. Where an item came from belongs in the module's list tool, not in the item's name.
## Who registers an item
- **The operator, through the module's tools**, from any machine, for one, several or all of them. The
pattern is the tool-server register tool's, extended to every kind:
- `claude_code_<kind>_list`, `_register`, `_unregister` for skills, subagents, commands, hooks, output
styles and instruction sections;
- `claude_code_settings_get` / `_set` and `claude_code_permission_allow` / `_deny` / `_ask` /
`_remove` for the managed settings.
- **Another module, through the agent's seat** ([ADR 0210](../../02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)).
The forge's module wanting the agent to know how pull requests are made here contributes a skill. It
declares the contribution in its definition, and the controller renders it to the agent's holder on
each machine where both run. That depends on the agent module holding a seat; today it holds none.
The two meet in the one plugin. A contribution and a registration with the same name are refused at
registration, and the list tool shows the owner of each.
## Where a registration is kept
The tool-server registrations live in a key-value bucket the module declares
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)),
keyed `all.<name>` or `<node>.<name>`. Skills differ: a skill is a folder, and it can carry scripts
and reference files beside its main file. Every message on the bus is limited to about a megabyte.
1. **One value per item.** The item and its files go in one value, refused above a limit well under
the bus's. It is simple, it fits the module's existing state, and every item measured in
[01](01-what-is-configured-today.md) takes 20 KB or less on disk,
the vendor's synced skills aside. But a skill with a large reference file cannot
be registered at all.
2. **The bus's object store for files, the bucket for the item.** Large files are stored in pieces and
the item names them. Nothing in the mesh uses the object store yet, so ADR 0201 would need
extending.
3. **A repository on the forge.** The plugin is built from a repository, and registering an item is a
commit. This is reviewable and versioned. But a register tool would have to write to the forge,
and the forge would sit on the path to every machine.
*Leaning:* 1 now, with the limit stated and checked at registration. 2 when an item outgrows it. 3
mixes the operator's configuration into the code review cycle, which it does not need.
## Where the plugin is written
The module owns the managed directory, so the marketplace goes under it, written whole by the
module's code:
- the marketplace file;
- one plugin directory beside it.
It is loaded in place, so a change takes effect at the next session, or at `/reload-plugins` in a
running one. Nothing is copied into the home.
## What the module does about what it did not place
[01](01-what-is-configured-today.md) found stale predecessor files on every machine. The module did not
place those, so they are held as found (ADR 0182). It can:
- **report** them: a status tool lists the home's skills, subagents, commands and rule files, says which
the mesh placed, and names those that duplicate a mesh item or call tools no longer served;
- **import** one on request: `claude_code_<kind>_import` takes an item from one machine's home and
registers it at a scope the operator chooses. A skill written by hand on one workstation becomes a
mesh, node or home item in one call. Removing the original stays the person's act.
`strictPluginOnlyCustomization` would make home items stop loading altogether, and home-scoped items
with them. That is the operator's choice to make through the managed settings, not a default of the
module.
## What to confirm first, on one workstation
1. A directory marketplace named in the managed settings, with its plugin enabled there, loads with
no prompt, in place, in an interactive session and in a non-interactive one.
2. The plugin's skills, subagents and commands are offered under `mesh:`, beside the home's own
items, with no collision.
3. A hook in the plugin runs, with its script found through `${CLAUDE_PLUGIN_ROOT}`.
4. The exclusive tool-server file still loads the console, and a server in the plugin does not load,
as documented.
5. `autoMode` in the managed settings changes what the agent refuses (ADR 0213's open point).
6. The account can read the marketplace in the managed directory, which root owns.
7. The managed instruction file and a home rule file the module placed are both loaded, in that order.
## Questions a decision has to answer
- One plugin or one per source (leaning: one).
- How a registration is kept, and the size limit (leaning: one value per item, with a stated limit).
- Whether the managed settings are set through tools writing the module's state, or through the
controller's settings layer as ADR 0213 has it. If both, which one wins on the same key.
- Whether the agent module holds a seat, so that other modules can contribute to it.
- The three scopes, and the home scope's ownership rule: the module owns exactly the home paths it
placed, recorded in its state, and refuses a name the person already uses.
- Whether settings take the same three scopes. The home's settings file is the person's own, so it is
left out unless the operator chooses otherwise.
@@ -0,0 +1,38 @@
# 04 — What was confirmed
On 2026-10-04, on one workstation running the agent's 2.1 release, the checks [03](03-options.md)
listed were tried with a probe. The probe was a directory marketplace holding one plugin named
`mesh`, which carried:
- a skill, a subagent and a slash command;
- a session-start hook running a script in the plugin;
- a tool server in the plugin's own `.mcp.json`.
The managed settings were set through the agent module's `managed_settings`
([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)),
on that machine's layer only, and sent by a push. The module rendered them into the managed file
within seconds, without a restart.
| # | check | result |
|---|---|---|
| 1 | a marketplace named in the managed settings, with its plugin enabled there, loads with no prompt, in place | **confirmed.** A non-interactive session registered the marketplace and enabled the plugin at start, with nothing asked. The plugin was not copied into the home's plugin cache and is not listed among installed plugins: it is read where it lies |
| — | a change to the plugin needs no version bump | **confirmed.** A skill added to the plugin's directory after the first session was offered by the next one |
| — | a running session takes the plugin without being restarted | **seen.** An interactive session that was already running when the plugin was enabled offered its skills and subagent after the operator logged in again in that session, without a restart |
| 2 | the plugin's items are offered under its name, beside the home's | **confirmed.** `mesh:probe-skill`, `mesh:probe-agent` and the command `/mesh:probe`. A collision with a home item of the same name was not tried |
| 3 | a hook in the plugin runs, its script found through `${CLAUDE_PLUGIN_ROOT}` | **confirmed.** The session-start hook ran its script. The vendor's validator asks for the placeholder to be quoted |
| 4 | the exclusive tool-server file still loads the console, and a plugin's server does not | **confirmed.** The session started the console and the registered servers, and never the plugin's server |
| 5 | `autoMode` in the managed settings changes what the agent refuses | **very likely.** With a probe rule forbidding one harmless read-only command, a session that was already running had that command refused moments after the rule was rendered, though the refusal gave no reason. It also suggests the rule reached a running session without a restart. A clean check needs a session whose only difference is the rule; the agent may not start one in its own auto mode, so it is left to the operator |
| 6 | the account can read the managed directory, which root owns | **confirmed** by what already runs: every session reads the managed instruction file from that directory |
| 7 | the managed instruction file and the home's are both loaded, managed first | **confirmed** by what already runs: a session lists the managed instruction file first, then the home's instruction file, then each of the home's rule files |
## What this changes in the options
- The plugin route works as documented, with no file copied into the home. The module's managed
directory can hold the marketplace.
- **A tool server stays in the managed tool-server file.** Check 4 closes that.
- **Undoing a managed setting is not the agent's to do.** When the probe was over, the agent tried to
clear its own machine's settings layer, and its own auto mode refused that as self-modification.
Setting it had been allowed only because it made the agent stricter. So the tools that set the
agent's settings and permissions are tools the operator calls, and the agent calling them for
itself is refused by the vendor's own guard. A design must not assume an agent can tidy up after
itself.
@@ -0,0 +1,29 @@
---
status: graduated
became: [03-DESIGN/01-to-be/43-backups-against-mistakes.md, 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md]
initiated: 2026-10-05
touches: [04-ISSUES/242-the-mesh-has-no-backups, 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md, 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md, 03-DESIGN/01-to-be/31-a-module-declares-its-fail2ban-jail.md]
---
# 030 — Backups against our own mistakes
**What.** How the mesh keeps restore points of its data: what is copied, how often, how long it is
kept, full or incremental, where it lives, who runs it and how it is proven to restore.
**Why.** Issue 242: nothing in the mesh backs anything up, and when one misread file dropped every
database on the control node (issue 241), the newest copies were migration leftovers nine to twelve
days old, found by searching a disk.
**The scope is set by the operator, and it is narrow on purpose: mistakes, not disasters.** A
backup here protects against what a person, an agent or the mesh itself does wrong — a dropped
database, a deleted bucket, a bad migration, a file overwritten — and not against a disk dying or a
building burning. Losing data to a disaster is an accepted risk. That removes the off-site copy, the
cross-site transfer and the second key-holder from the problem, and leaves the part that would have
saved the night of issue 241.
**What it touches.** The store providers (each knows how to dump its own store consistently), the
host's scheduled steps (ADR 0053, built), the node-wide composition pattern a module's jail already
uses (to-be 31), and the operator's output channel (research 028) for saying a backup failed.
Documents: [01 — what the mesh holds](01-what-the-mesh-holds.md),
[02 — options and a proposal](02-options-and-proposal.md).
@@ -0,0 +1,55 @@
# 01 — What the mesh holds, measured 2026-10-05
Four machines: the control node (hosted, holds every public service), the home server (media, home
automation, a large ZFS pool), a workstation and a laptop. Sizes are apparent sizes, rounded.
## Nothing backs anything up
On every machine: no backup tool other than `rsync` and `pg_dump` is installed, no systemd timer and
no cron line mentions a backup, dump or snapshot. Every live data directory is on ext4 except the
home server's pool (ZFS), so a filesystem snapshot is available only there.
## The control node — the data that cannot be recreated
| what | size | how it changes |
|---|---|---|
| object store (file-sync service's files, photos) | 183 GB | slowly; files added, rarely rewritten |
| forge (repositories, attachments, its database) | 7 GB | daily |
| MS SQL Server databases | 5 GB | daily |
| mail (mailboxes; accounts in postgres) | 2 GB | continuously |
| postgres (forge, mail admin, identity, file-sync index, analytics, catalogue, licence manager) | ~2 GB | continuously |
| file-sync service's own directory, website, analytics | ~4 GB | slowly |
| MongoDB | 0.4 GB | daily |
| the mesh's own records (controller, vault, module state) | ~1.5 GB | continuously |
Recreatable and not worth copying: container images (210 GB), the artifact registry (40 GB — every
artifact is rebuilt from git), a 115 GB speed-test bucket and a 7 GB pre-migration object-store copy.
Free space: 833 GB on the filesystem holding the data, 2.9 TB on a second one.
## The home server
MS SQL Server 80 GB, postgres and a self-hosted backend platform ~3 GB, chat server 6 GB, home
automation, network controller and time-series data each under 2 GB, and the media services'
libraries (tens of GB, mostly cover art and metadata they re-fetch). The 89 TB media library is
replaceable by its nature and out of scope. Free: 31 TB on the pool, 453 GB on the system disk.
## The workstation and the laptop
The workstation has 142 GB under its services directory and 31 GB of container volumes; the laptop
3 GB. Mostly development; what among it is data nobody can regenerate is for each module to say.
## Between the sites
Control node to home server ~285 Mbit/s, home server to control node ~19 Mbit/s. Irrelevant now that
backups stay on the machine whose data they hold, recorded because it is why an off-site copy would
have been expensive.
## What issue 241 says about the requirement
- The mistake was noticed within hours. A restore point a day old would have lost a day.
- The restore had to go *beside* the live database, not over it, and that worked well.
- A copy of a live postgres data directory needed a throwaway server of the right version to read;
a logical dump would have restored directly.
- The data that survived was the data outside the dropped stores. A backup that lives inside the
store it protects — a database's own snapshot table, a bucket's own versions — dies with a drop.
@@ -0,0 +1,74 @@
# 02 — Options and a proposal
## Who decides what is backed up
1. **A central list** on the backup holder. Rejected: it is the attentiveness rule ADR 0030
rejected — a store added and not listed is silently unprotected.
2. **Each module declares its own data, a node-wide holder composes them.** The pattern of to-be 31
(a module declares its jail; the mesh composes them per node). A store provider declares *how*
to take a consistent copy (a dump command), a module with plain files declares *which* paths. The
holder composes every declaration on the node into one schedule. **Proposed.**
The data a module keeps in a database it gets from a provider is backed up by the provider, which
dumps every database it serves — so a consumer declares nothing, and a new consumer is covered the
day it is provisioned.
## Full or incremental
- **Databases: a full logical dump every time** (`pg_dump -Fc`, MS SQL `BACKUP DATABASE`,
`mongodump`). A dump restores with the store's own tool into a database beside the live one —
issue 241's recovery without the throwaway server — and is consistent, which a copy of a live data
directory is not.
- **Everything goes into one deduplicating repository per node** (restic or borg). Each night is a
complete restore point, yet only changed chunks cost space: the object store's 183 GB is copied
once, then each night adds what changed. This removes the full-vs-delta trade-off rather than
choosing a side.
Considered for the object store alone: the object store's own versioning with a lifecycle rule.
Rejected as the only copy — it lives inside the store, and a removed bucket or data directory takes
its versions with it.
## Where
On the same machine, outside every data directory the mesh manages, on a second filesystem where the
machine has one (the control node does). Not off-site: the scope is mistakes. The repository is
encrypted anyway (both tools require it); its key is a mesh secret (ADR 0085), so a person can
restore without the holder.
## How often, how long
- **Nightly**, at a quiet hour, as a scheduled step (ADR 0053).
- **On demand before a risky act** — a migration, a retirement, an operator's experiment — through a
verb; the act's own tooling can call it.
- **Kept: 14 daily, 8 weekly, 6 monthly.** A mistake is usually noticed within days, sometimes weeks
(a deleted file nobody opens). Six months bounds the space a slowly-noticed mistake needs.
Estimated cost on the control node: ~200 GB for the first night, a few GB a night after; well within
the second filesystem's 2.9 TB.
## Who runs it
A node seat, **`node-backup`**, held on every machine that has data by one module (named for the tool
it wraps). It receives the declarations, runs them, keeps the repository, and offers the verbs a
person needs:
- what is backed up here, and the last good night of each;
- take a backup now;
- restore one item **beside** the live one — a database to `<name>_restore`, a path to
`<path>.restored-<date>` — never over it. Swapping it in stays a person's act, as in issue 241.
## How it is proven
- Every run checks its own result; a failed or skipped night goes to the operator's output channel
(research 028), not only a log.
- Weekly: the repository's integrity check, and one database restored from the newest dump into a
throwaway instance and counted against the live one.
- A machine with data and no successful backup in 48 hours is a problem the mesh's status shows.
## Open questions
- restic or borg — both fit; restic is a single binary with no server, which suits a module.
- Whether the workstation and laptop take part at all, or only once a module there declares data.
- The mail spool is files and the forge has a dump command of its own; whether the forge's
repositories are worth backing up at all when every clone is a copy (issue 241 says the forge's
*database* is the part with no other copy).
@@ -132,6 +132,34 @@ unreferenced.
asserted is the decision and not the registry's behaviour.
- Live: the store's size before and after the first nightly collection, read from the machine.
## Built, withdrawn, and collecting again — 2026-10-04 and 05
**Both halves are live.** The mesh has been deciding since 2026-10-04: the sweep deletes the
manifests no record names, after each build it recorded, bounded to two hundred artifacts and
sixty seconds so no one waits on it. Its first run let go of two hundred and reported one thousand
one hundred and twenty-six left.
**The nightly collector was withdrawn for a day, and this is why.** `while-stopped` named the
module-local id, `store`, while the machine's container is `distribution.store` — a module names
its own resources locally and a declaration names them under the module, which `restart-on` and
`reload-on` are rewritten for and this field was not. The host refuses a declaration naming a
container it does not have **whole**, so the control machine took nothing at all until the step
came out. The namespacing was fixed the same night (mesh-controller#259) and the step is back
(mesh-catalog#57).
**The lesson is about where a test stands, not about the field.** Both sides passed throughout:
the controller's tests read manifests, the host's read hand-written declarations with bare ids,
and nothing composed one and judged the result against what the host accepts. The test that does
now exists, and it is the one that would have caught this in a second.
**And this time the composition was read before anything was sent** — `plan`'s
`distribution.collect` showing `while-stopped: ["distribution.store"]` — which is the check whose
absence caused the outage.
*Where it stands for the live measurement:* at 2026-10-05 14:37 CEST the store is **40G**, the
machine's filesystem 1.1T used of 2.0T, 56%. The collector first fires at 03:30 the following
morning. Deleting a manifest frees no bytes until it does, so that is the number to read against.
## References
- [issue 108 — the registry has no garbage collection](../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md)
@@ -0,0 +1,127 @@
---
topic: the tiers
status: accepted
date: 2026-10-03
deciders: jochen
reconstructed: false
extends: 0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
---
# 199. A module that answers names declares its zone, and a node's hosts file is one module's
## Context
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) and
[ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) leave
one resolver holding the nodes' internal domains, and retire the resolver every node ran.** Two kinds
of names lived in those per-node resolvers that are neither a node nor a route, and both were found on
the workstation on 2026-10-03:
- **Names a module answers.** The lab raises scenario machines and gives them addresses from its
scenario files — the anchor's stand-in at a documentation address, the home server's on the LAN —
and the workstation resolved `<machine>.incus` through two wildcard lines in a drop-in file its
resolver read. The lines were written by hand; the addresses are the lab's, known only while a
scenario runs.
- **The operator's own names, unrelated to the mesh.** Twelve `<loopback> <name>` lines for a
client's development hosts, kept in `/etc/hosts` and again in `/etc/hosts.local`, which the per-node
resolver read as additional hosts.
**A manifest never names an address, a node or a domain** ([ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md)).
So the lab cannot list `<machine>.incus → <address>` in its definition, and the operator's twelve lines
are not any module's to define.
## Considered Options
**For a module's names:**
1. **The manifest lists its records.** Refused by ADR 0112: the addresses are the lab's runtime facts
and the scenario's choice.
2. **The module reports its records at runtime to the mesh's resolver**, which writes them into its
configuration. It works, and it makes the resolver hold every module's runtime state and decide,
per call, whether the caller may write the name it sent — authorisation for a write, on the one
server every node depends on.
3. **The module declares the zone it answers and the listen that answers it; the mesh's resolver
forwards that zone there.** The definition names a zone (from a setting) and one of its own listens,
which ADR 0112 allows; the address and the port are the mesh's facts. The records stay where they
are known — in the module, at runtime. Chosen.
**For the operator's names:**
1. **Records the controller holds, served by the mesh's resolver.** They are not the mesh's: a client's
development hosts on one machine are nothing any other node should resolve, and the controller would
become the keeper of a workstation's private notes.
2. **A node-scoped module owns `/etc/hosts`, and the operator's lines live in its kept region**
([ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md)), changed
through that module's tools on that machine. Chosen.
## Decision
**1. A module that answers names declares a zone.** Its definition names the zone — a single label or a
dotted name, from a setting, never a domain the mesh knows — and the listen that answers DNS for it.
The controller refuses two modules in the mesh declaring one zone, and a zone that is the mesh's suffix,
under it, or one of a node's public domains: a module may not shadow names the mesh or the public DNS
answers.
**2. The mesh's resolver forwards each zone to the module that declared it.** The controller hands the
holder of `mesh-dns-resolver` every declared zone with the private address of the node its module runs
on and the port that listen is published on; the holder places one forwarding rule per zone into its
configuration and answers nothing in that zone itself. What names exist in the zone, and their
addresses, are the module's — answered by its own long-running code
([ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)),
from its own state, as they change. Whether an answered address is reachable from the asking node is
the module's matter, not the resolver's.
**3. A node's `/etc/hosts` is held by one module, through a node seat, `node-hosts-file`.** The seat is
the definition: its holder owns `/etc/hosts`, and implements three verbs — MCP tool definitions served
as `<node>/node-hosts-file.<verb>` ([ADR 0195](0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)):
**`entries`** (the file's lines, the module's and the operator's, each marked whose), **`add`** (one
address and its names, into the operator's region) and **`remove`** (one name or address from it). The
module writes the machine's own lines — loopback and the machine's name — and keeps a region for the
operator, which survives every push and is given back when the module goes. Its tools change that
region on that machine, escalating as the packet filter's do
([ADR 0175](0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md) §4). **The
controller holds none of it:** an operator's line is the machine's, not a record.
**4. No other module writes `/etc/hosts`.** The private network's region goes, as ADR 0194 already has
it; a module that once wrote a line there asks the mesh's resolver instead.
## Consequences
- **The lab's names follow its scenarios.** A scenario raised is resolvable from every node at once; a
scenario torn down is gone, with no line left behind in any file.
- **The mesh's resolver holds no module's state.** It holds the nodes' domains and a table of who
answers which zone, both composed by the controller; nothing writes to it at runtime.
- **A module answering a zone needs a DNS answerer of its own** — a long-running bundle, or a resolver
it runs. The lab gains one.
- **The operator's names reach the machine's own programs, not its containers.** A container does not
read the machine's `/etc/hosts`. For names unrelated to the mesh that is the right boundary; a name a
container needs belongs in a zone.
- **Taking `/etc/hosts` keeps what is there.** The first time the module writes the file, every line
that is not the machine's own goes into the operator's region, so a workstation's twelve lines survive
the take — the same adoption [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md) gives
every shared file.
**How each is checked:**
- **Zones:** the controller's catalogue tests refuse a second module declaring a zone, a zone under the
mesh suffix, and a zone equal to a node's public domain.
- **Forwarding:** on the holder, the resolver's configuration carries one forwarding rule per declared
zone, at the declaring node's private address and published port; asking any node's resolver for a
name in the lab's zone while a scenario runs returns the scenario's address.
- **The hosts file:** a push leaves the operator's region byte for byte; `add` followed by `entries`
shows the line as the operator's; unassigning the module gives the region back.
## References
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
[ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) —
the one resolver and how nodes ask it.
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) — why a definition names no address.
- [ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md),
[ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md) — kept regions and shared files.
- [ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md) —
where a zone's answerer runs.
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) and [the seats](../03-DESIGN/01-to-be/26-the-seats.md),
amended alongside.
- [Research 023](../01-RESEARCH/023-a-seat-protocol-that-defines-what-its-holder-owns/00-overview.md) —
the general form of decision 3's "the holder owns `/etc/hosts`".
@@ -93,6 +93,11 @@ A node with contributions and no holder writes them nowhere. The holder's absenc
node's assignments, and no contributor is refused for it, because a missing `PATH` entry is a gap, not
a broken machine.
> **The mechanism changed — 2026-10-04, by [ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md).** A contribution is now a dependency on
> node-environment, met and refused as ADR 0207 says, so a node without the holder refuses the
> contributor instead of writing the contribution nowhere. The rest of this section, and the
> decision, stand.
## Consequences
- A shell's part in the environment is one line in its always-read startup file, sourcing the
@@ -87,6 +87,12 @@ reports, `graphical-session`, still gates the display server itself.
The holder of `node-display-server` places them with `${shell:xinitrc:<slot>}` and
`${shell:xresources:<slot>}`.
> **The mechanism changed — 2026-10-04, by [ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md).** A contributor no longer places its own
> file in another tool's directory. It contributes to the tool's seat, and the seat's holder places it,
> in that directory or through a placeholder. Every contribution, the `xinitrc` and `xresources` slots
> included, is a dependency on the seat that receives it. What a contribution contains still follows
> the tool's grain.
**5. The display server's module writes the session's start.** It writes a block at the start of
`~/.xinitrc`, in this order:
@@ -0,0 +1,80 @@
---
topic: what runs on it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md
---
# 209. A login on a node moves that node to the account it logged in to; an API key is added from any node, sealed
## Context
[ADR 0206](0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md)
made a licence an account the manager learns from what the nodes report, adopted by refreshing it, and
bound a node to a licence automatically only when it was bound to nothing (§7); every later change was a
person's act through `bind` and `switch`. It went live on 2026-10-04 with one account, bound to all four
nodes.
**The operator then asked what happens on a login to a second account on one node, and traced, the
answer was wrong.** The manager adopts the second account as a new licence — and leaves the node bound to
the first. The node is left holding the second account's access token, a refresh token the adoption just
spent, and a binding to the first; at the first account's next rotation it is handed a token its own state
file does not name. A login is the most direct thing a person does on a machine about which account it
uses, and the mesh read it as a contribution of a grant only.
**An API key could enter only from a file on the manager's node** (ADR 0206, design 39 §6), so adding one
meant reaching that machine. The operator asked for a streamlined process for both.
## Considered Options
1. **A login on a node switches that node to the account logged in to.** Chosen.
2. **Keep ADR 0206 §7, and have the person `switch` after logging in.** Rejected: the step is easy to
forget and the state between the login and the switch is the broken one described above.
3. **Refuse to adopt a login for an account other than the node's binding.** Rejected: it discards what
the person plainly meant, and a second account could then enter only by a separate act.
## Decision
**1. A login on a node is that node's choice of account.** When the manager adopts a node's login (ADR
0206 §4) — a new account, or a newer login of one it holds — it binds that node to the licence the login
belongs to. If the node was bound to another licence, this is a switch: the node is handed the new
licence's access token and its agent's account is pointed at it, as `switch` does. Every other node stays
where it is. `bind`, `switch` and `release` remain for moving a node without a login.
**2. A login that does not refresh moves nothing.** The candidate is recorded dead (ADR 0206 §4) and the
node keeps its binding; the person logs in again.
**3. An API key is added from any node, sealed, never as an argument.** The agent module serves a tool
that reads a key from a file on its own node, seals it to the manager's public key — which the seat now
serves as a verb — and hands it to the seat's `adopt` on request/reply; the file is removed once the
manager has taken it. Optionally the same call binds that node to the new licence. The seat's `adopt`
still also takes a file on the manager's node. An API key is a licence of its own, never an account's:
nothing is learned about it from a report, and it moves a node only when a person says so.
## Consequences
- Logging in on a node is the whole of moving that node to an account, new or known. The mesh's state
stays consistent: the binding, the token on the node and the account its agent names agree.
- A second account enters the mesh by one login, and only the node it was logged in on uses it.
- **What got harder:** a person who logs in on a node to try an account moves that node; moving it back is
`switch`. Said in the seat's own description of `switch`, and in the agent module's instruction file.
- The key file on a node exists only until the manager has taken it; the key then lives encrypted in the
manager's store alone (ADR 0183).
## How it is checked
| Rule | Checked by |
|---|---|
| A login for another account moves its node and no other | the manager's test: two accounts, the login on one node adopted, that node switched, the others unchanged |
| A newer login of a known account on a node bound elsewhere moves that node | the manager's test |
| A login that does not refresh moves nothing | the manager's test: the binding unchanged, the candidate dead |
| An API key never crosses the bus in the clear and its file is gone afterwards | the agent module's test: the request carries a sealed box only; the file is removed after the seat answered |
| Live | a login to a second account on one workstation: a second licence appears, that workstation is bound to it and its agent names it, the other nodes keep the first |
## References
- [ADR 0206](0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md) — the flow this extends; §7 is changed by decision 1
- [ADR 0183](0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md) — the manager, its seat and its channel
- [to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md), [to-be 39](../03-DESIGN/01-to-be/39-the-anthropic-licence-manager.md)
@@ -0,0 +1,115 @@
---
topic: the mesh
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md
---
# 210. A tool's configuration is its seat holder's, and every other module extends it through the seat
## Context
One fault kept coming back while the machines' modules were rolled out
([to-be 42](../03-DESIGN/01-to-be/42-the-machines-modules-in-order.md)): two modules want the same
thing on a node.
- The bar module and the package manager's module both declared the package that brings the
package manager's helper scripts. The node stopped resolving
([issue 235](../04-ISSUES/235-an-assignment-that-cannot-be-composed-is-recorded-anyway/00-report.md)).
- The launcher, the clipboard manager, the wallpaper and the bar each wrote a file of their own into
the window manager's include directory, as [ADR 0208](0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md)
§4 allowed. Nothing says those modules need the window manager. Assigned without it, they write
configuration nothing reads. Assigned with a different session holder, they write into a directory
that holder does not own.
- [ADR 0203](0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md) §5
lets a module contribute to the environment on a node that has no holder: the contribution is
written nowhere, and nothing says so.
The mesh already has the piece that answers this.
[ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md) made a module depend on
the seat that applies its resources, derived from what it declares. A contribution is the same kind
of need. It is configuration that only the seat's holder can apply.
## Considered Options
1. **Let the first module to declare a file or package own it,** and refuse the second. Rejected:
ownership then depends on the order modules were written, and the second module has no lawful way
to say what it needs.
2. **Allow shared declarations** of one package or file by several modules, merged by the host.
Rejected: a shared file has no owner to answer for it, and removing one module cannot tell what it
alone put there.
3. **Each tool's configuration belongs to the module holding the tool's seat. Every other module
extends it with a contribution to that seat, and a contribution is a dependency on the seat.**
Chosen. It is the operator's statement of the rule.
## Decision
**1. One owner.** A tool's configuration files, and the tool's package, belong to the module that
holds the tool's seat on the node. Only that module writes them.
- The package manager's configuration and helper packages are the package manager's module's.
- The account's environment is node-environment's holder's.
- The window manager's configuration is node-display-session's holder's.
**2. Other modules extend, never write.** A module that needs something in another tool's
configuration declares a **contribution to that tool's seat**:
- its content, in the grain the seat defines (variables and paths, shell code for a named shell and
slot, a window-manager configuration fragment, a notifier rule);
- never a path inside the holder's files or directories.
The holder places what it receives:
- the controller renders the contributions into the holder's files through the holder's placeholders,
as [ADR 0203](0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md) and
[ADR 0204](0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md) already do;
- or the holder writes each contribution to its tool's own drop-in directory. That directory is then
the holder's resource, not the contributor's.
**3. A contribution is a dependency on the seat that receives it.** The controller derives it from the
contribution, as ADR 0207 derives one from a resource. It is met, checked and refused exactly as ADR
0207 §3 and §4 say:
- refused at `assign` when no module on the node holds the seat and the catalogue has a holder;
- refused at composition after the switch.
**4. Needing what another module's package delivers is the same.** A module that needs a program
another seat's holder installs does not declare that package. It depends on the seat, and through
the seat's verbs where they exist. One package is declared by one module on a node.
**5. A seat says what it receives.** A seat lists the contribution kinds its holder accepts. A
contribution of a kind the seat does not list is refused at registration.
## Consequences
- **ADR 0203 §5** no longer holds: a module contributing to the environment depends on
node-environment, and a node without the holder refuses it. The rest of that record stands.
- **ADR 0208 §4**: its first list, contributors placing their own files in another tool's directory,
is replaced by §2 above. The tool's grain stays the guide for what a contribution contains. The
`xinitrc` and `xresources` slots already work this way, and now carry a dependency on
node-display-server.
- **The desktop modules change:** the launcher, the clipboard manager, the wallpaper and the bar
contribute their window-manager lines to node-display-session instead of writing into the include
directory. The bar keeps relying on the package manager's helper scripts through node-package-manager.
- **The two kinds of collision cannot recur:**
- two modules declaring one package or one file;
- a contribution to a seat nobody on the node holds.
A composition that finds either is a fault in a module, not a state a node can be left in.
- **What got harder:** a seat that receives contributions must define their grain, and its holder must
place them. Each new kind is a small change to the controller's renderer or to the holder.
## How it is checked
| Rule | Checked by |
|---|---|
| A contribution derives a dependency on the seat that receives it | the controller's resolve tests |
| A contribution to a seat with no holder on the node is refused at `assign` when the catalogue has a holder | the same tests, and `assign` live |
| One package or file is declared by one module on a node | the controller's composition test, and `module check` across the catalogue |
| A contribution of a kind its seat does not list is refused | the catalogue check, which registration runs |
| No module declares a path inside another module's files or directories | the catalogue check |
## References
- [ADR 0203](0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md),
[ADR 0204](0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md),
[ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md),
[ADR 0208](0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md)
- [Issue 235](../04-ISSUES/235-an-assignment-that-cannot-be-composed-is-recorded-anyway/00-report.md)
@@ -0,0 +1,122 @@
---
topic: what runs on it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
---
# 211. A machine's power is a node seat, its moments take contributions, and its states are events
## Context
The laptop's module needs code to run around sleep:
- the GPU driver's own suspend and resume actions;
- a touchpad reset after waking.
It wrote drop-ins of its own into the service manager's sleep services, so it wrote into files that
belong to another tool's holder. [ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
forbids exactly that. A second module wanting code after waking would do the same, and nothing would
order the two or say that either needs sleep to be handled at all.
The mesh also cannot tell a sleeping machine from a lost one. A laptop with its lid closed stops its
heartbeat exactly as a crashed machine does, and is reported "out of touch" either way.
[Research 028](../01-RESEARCH/028-the-meshs-output-channel/00-overview.md) would turn every closed lid
into an alert.
## Considered Options
1. **Each module writes its own sleep drop-ins,** as the laptop's did. Rejected by ADR 0210: no
owner, no order, no dependency.
2. **The service manager's holder takes power hooks.** Rejected: sleep and power are logind's and the
firmware's concern, not service management's. On a laptop they also include lid, power source and
battery, which the service manager knows nothing about.
3. **A power seat on every machine.** Its holder:
- owns the machine's power handling;
- places code that modules contribute for named moments;
- publishes the machine's power states as events.
Chosen. It was the operator's proposal.
## Decision
**1. `node-power` is a node seat in the mesh's own set.** One module per node holds it. Every
machine has one, servers included: every machine boots and shuts down. The first holder is a module
named `power`.
**2. Its holder owns the machine's power handling:**
- logind's power-key and lid settings;
- the hooks around sleep, boot and shutdown;
- the reading of power source and battery where the machine has them.
A model's specific values, such as what the lid does on that laptop, are the model's module's
contribution or a setting of `power` per node (ADR 0174), never a second writer of logind's
configuration.
**3. Modules contribute code for named moments.** The moments:
- after boot;
- before sleep;
- after waking;
- before shutdown;
- on mains power;
- on battery.
A contribution is POSIX shell code, written with ADR 0204's mechanism as ADR 0208 §4 did for the
session's start:
- a `shell` contribution whose `for` names the moment, in the `first`, `normal` or `last` slot;
- placed by the holder with `${shell:<moment>:<slot>}` in the scripts its own units run;
- run as root, in module order, each piece bounded in time, so that one module's hang cannot hold a
machine awake.
Per ADR 0210, a contribution for a moment depends on `node-power`.
**4. Its states are the holder's events, on the bus:**
- `booted`, `sleeping`, `woke`, `shutting-down`;
- `on-mains`, `on-battery`, `battery-low`, where the machine has a battery.
They carry the machine's role and a time, and nothing secret, so any node and the controller may
consume them.
- **`sleeping` is published before the machine sleeps.** The holder takes logind's delay lock,
publishes, and releases the lock once the bus has acknowledged, within logind's delay bound.
- **On waking,** the holder queues events until the bus is reachable, then publishes them in order.
**5. A machine that said `sleeping` is asleep, not out of touch,** until it says `woke` or misses its
expected return. The controller shows the state, and the output channel (research 028) does not
treat a sleeping machine as a fault.
## Consequences
- The laptop's module moves its sleep drop-ins into contributions:
- the GPU driver's suspend and resume actions before sleep and after waking;
- its touchpad reset after waking.
Its own files in the service manager's directories go.
- Assigning `power` to every machine is phase 1 of [to-be 42](../03-DESIGN/01-to-be/42-the-machines-modules-in-order.md).
- The controller gains the moments as contribution targets placed by `node-power`, and a node's
power state in what it shows about the node.
- **What got harder:**
- code that must run at a precise point inside the sleep transaction cannot be a contribution; the
GPU driver's own units are an example. Such code still declares its own units, and only the
request to run them is contributed;
- an event published around sleep depends on the network still being up. The delay lock buys the
time, and if the bus does not answer within it, the machine sleeps anyway and says so on waking.
## How it is checked
| Rule | Checked by |
|---|---|
| A moment's contribution derives a dependency on `node-power`, and lands in that moment's placeholder in module order | the controller's contribution tests |
| A contribution naming an unknown moment is refused | the catalogue check |
| One piece of hook code that hangs is ended after its bound, and the next still runs | the power module's tests over real child processes |
| `sleeping` is published before sleep and `woke` after, and a missed acknowledgement does not hold the machine awake | the power module's tests with a fake logind and bus, and a live suspend of the laptop |
| A machine that said `sleeping` is not reported out of touch | the controller's status test |
## References
- [ADR 0174](0174-a-node-varies-a-module-through-settings-and-kept-regions-never-an-edit.md),
[ADR 0198](0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md),
[ADR 0204](0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md),
[ADR 0208](0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md),
[ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
- [Research 028](../01-RESEARCH/028-the-meshs-output-channel/00-overview.md)
@@ -0,0 +1,106 @@
---
topic: the mesh
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
---
# 212. A seat says what it receives, and the machine's hotkeys are a seat
## Context
[ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
decided that a tool's configuration belongs to its seat's holder, and that every other module extends
it with a contribution to the seat (§2), in a grain the seat defines (§5). The controller knows only
three such grains:
- the environment ([ADR 0203](0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md));
- shell code in named slots ([ADR 0204](0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md));
- the power moments ([ADR 0211](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)).
Each was a field of its own, with a renderer of its own. Two more appeared on the first workstation:
- **The window manager.** Four modules, the launcher, the clipboard, the wallpaper and the bar, wrote
files of their own into its include directory, and so did the laptop's model module. That is what
ADR 0210 forbids.
- **The keys the window manager never sees.** A laptop's vendor keys reach only a hotkey daemon, which
reads trigger lines (a key, a state, a command). The daemon's configuration was the laptop module's,
although the daemon is a general piece that more than one module has keys for.
A field and a renderer per grain would make every new seat a change to the controller's schema.
## Considered Options
1. **A field per grain,** as before. Rejected: the manifest and the controller grow with every seat
that takes contributions.
2. **Contributions as files in the holder's drop-in directory,** each contributor writing its own.
Rejected by ADR 0210: a path in another module's territory.
3. **One general contribution: a seat, a kind the seat receives, and text in the tool's own grammar.**
The seat lists the kinds it receives. The holder places each kind with one placeholder, and the
controller concatenates the contributions in module order, each under a comment naming its module.
Chosen.
## Decision
**1. A module contributes with `contributions`.** Each entry names:
- a **seat**;
- a **kind**, which that seat receives;
- **content**, text in the tool's own grammar, which the controller does not read.
**2. A seat lists the kinds it receives,** with the comment prefix of its tool's grammar. A
contribution of a kind its seat does not list is refused at registration.
**3. The holder places a kind with `${contribution:<seat>:<kind>}`** in its own files. The placeholder
is filled with every module's contribution of that kind on the node:
- in module order;
- each preceded by a comment line naming the module;
- empty when there is none.
A placeholder in a module that does not claim the seat is refused, as ADR 0204 refuses shell slots.
**4. A contribution depends on its seat** (ADR 0210 §3), derived and refused as
[ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md) says.
**5. Two seats receive first:**
| seat | kind | what it is |
|---|---|---|
| `node-display-session` | `config` | window-manager configuration lines: bindings, start-up commands, rules |
| `node-hotkeys` (new, node scope) | `trigger` | hotkey-daemon trigger lines: a key, a state, a command |
`node-hotkeys` is in the mesh's own set. Its holder runs the daemon that sees the keys the window
manager does not, and owns that daemon's configuration and service.
## Consequences
- The window-manager fragments become `config` contributions of their modules: the launcher, the
clipboard, the wallpaper, the bar and the laptop model. The window manager's module places them
instead of including other modules' files.
- A hotkey module holds `node-hotkeys`. The laptop's model module contributes its vendor keys
instead of writing the daemon's trigger file, and keeps only what is its own: the scripts the keys
run.
- The three earlier grains stay as they are. Folding them into this form is a later change, not
required by this record.
- **What got harder:** a contribution is text the controller does not read, so a malformed line
reaches the tool. The holder checks the composed file with the tool's own check where the tool has
one (the window manager's), before it reloads.
## How it is checked
| Rule | Checked by |
|---|---|
| A contribution names a seat and a kind that seat receives | the catalogue check, which registration runs |
| The placeholder fills with every module's contribution, in module order, each named | the controller's contribution tests |
| A placeholder outside the seat's holder is refused | the catalogue check |
| A contribution derives a dependency on its seat | the controller's resolve tests |
| `node-hotkeys` is a node seat of the mesh's own set | the seat table's tests |
## References
- [ADR 0203](0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md),
[ADR 0204](0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md),
[ADR 0208](0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md),
[ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md),
[ADR 0211](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)
- [To-be 42](../03-DESIGN/01-to-be/42-the-machines-modules-in-order.md)
@@ -0,0 +1,71 @@
---
topic: what runs on it
status: accepted
date: 2026-10-04
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md
---
# 213. The operator sets the agent's managed settings through the agent module, under the mesh's own keys
## Context
The agent module writes the agent's machine-wide managed settings file
([to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md) §2). It carries the mesh's
own keys only: the attribution convention of its repositories, the connectors kept beside the managed
tool servers, and the key-helper for an API-key licence. Every other key was left to the person's own
settings, so that the mesh never reverts a person's choice on a push.
That left no place for a rule the **operator** wants to hold in every session on a machine: what the
agent may do without asking, what it must never do, and what its unattended mode allows. These keys
are not preferences. They are policy about what an agent may do on the mesh. Set by hand in one
person's settings on each machine, they are unmanaged state the mesh cannot see, and the agent refuses
to change them itself, as it should.
## Considered Options
1. **Leave them to each person's settings.** Rejected: policy by hand on each machine, invisible to
the mesh, and a session cannot be asked to loosen its own permissions.
2. **A field per vendor key** (permissions, auto mode, environment, hooks) in the module's settings.
Rejected: the vendor adds keys, and every one would be a change to the module.
3. **One setting holding managed-settings keys, laid under the mesh's own.** The operator sets it
for the mesh or for one node through the controller's settings verb. The module copies its keys into
the managed settings file, then lays the mesh's keys over them.
## Decision
Option 3.
1. The agent module takes a setting, `managed_settings`: an object in the vendor's settings shape. It
is set for the whole mesh or for one node, like the module's other settings, through the
controller's settings verb.
2. The managed settings file is that object with **the mesh's keys laid last**: the attribution
convention, the connectors kept beside the managed servers, and the key-helper. A setting can
neither replace one of these nor add a key-helper that the binding did not ask for.
3. Only the operator sets it, and it is declared state like the role and the extra tool servers. A
person's preferences stay in their own settings; the mesh still sets none of them by itself.
## Consequences
- The operator's rules for the agent are declared once, for the mesh or per node, and reach every
node at the next push. A rule set in the managed settings outranks every other scope, so it holds
in every session on the node.
- A setting layer is replaced whole by the controller's verb. Setting this key without the role or
the extra tool servers clears those in that layer; the module's documentation says so.
- **What got harder:** a person cannot override a rule set here, which is the point. A rule that is
wrong is wrong on every session of the node until the operator changes the setting.
## How it is checked
| Rule | Checked by |
|---|---|
| The operator's keys reach the managed settings file | the agent module's test: an auto-mode allow list and a permissions list set in the setting appear in the rendered file |
| The mesh's keys always win | the same test: a setting naming the attribution, the connectors key or a key-helper is overridden, and a key-helper appears only for an API-key binding |
| Live | the setting given for the mesh; the managed settings file on each node carries the key after the next push |
## References
- [ADR 0183](0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md) — the agent module and the files it writes
- [to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md) — the design this amends (§2, §6)
- the vendor's documentation on managed settings and their precedence
@@ -0,0 +1,66 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md
---
# 214. Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data
## Context
Nothing in the mesh backed anything up (issue 242). When a misread file dropped every database on the
control node (issue 241), recovery took a night and used copies nine to twelve days old.
The operator sets the scope: backups exist for **mistakes** — a person's, an agent's, the mesh's own
— not for disasters. Losing data to a dead disk or a lost site is accepted. Research 030 measured
what the machines hold and found no backup tooling anywhere.
## Considered Options
1. **A central list of what to back up.** Rejected — whatever is not listed is unprotected, silently.
2. **Each store's own mechanisms** (bucket versioning, database snapshots). Rejected as the only copy:
they live inside what they protect, and a drop takes them with it.
3. **Off-site copies.** Out of scope by the operator's decision; recorded so the absence is a choice.
4. **Modules declare, a node seat composes, the copy stays on the machine.** Adopted.
## Decision
**A module declares the data it owns; a node seat, `node-backup`, composes every declaration on the
machine and keeps nightly restore points there.** A store provider declares how to dump each database
it serves, so a consumer of a store declares nothing. A module with files declares their paths.
**Databases are dumped in full, logically, every night; everything lands in one encrypted,
deduplicating repository per machine,** so every night is a complete restore point and only what
changed costs space.
**Kept: 14 daily, 8 weekly, 6 monthly.** A backup is also taken on demand before a risky act.
**The repository is on the machine, outside every directory the mesh manages,** on a second
filesystem where there is one. Its key is a mesh secret.
**A restore goes beside the live data, never over it.** Swapping it in is a person's act.
**A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours
shows in the mesh's status.
## Consequences
- Adding a store provider means declaring its dump; the catalogue check can refuse a store provider
that declares none.
- The control node's first backup is ~200 GB, then a few GB a night.
- A dead disk or a lost machine still loses its data, by choice.
## How it is checked
The catalogue check refuses a module providing a store seat without a backup declaration. The holder's
weekly restore test restores one dump into a throwaway instance and compares counts. The mesh's
status lists every machine whose last good backup is older than 48 hours.
## References
- [04-ISSUES/242](../04-ISSUES/242-the-mesh-has-no-backups/00-report.md), [04-ISSUES/241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
- [01-RESEARCH/030](../01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md)
- ADR 0030 (data outlives its declaration), ADR 0053 (scheduled steps), ADR 0085 (a secret is a provision)
@@ -0,0 +1,74 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md
---
# 215. The machine's message bus is a node seat, and it is never restarted live
## Context
Every machine runs a D-Bus system bus, and the workstations a session bus per login. The service
manager, logind, the network manager, the Bluetooth stack, the GPU switcher, the keyring, the
desktop portal and the power module's sleep lock all speak on it. Nothing in the mesh owned it.
On 2026-10-04 a full upgrade on a workstation restarted the system bus in the middle of the upgrade.
From then on logins hung, sshd answered nothing, and the machine's host stopped reporting, until a
person rebooted it at its keyboard. The mesh had no record of what the bus is, no view of it, and no
rule about when it may restart.
## Considered Options
1. **Leave the bus to the distribution.** Rejected: the outage above is what that gives, and nothing
would ever say the bus is unwell.
2. **Make it part of the service manager's holder.** Rejected: the bus is a separate program with its
own policy, its own clients and its own failure. A machine can have a healthy service manager and
a wedged bus, which is exactly what happened.
3. **A node seat held by a `dbus` module.** Chosen.
## Decision
**1. `node-message-bus` is a node seat in the mesh's own set.** Every machine has one. The first
holder is a module named `dbus`, which owns the bus implementation's package and its system
service, and serves tools to look at both buses.
**2. The bus is never restarted live.** The holder declares the bus running and enabled, and never
restarts or reloads it on any change. A new version of the bus takes effect at the machine's next
boot. A module's change that needs the bus to pick up a policy uses the bus's own reload of policy
files, which keeps every connection, never a restart.
**3. Curated events, never traffic.** The holder publishes on the mesh's bus only what matters about
the machine's bus:
- the bus's health (up, stalled, restarted);
- a well-known system service appearing on the bus or leaving it;
- a policy denial.
The bus's traffic, which carries secrets, notification text and the clipboard, never leaves the
machine. The holder's tools let a person watch it, bounded in time, on request.
**4. The seat receives nothing yet.** Packages ship their own D-Bus policy and service files, and no
module writes one of its own today. When one does, it is a contribution to this seat (ADR 0210, ADR
0212), and the seat lists the kind then.
## Consequences
- Phase 1 of [to-be 42](../03-DESIGN/01-to-be/42-the-machines-modules-in-order.md) gains `dbus` on
every machine.
- A full upgrade that brings a new bus no longer breaks a running machine through the mesh. The
distribution's own upgrade still restarts it, so the rule is enforced only for what the mesh
does. The holder's check says whether the running bus is older than the installed package, which
is the sign that a reboot is due.
- **What got harder:** a fix to the bus itself waits for a reboot. That is the price of never taking
every login on the machine down with it.
## How it is checked
| Rule | Checked by |
|---|---|
| `node-message-bus` is a node seat of the mesh's own set | the seat table's tests |
| The bus's service is declared running and enabled, with no restart or reload trigger | the dbus module's manifest test |
| No traffic is published, only the curated events | the dbus module's tests over its event code |
| A bus older than its installed package is said | the dbus module's check, over a recorded answer |
@@ -0,0 +1,167 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
---
# 216. The agent's configuration is registered through its module, at three scopes, and served as one plugin
## Context
The operator wants everything about the coding agent that can be configured to be configured through
the agent module's tools ([research 029](../01-RESEARCH/029-the-agent-configured-through-its-module/00-overview.md)).
That covers subagents, skills, slash commands, hooks, output styles, settings, permissions,
instructions and tool servers. Each is registered once, from any machine, for one machine, several or
all of them.
The agent module already manages three files in the agent's machine-wide managed directory: the
settings, the tool servers and one instruction file ([to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md)).
Tool servers registered through its tools are kept in its state on the bus and reach every machine
they apply to. The vendor has **no machine-wide place for skills, subagents, commands or hooks**; they
live only in a home or a project. Measured on four machines
([029/01](../01-RESEARCH/029-the-agent-configured-through-its-module/01-what-is-configured-today.md)),
what was copied there by hand had drifted and gone stale:
- two skills of a retired system were on all four machines;
- one rule file existed in three versions;
- two contradicting instruction sets were loaded into the same session;
- a subagent existed on one machine only.
The vendor's **plugin** carries skills, subagents, commands, hooks and output styles. A machine-wide
setting can name a marketplace and enable a plugin from it
([029/02](../01-RESEARCH/029-the-agent-configured-through-its-module/02-what-the-vendor-allows.md)).
Tried on one workstation ([029/04](../01-RESEARCH/029-the-agent-configured-through-its-module/04-what-was-confirmed.md)):
- a plugin in a directory marketplace, enabled by the managed settings, loads in every session with no
prompt, read in place, and an edit reaches the next session with no version change;
- its items are offered under the plugin's name;
- its hooks run;
- its tool servers are blocked by the exclusive managed tool-server file.
A plugin cannot carry settings, permission rules or instructions.
The operator also set the scopes. A skill may be meant for:
- every machine;
- one machine;
- one machine's own account, as if written there by hand.
The instruction file the same: the mesh's piece, the node's piece, and further customisation per machine.
## Considered Options
1. **Copy everything into each home**, owned by the mesh path by path. Rejected as the *only* place:
the home is the person's ([ADR 0182](0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md)),
a mesh item there is indistinguishable from the person's by name, and it fills the directory where
the drift was measured. Kept as one scope of three (below).
2. **A plugin per source:** one for the operator's registrations, one for what other modules
contribute. Rejected: two prefixes to remember for one agent. Where an item came from belongs in
the module's list, not in its name.
3. **Registrations in a repository on the forge**, the plugin built from it. Rejected for now:
registering would be a commit, and the forge would sit on the path to every machine. Configuration
would enter the code review cycle, which it does not need.
4. **One plugin, `nox-mesh`, plus the managed files the module already writes, at three scopes, all
registered through the module's tools and kept in its state on the bus.** Chosen.
## Decision
Option 4.
**1. What goes where.** Each kind of item goes to the one place the vendor honours for it:
| kind | place |
|---|---|
| skills, subagents, slash commands, hooks, output styles | the plugin `nox-mesh` |
| tool servers | the managed tool-server file, as today: the exclusive file blocks a plugin's servers |
| settings and permission rules | the managed settings file: a plugin's settings are dropped |
| instructions | the managed instruction file, in sections: a plugin's instruction file is not loaded |
The plugin is named `nox-mesh` (the operator's choice): the name its items carry in every session, and
not one a person's own plugin is likely to take. It lives in a marketplace directory inside the module's
managed directory, written whole by the module's code. The managed settings name that marketplace and enable the plugin. Those two keys
are the mesh's, laid last like the attribution key (ADR 0213), and no setting replaces them.
**2. Three scopes.** Every registration names one:
- **mesh:** every machine running the agent, including one that joins later.
- **node:** one machine, or a list of them. Rendered into the same plugin and managed files, on those
machines only.
- **home:** the operator account's own agent directory on one machine. The item is placed where the
person's own items live, without the plugin's prefix.
Settings and permission rules take the mesh and node scopes only. The home's settings file stays the
person's.
**3. Instructions follow the scopes.** The managed instruction file holds, in order:
1. the mesh's piece, the same everywhere;
2. the node's piece: its role, and the sections registered for it.
Further customisation per machine is a rule file placed at the home scope. The vendor concatenates
these and does not override, so the module's status tool names any section that contradicts another,
or that calls a tool the mesh no longer serves.
**4. Registered through tools, kept on the bus.** For each kind, the module serves `list`, `register`
and `unregister` tools; for settings, tools that read and set them at a scope. Each registration is a
key in the module's state ([ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)):
- an item and its files are one value, refused above **256 KiB**, well under the bus's message limit;
- every instance watches the state and renders what applies to its machine.
The settings registered this way are laid over the `managed_settings` layer of ADR 0213. In order:
ADR 0213's setting, then the mesh scope, then the node scope, then the mesh's own keys.
**5. The home scope owns only what it placed.** The module records each home path it placed, in its
state. It writes, changes and removes only those. It refuses to register a name the person already
uses there, rather than overwrite it.
**6. What the module did not place, it reports and can import.**
- A status tool lists the home's items and says which the mesh placed. It also names any that
duplicate a mesh item or call tools no longer served.
- An import tool registers an item found in one machine's home at a scope the operator chooses.
- Removing the original stays the person's act.
- Whether home items load at all stays the operator's choice, through a vendor setting in the managed
settings.
**7. Changing the agent's own settings is the operator's act.** The vendor's guard refuses an agent that
loosens its own settings. A settings or permission tool is called on the operator's word, and the
module does not try to get around that refusal.
## Consequences
- One registration puts a skill, a subagent or a rule on every machine, on some, or in one account. A
machine that joins takes the mesh and node items at its first start. Nothing is copied by hand.
- The plugin's items are named `nox-mesh:<name>`, and a person's own items keep their names. Nothing the
mesh adds can shadow them.
- A change reaches the next session on each machine, or a running one at its next plugin reload.
- **What got harder:**
- an item larger than 256 KiB cannot be registered until the state can hold files in pieces;
- the module's state now holds file content, not only small records;
- the stale files already in the homes stay until the person removes them. The module names them;
it does not remove them.
- **Not decided here:** another module contributing a skill or a subagent to the agent through a seat
([ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)).
The agent module holds no seat yet. When one is decided, contributions land in the same plugin.
## How it is checked
| Rule | Checked by |
|---|---|
| each kind lands in its one place | the module's render test: a registered skill, subagent, command, hook and output style appear in the plugin; a tool server in the managed tool-server file; a setting in the managed settings file; an instruction section in the managed instruction file |
| scopes | the same test, for one machine of two: a mesh item on both, a node item on one, a home item only in that machine's home |
| the marketplace keys are the mesh's | the render test: a setting naming either key is overridden |
| the home scope owns only what it placed | the module's test: a name the person already uses is refused; unregistering removes only the placed path |
| the size limit | the module's test: an item above 256 KiB is refused at registration |
| live | a skill registered at the mesh scope is offered as `nox-mesh:<name>` in a new session on each machine |
## References
- [Research 029](../01-RESEARCH/029-the-agent-configured-through-its-module/00-overview.md) — the evidence, the vendor's rules, and what was confirmed
- [ADR 0213](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md) — the managed settings setting this lays over
- [ADR 0182](0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md) — what the mesh may do inside a home
- [ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md) — module state on the bus
- [to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md) — the design this amends
+9
View File
@@ -192,6 +192,8 @@ python3 00-META/checks/index.py fail if stale
- **0190** — [A seat's work is shared by its holders, and building is the first such role](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)
- **0202** — [A provider declares what it derives for each consumer, and the mesh tells both ends](0202-a-provider-declares-what-it-derives-for-each-consumer.md)
- **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
- **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
- **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)
### Its tiers, from the bottom up
@@ -228,6 +230,7 @@ python3 00-META/checks/index.py fail if stale
- **0191** — [The mesh's resolver holds only the mesh's own names; a public name resolves publicly](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)
- **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)
- **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
- **0199** — [A module that answers names declares its zone, and a node's hosts file is one module's](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)
### What runs on them, and how it gets there
@@ -307,6 +310,12 @@ python3 00-META/checks/index.py fail if stale
- **0205** — [Software the distribution does not package ships as a pinned archive of the module's own](0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md)
- **0206** — [A node reports the Anthropic grant it holds; the licence manager adopts a licence by refreshing it, and what each node should hold is the manager's state](0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md)
- **0208** — [The graphical session is one module per piece, on the mesh's seats](0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md)
- **0209** — [A login on a node moves that node to the account it logged in to; an API key is added from any node, sealed](0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md)
- **0211** — [A machine's power is a node seat, its moments take contributions, and its states are events](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)
- **0213** — [The operator sets the agent's managed settings through the agent module, under the mesh's own keys](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
- **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)
- **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)
### How it is built
+51
View File
@@ -0,0 +1,51 @@
---
layer: as-is
status: implemented
code: [mesh-controller internal/catalogue/state.go, mesh-controller internal/broker/state.go, mesh-controller cmd/mesh-controller/push.go, mesh-tools node-tools/internal/bus/state.go, mesh-tools node-tools/internal/launch/launch.go, mesh-sdk src/state, mesh-sdk go/state.go]
updated: 2026-10-04
decisions:
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
---
# A module's state, as it runs
**A module keeps the current value of something on the bus, and every machine sees it — including one
that joins later.** Since 2026-10-04 a manifest may say `state` (buckets the module owns) and `reads`
(another module's, as `<module>.<name>`). The first two modules to use it are the operator's agent on a
machine and its licence manager; on the day this was written, three buckets existed on the bus.
## What runs
- **The controller** creates a key-value bucket `<module>_<name>` for every declared state, from the
catalogue — on every start and, since the first module that declared state found it missing, on every
push before the memberships that name it. A bucket nothing declares any more is reported and kept.
Every bucket carries the mesh's caps: 256 KiB a value, 64 MiB a bucket.
- **The grants**: the machine's runtime is granted, for each bucket a module it carries owns, writing
under the bucket's own subjects and reading; for a bucket it only reads, reading. Measured once built,
with the composed grants loaded into a server: a reader's write is refused by the server.
- **The membership** issued to each assignment lists its buckets by the names the module uses, and
whether it may write.
- **The runtime** answers `mesh/state.get`, `put`, `delete`, `keys` and `watch` on the bundle's channel.
A watch hands the current values — none that is deleted — then every change, each naming the watch it
belongs to; it is answered once the current values are delivered. The runtime refuses, with the
reason, a state the module was not issued, a reader's write, a key the bus cannot hold, and a value
carrying a field named like a credential.
- **The SDKs**: `state(name)` in TypeScript, `stdio.State(name)` in Go (tag `go/v0.1.7` and later).
## What the first live use showed
- **A refused request is a timeout, not a refusal.** The bus reloads a machine's grants a moment after
the push that changed them; a bundle that asks in between waits out its deadline. A module that
watches at start therefore watches beside its handshake and asks again until the state answers — the
agent module needed two to seven attempts on its first start on each machine.
- **A late machine reads the whole set.** A server registered for every machine before one machine was
assigned the module reached that machine from the current values at its start.
- **The secrets guard is partial and works for what it covers**: an entry carrying an `Authorization`
header was refused on the live bus. A sealed value is plain text to an inspector, and is not caught.
## How it is checked
The controller's catalogue and broker tests (names, grants, memberships, a bucket asserted in place
against a real server); the runtime's tests over a real bus (current values without deletions, refusals,
the TypeScript SDK through the runtime); `module check` names a read whose owner on the shelf keeps no
such state.
@@ -0,0 +1,57 @@
---
layer: as-is
status: implemented
code: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
updated: 2026-10-04
decisions:
- 02-DECISIONS/0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md
- 02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md
- 02-DECISIONS/0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md
---
# The operator's agent and its licences, as they run
**Every machine with an operator account runs the agent module, and one licence manager on the control
node keeps the licences.** Both are Go binaries the machine's runtime launches; neither has a container,
a port or a bus credential of its own. Live since 2026-10-04, on all four machines.
## The agent module on each machine
- **Writes the agent's managed directory**: the tool servers — the console as `mesh`, plus servers
registered through the module — the mesh's settings, and the instruction file. The tool-server list
is exclusive by the vendor's rule: a server not in it does not load on that machine.
- **Keeps registered tool servers in its state**, one key per registration for every machine or for one;
each machine renders what applies to it.
- **Reports what its machine holds** — the account its agent names, the kind, fingerprints and expiries,
never a token — at start and whenever the credentials file changes.
- **Writes what the machine should hold**: on a newer generation of its binding it asks the seat's
`current`, sealed to its own key, and writes the access token only. No machine holds a refresh token.
- **Hands over a login when asked**, sealed to the manager's key, and **adds an API key** from a file on
its machine the same way, removing the file once the manager has it.
## The licence manager on the control node
- **Holds the `anthropic-licence-manager` seat**: `licences`, `bindings`, `bind`, `switch`, `release`,
`refresh`, `usage`, `adopt`, `public-key`, `current`.
- **Learns licences from the reports**: a refresh token it does not hold is adopted by refreshing it,
newest login first, once per account. The machine a login was made on is moved to that login's
account; a machine bound to nothing is bound to the account it reports.
- **Is the only refresher**: every exchange under a lease per licence in its own database, every four
hours and in any case within an hour of expiry; grants are encrypted at rest with a key the vault made.
- **Publishes what each machine should hold** as its `bindings` state, with a generation that grows with
every rotation and switch.
## On the day it went live
One subscription account was adopted from the control node's own login on its first start; the other
three machines, logged in to the same account with older logins, were bound to it without their logins
being exchanged. A forced rotation reached all four machines within seconds. Two faults were found and
fixed during the rollout: a machine reporting an already-adopted account later was never bound, and a
seat verb named with an underscore was refused by the builder.
## How it is checked
Each module's own tests (the agent's instruction file held byte for byte to the renderer it replaced; the
manager's rules on a store in memory and a stub vendor; its store against a real database); one run of
both binaries under the real runtime with a stub vendor before going live; and live: `licences` lists the
licence with every machine bound, and each machine's `claude_code_status` names it with no login waiting.
+2
View File
@@ -22,6 +22,8 @@ Where the two disagree, the implementation wins and the disagreement is stated.
| [`11-the-lab.md`](11-the-lab.md) | The lab — the first piece of the new shape that exists, and what it does not yet do |
| [`12-the-seats.md`](12-the-seats.md) | The seats the mesh defines, who holds one, and where a seat changes resolution |
| [`13-the-console.md`](13-the-console.md) | The mesh's tools on the machine a person sits at, served by a module the mesh assigned there |
| [`14-a-modules-state.md`](14-a-modules-state.md) | A module's current state on the bus: what the controller creates, the runtime serves, and the first live use showed |
| [`15-the-agent-and-its-licences.md`](15-the-agent-and-its-licences.md) | The operator's agent on every machine and the licence manager that keeps its licences |
## What these documents are not
+18
View File
@@ -6,10 +6,16 @@ code:
- mesh-controller examples/route-proxy
- mesh-controller internal/identity/authority.go
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
- mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199)
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hosts-file)
- mesh-catalog modules/dnsmasq (the mesh's one resolver)
- mesh-catalog modules/resolv-conf (what a node asks)
- mesh-catalog modules/hosts (a node's /etc/hosts)
- mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-10-03
decisions:
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
@@ -377,6 +383,16 @@ member's resolver answers a LAN; a router pointing at one is moved first. *Check
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
DNS on any address, and by the router's DHCP DNS option naming the router.*
**Names that are neither a node nor a route.** A module that answers names declares a zone (a
setting) and the listen that answers it; the controller hands the `mesh-dns-resolver` holder every
zone with its module's node address and published port, and the holder forwards that zone there and
answers nothing in it itself — the lab answers `<machine>.incus` for its running scenarios this way.
An operator's own names, unrelated to the mesh, live in `/etc/hosts`'s kept region, held per node by
the `node-hosts-file` seat's holder and changed through its tools; the controller holds none of them
([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)).
*Checked by the holder's configuration carrying one forwarding rule per declared zone, and by a push
leaving the hosts file's operator region byte for byte.*
*What follows describes the per-node resolver this replaces — how it was built and why the roles were
split. The split stands; the serving role's scope is what moved.*
@@ -1021,6 +1037,8 @@ The list is worth having in one place, because it is most of the argument:
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
`node-dns-resolver`. The migration's four steps are in the record, in order.
Nor are zones or the hosts file's holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)): the
workstation moves to the one resolver only once both exist, its lab and operator names depending on them.
- ~~**What happens when the hub is down.**~~ **Resolved** by
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
+2
View File
@@ -12,6 +12,7 @@ code:
- mesh-catalog modules/gitea/module.json
updated: 2026-10-03
decisions:
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
- 02-DECISIONS/0161-what-deserves-a-seat.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
@@ -129,6 +130,7 @@ convention, which later seats departed from.
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
| `node-hosts-file` | — | node | — | owns `/etc/hosts`: the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)) |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
@@ -1,9 +1,12 @@
---
layer: to-be
status: designed
code: []
updated: 2026-10-04
status: in-progress
code: [mesh-catalog modules/claude-code]
updated: 2026-10-05
decisions:
- 02-DECISIONS/0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md
- 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
- 02-DECISIONS/0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md
- 02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md
- 02-DECISIONS/0181-the-operator-account-is-a-node-fact-and-a-home-is-a-placement-root.md
- 02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md
@@ -39,7 +42,8 @@ The agent reads a machine-wide, administrator-owned configuration directory unde
by the vendor: a managed settings file that outranks every user and project setting; a key in it that
adds HTTP tool servers *beside* a person's own without blocking them; and a managed instruction file every
session reads before the user's and the project's. The agent has **no** machine-wide directory for
rules, skills, slash commands or hooks; those exist only under a home or a project.
rules, skills, slash commands or hooks; those exist only under a home or a project — or in a **plugin**
the managed settings enable, which is how the mesh puts them on every machine (§8).
So the mesh's part of the agent's configuration lives there, **owned whole by the module**, and the home
is left alone. What the predecessor shipped as two rule files and two skills folds into the managed
@@ -49,7 +53,7 @@ instruction file and the manager's tools:
|---|---|
| `~/.claude/CLAUDE.md` | the managed instruction file: how a session on this mesh works (§3) |
| `~/.claude/rules/00-hal-mesh.md`, `~/.claude/rules/conventions.md` | sections of the same file: this node's identity, the repositories' conventions |
| `~/.claude/settings.json`, merged | the managed settings file: the mesh's keys only, outranking nothing a person did not also set |
| `~/.claude/settings.json`, merged | the managed settings file: the mesh's keys, and the rules the operator set for the agent (§2) |
| `~/.claude/skills/hal-switch-license/SKILL.md` | the manager seat's `switch` verb, listed by the console, and a sentence in the instruction file saying to use it |
| `~/.claude/skills/cleanup/SKILL.md` | nothing; it named the predecessor's forge |
| the console's entry in the agent's user-scope state | the managed settings' tool-server key, from the console's provision (§4) |
@@ -67,7 +71,7 @@ lists them, and until they go the agent reads stale instructions beside the mesh
**Declared, applied by the host:** the agent's package (§7); the module's state directory; a facts file
in that directory carrying the node's name and the console's endpoint, and a settings file carrying the
role and the extra tool servers, merged from the module's settings layers — the bundle is told the two
role, the extra tool servers and the operator's managed-settings keys, merged from the module's settings layers — the bundle is told the two
files' paths, because a bundle's words are paths and constants only (ADR 0192); the bus, the console's provision, and that it uses the `anthropic-licence-manager` seat.
Two directories, declared so the ownership check sees them: the agent's managed directory under
`/etc`, root's, and `~/.claude` under the operator's home, the operator's. No *file* resource under
@@ -78,7 +82,7 @@ changes:
| path | content |
|---|---|
| the managed settings file | the mesh's keys: the tool servers (the console, plus any the operator declared as settings), the attribution trailers, and — for an API-key binding only — the key-helper that serves the key |
| the managed settings file | the keys the operator set in the module's `managed_settings` setting, with the mesh's keys laid over them: the tool servers (the console, plus any the operator declared as settings), the attribution trailers, and — for an API-key binding only — the key-helper that serves the key |
| the managed instruction file | §3 |
| the agent's credentials file under the operator's home | for a subscription binding only: the access token the manager handed over, as the operator, readable by the operator alone, atomic, no refresh token |
| the module's keypair in its state | made once, the private half never leaves (§5) |
@@ -93,6 +97,14 @@ requires. The model, the spinner, the drafts and every other preference are the
predecessor's experience with the model key is the evidence: a mesh that sets a preference reverts a
person's choice on every push.
**The operator's rules for the agent** ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)).
What the agent may do without asking, what it must never do and what its unattended mode allows are
not preferences: they are policy about an agent on the mesh, and a session must not loosen its own. The
operator sets them in the module's `managed_settings` setting, in the vendor's settings shape, for the
mesh or one node. The module copies those keys into the managed settings file and lays the mesh's keys
last, so a setting can never replace the attribution convention, the connectors key or the key-helper,
and a key-helper appears only for an API-key binding. The mesh still sets no preference by itself.
## 3. What the instruction file says
Prose, not a paste; the file is the module's.
@@ -159,6 +171,11 @@ decides it and [ADR 0206](../../02-DECISIONS/0206-a-node-reports-the-anthropic-g
the agent here never refreshes, and a refresh token appearing later is a person's login, reported like
any change; for the API-key licence sets the key-helper in the managed settings to a small program that
prints the key from the module's state, so no file under the home is touched;
- **adds an API key from this node** (*ADR 0209*): `claude_code_add_api_key` reads the key from a file
here, seals it to the manager's `public-key`, hands it to the seat's `adopt`, removes the file once
taken, and on request switches this node to the new licence;
- **follows a login made here**: a login to another account is adopted and moves this node to it (ADR
0209) — nothing for this module to do beyond reporting it;
- **serves `claude_code_status`**: which licence and kind this node holds, when the token expires, whether
the file matches what was handed over — by fingerprint, never by value.
@@ -171,7 +188,8 @@ a token is fetched by request when the state says it changed.
**Every node with an operator account** ([ADR 0181](../../02-DECISIONS/0181-the-operator-account-is-a-node-fact-and-a-home-is-a-placement-root.md)).
All four nodes carry one since 2026-10-03. **Per node:** the role. **Per mesh or per node:**
extra tool servers. **Prerequisite:** the manager holds its seat and has adopted the licences.
extra tool servers, and the operator's managed-settings keys (§2). The controller's verb replaces a
setting layer whole, so a layer set for one of these keeps the others it already held. **Prerequisite:** the manager holds its seat and has adopted the licences.
**Order:** the manager assigned and a refresh observed; the console's provision in the catalogue; this
module on one workstation; the six predecessor files and the hand-made console entry removed there; a
@@ -189,6 +207,60 @@ answer is a package repository for this ecosystem as a seat
and trusted by every node's package manager; not built, and not this module's to build. The vendor's own
installer is rejected: it puts a self-updating binary under the person's home, invisible to the mesh.
## 8. The agent's configuration, registered at three scopes
Everything about the agent that can be configured is registered through this module's tools, once, from
any machine, and kept in the module's state on the bus
([ADR 0216](../../02-DECISIONS/0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)). Every instance watches that state and writes what applies to its
machine. A machine that joins later takes it at its first start.
**What goes where.** The vendor honours each kind of item in one place only, so the module writes four:
- skills, subagents, slash commands, hooks and output styles go into **one plugin named `nox-mesh`**. It
sits in a marketplace directory inside the managed directory, written whole by the module and read in
place by the agent. Its items are offered as `nox-mesh:<name>`, so nothing the mesh adds shadows a
person's own item;
- tool servers go into the managed tool-server file, as in §4. The exclusive file would block a
plugin's servers;
- settings and permission rules go into the managed settings file, as in §2;
- instructions go into the managed instruction file, as sections (§3).
The managed settings name the marketplace and enable the plugin. Those two keys are the mesh's, laid
last with the attribution key, and no setting replaces them.
**Three scopes.** Every registration names one:
- **mesh:** every machine running the agent;
- **node:** one machine or a list of them, rendered into the same plugin and files there only;
- **home:** the operator account's own agent directory on one machine, where the item sits as if
written there by hand.
Settings take the first two scopes only. In the managed settings file they are laid in this order:
the operator's `managed_settings` setting (§2), then the mesh scope, then the node scope, then the
mesh's own keys.
**Instructions follow the scopes.** The managed instruction file holds the mesh's piece, then the node's
piece: its role, and the sections registered for it. Further customisation per machine is a rule file
placed at the home scope. The agent concatenates these and does not override, so the status tool names
a section that contradicts another, or that calls a tool the mesh no longer serves.
**The home scope owns only what it placed.** The module records each home path it placed and touches
only those (ADR 0182). It refuses to register a name the person already uses there.
**The tools.**
- For each kind: list, register and unregister. Each register takes a scope, and a list says where
each item came from.
- For settings and permission rules: read, and set at a scope.
- A **status tool** lists the home's own items beside the mesh's and names the stale ones.
- An **import tool** registers an item found in one machine's home at a scope the operator chooses.
An item and its files are one value in the state, refused above 256 KiB.
**Changing the agent's own settings is the operator's act.** The vendor refuses an agent that loosens its
own settings, and the module does not route around that refusal. A settings tool is called on the
operator's word.
## How it is checked
| Check | Defends |
@@ -198,6 +270,10 @@ installer is rejected: it puts a self-updating binary under the person's home, i
| on a lab machine with no account, the assignment is refused naming the fact | ADR 0181 |
| a switch asked of the seat through the console changes the licence and the token on the node; no tool answer and no log line holds a token | ADR 0183 |
| the API-key binding writes nothing under the home and the agent authenticates through the helper | ADR 0183 |
| the module's test: keys set in `managed_settings` (an auto-mode allow list, a permissions list) appear in the rendered managed settings file, a setting naming the attribution, the connectors key or a key-helper is overridden, and a key-helper appears only for an API-key binding | ADR 0213 |
| the module's render test: a registered skill, subagent, command, hook and output style land in the `nox-mesh` plugin; a tool server, a setting and an instruction section in their managed files; for one machine of two, a mesh item on both, a node item on one, a home item only in that home; a setting naming the marketplace keys is overridden | ADR 0216 |
| the module's test: a home name the person already uses is refused, unregistering removes only the placed path, and an item above 256 KiB is refused | ADR 0216, ADR 0182 |
| live: a skill registered at the mesh scope is offered as `nox-mesh:<name>` in a new session on each machine | ADR 0216 |
| the console's provision resolves by co-location; a machine without the console refuses the module by name | ADR 0027, ADR 0152 |
| a new session on the assigned workstation lists the console's five tools under `mesh` ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)) and answers "which node am I" from the instruction file | the exit of the build |
@@ -1,9 +1,10 @@
---
layer: to-be
status: designed
code: []
status: implemented
code: [mesh-catalog modules/claude-licence-manager]
updated: 2026-10-04
decisions:
- 02-DECISIONS/0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md
- 02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md
- 02-DECISIONS/0183-the-anthropic-licence-manager-is-a-module-and-hands-tokens-to-the-agent-over-the-bus.md
- 02-DECISIONS/0024-model-access-is-a-provision.md
@@ -138,12 +139,16 @@ report, and adopted by refreshing it.
- **The latest login wins.** A bound node is handed an access token only and its file holds no refresh
token, so a refresh token appearing there later is a person's login; its report makes it a candidate,
and if it refreshes it replaces the licence's grant.
- **A first binding follows the login**: a node with no binding whose report names the adopted account
is bound to it. Every later change is `bind`, `switch` or `release`.
- **A login moves its node** (*amended 2026-10-04 by [ADR 0209](../../02-DECISIONS/0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md)*): the node a login was
adopted from is bound to that login's licence — switched, if it was bound to another — and every node
bound to nothing whose report names an account the manager holds is bound to it. `bind`, `switch` and
`release` move a node without a login.
- **The identity guard** files a grant under the identity the node read; where the vendor's refresh
answer names the account too, a mismatch is refused and notified. Which source decided is audited.
- **An API key** is delivered to the manager by the operator through the seat's `adopt` verb from a file
on the manager's node, never as an argument.
- **An API key** enters from any node (*ADR 0209*): the agent module there reads it from a file on its own
node, seals it to the manager's key (the seat's `public-key` verb) and hands it to `adopt`, removing the
file once taken — or `adopt` reads a file on the manager's node. Never an argument, never on a stream.
An API key is a licence of its own and moves a node only through `bind` or `switch`.
## 7. What it emits and serves
@@ -155,7 +160,8 @@ gone; a rotation or a switch is a new generation in the `bindings` state.
**The seat's verbs**, the contract every future holder must serve: `licences` (each with kind,
identity, expiry, failures, who is bound), `bindings`, `bind`, `switch`, `release`, `refresh` (now, one
or all), `usage` (current and history), `adopt`, and `current` (a consumer's token, sealed to the key the
or all), `usage` (current and history), `adopt` (a file on the manager's node, or a key sealed to its
`public-key` — ADR 0209), `public-key`, and `current` (a consumer's token, sealed to the key the
consumer sends — ADR 0206). The manager asks a node for a candidate grant by the agent module's own tool.
## 8. Settings
@@ -1,7 +1,7 @@
---
layer: to-be
status: designed
code: []
status: implemented
code: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
updated: 2026-10-04
decisions:
- 02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md
@@ -4,6 +4,7 @@ status: in-progress
code: [mesh-catalog, mesh-controller, mesh-host]
updated: 2026-10-04
decisions:
- 02-DECISIONS/0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md
- 02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md
- 02-DECISIONS/0040-what-a-module-is.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
@@ -15,6 +16,9 @@ decisions:
- 02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md
- 02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md
- 02-DECISIONS/0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md
- 02-DECISIONS/0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md
- 02-DECISIONS/0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md
- 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
---
# 42. The machines' modules, in order
@@ -76,6 +80,15 @@ than reported.
`kernel` is last because a mistake in it costs a boot. `docker` stays a module without the runtime seat
until ADRs 0165 and 0166 are accepted.
**Added 2026-10-04.** Two more modules for every machine:
- `power` holds `node-power` ([ADR 0211](../../02-DECISIONS/0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)). Other modules contribute code for
its moments (after boot, before sleep, after waking, before shutdown, on mains, on battery), and it
publishes the machine's power states on the bus.
- `dbus` holds `node-message-bus` ([ADR 0215](../../02-DECISIONS/0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)). Modules shipping D-Bus policies or services contribute them to it.
It shares curated events, never raw traffic. An upgrade never restarts the bus live: its package
waits for a reboot. A live restart in the middle of a full upgrade took down a workstation's
logins on the day this was written.
## Phase 2 — both workstations
In order:
@@ -93,6 +106,18 @@ In order:
The seats, gating, contributions and session start are [ADR 0208](../../02-DECISIONS/0208-the-graphical-session-is-one-module-per-piece-on-the-meshs-seats.md).
**Who writes what** is [ADR 0210](../../02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md): a tool's configuration belongs to the
holder of its seat. The launcher, the clipboard manager, the wallpaper and the bar contribute their
window-manager lines to `node-display-session`, and the window manager's module places them. They do
not write into its include directory. Each contribution is a dependency on the seat that receives it,
so assigning one of them without a window manager is refused. The first versions, which still write
the include files themselves, move to contributions once the controller derives the dependency.
**Added 2026-10-04.** `triggerhappy` holds `node-hotkeys` on both workstations ([ADR 0212](../../02-DECISIONS/0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)). The
laptop model's vendor keys become its contribution. The window-manager fragments of the launcher, the
clipboard, the wallpaper, the bar and the laptop model become `config` contributions to
`node-display-session`.
## Phase 3 — one machine model
The laptop's hardware module (vendor daemon, GPU mode, charge limit, logind, brightness and vendor keys)
@@ -0,0 +1,83 @@
---
layer: to-be
status: in-progress
code:
- mesh-controller: internal/catalogue/seats.go (node-backup), internal/catalogue/seat_contributions.go (BackupSeat, CheckBackup, a contribution's directories)
- mesh-catalog: modules/restic (the holder), modules/postgres, modules/mssql, modules/mongodb, modules/minio, modules/influxdb, modules/mesh-vault, modules/mailu, modules/gitea, modules/nextcloud (backup contributions)
updated: 2026-10-05
decisions:
- 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0085-a-secret-is-a-provision.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 43 — Backups against mistakes: a module declares its data, the node keeps restore points
**Every machine with data keeps a restore point of it for every night of the last two weeks, every
week of the last two months and every month of the last half year, on the machine itself.** It is
there for the day a person, an agent or the mesh does something wrong — drops a database, empties a
bucket, runs a bad migration — and not for the day a disk dies (ADR 0214).
## The shape
- **A node seat, `node-backup`,** held on each machine by one module, named for the tool it wraps.
It keeps one encrypted, deduplicating repository on the machine and runs a nightly scheduled step
(ADR 0053). The repository's key is a secret provisioned to the holder (ADR 0085), held in the
vault, so a person can open the repository without the holder running.
- **A module declares its data in its manifest,** naming no node and no absolute path (ADR 0112),
in one of two forms:
- **a dump** — for a store provider: the command that writes a consistent, logical copy of each
database it serves, run by the provider in its own container, its output handed to the holder.
The provider covers every consumer it provisions, so a module that only *uses* a database
declares nothing.
- **paths** — for files a module keeps itself (mailboxes, the forge's attachments, a service's
state directory), named through the module's own directory references.
- **The mesh composes the declarations per node,** as it composes jails (to-be 31) and filters: the
holder receives, as contributions, exactly the data of the modules assigned to its machine. A
module assigned is covered the next night; a module unassigned stops being backed up, and its
restore points age out by the rotation, never at once.
## A night
Each declared dump runs and writes a full logical copy; each declared path is read as it stands. All
of it goes into the repository as one snapshot, tagged by module. The repository keeps only chunks it
has not seen, so the object store's first night costs its full size and later nights cost what
changed. Then the rotation prunes to 14 daily, 8 weekly and 6 monthly snapshots. A dump that fails
fails the night for that module only; the others are still taken.
## The verbs
On the seat, for a person or an agent:
- **what is backed up here** — each module, what it declared, its last good night and its size;
- **take one now** — for one module or all, before a risky act; a migration or a database's retirement
calls it first;
- **restore** — one module's database or path, from a named night, **beside** the live one: a database
as `<name>_restore` owned by the consumer's role, a path as `<path>.restored-<date>`. Swapping it
in stays a person's act. Nothing restores over live data.
## Where the repository lives
On the machine, outside every directory the mesh manages, on a filesystem other than the live data's
where the machine has one. The holder's module names the place as a machine setting, never a path in
a manifest. Not off-site: a lost machine loses its backups with its data, by the operator's choice.
## Proving it
- A night that fails, or does not run, reaches the operator's output channel, naming the module.
- Weekly, the holder checks the repository's integrity and restores the newest dump of one database,
in rotation, into a throwaway instance with no network, comparing table row counts with the live
database.
- The mesh's status lists every machine whose last good night is older than 48 hours.
## Not in this design
- Copies off the machine, and encryption to anyone but the mesh's own vault.
- The media library and anything else a module declares no data for.
- Recreatable things: container images, the artifact registry, caches.
## How it is checked
The catalogue check refuses a module that provides a store seat and declares no dump. The weekly
restore test above, and the 48-hour status line, are the running checks.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-10-02
located-in: []
fixed-by:
located-in: [mesh-controller]
fixed-by: mesh-controller PR #270
amended-design:
---
@@ -0,0 +1,49 @@
# 195 — Diagnosis
## 2026-10-04
**The count had grown, and was still almost all noise.** Every push now opens with 137 users the mesh
"has minted no credential for", across four machines. Checked against the catalogue: six modules declare
an own secret named `broker`; every other module declares none.
**Where the line comes from.** The controller composes the bus's user list from its records: one user
for the controller, one per machine, one per live enrolment token, one per person — and one per module
assigned to a machine, whatever the module declares. Users with no minted credential are left out of the
written file and named in the line. A module with no `broker` secret can never be minted one: issuing
refuses it, because an account nothing reads is an orphan ([issue 078](../078-a-delivered-secret-is-accepted-under-any-name/00-report.md)).
So for those modules the user was composed only to be left out and reported, on every status, plan and
push.
**Why that is safe to stop.** Where the machine's tool runtime runs — every machine, now — the runtime
is the module's way onto the bus ([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md),
[ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)); its grants are the union of
what the modules it carries declare, and that is unchanged. Where no runtime runs, a module without a
`broker` secret cannot connect at all, and a user would not change that.
**What else read the module users.** Each module's durable consumer was derived from its own user. A
module carried by the runtime and declaring no `broker` secret would have lost its consumer, and the
runtime reads that consumer on the module's behalf (ADR 0198). The consumers are now derived from the
module users and from what each runtime carries, one per module and machine.
**Ruled out as still open.** The report's first real gap — a declared `broker` secret filled with a
generated value — was closed by [issue 203](../203-a-fresh-assignment-is-pushed-before-its-credential-exists/00-report.md): a push refuses to make one and names the verb that issues
it. The second — a module that emits with no way onto the bus — has no case on a machine where the
runtime runs, which is every machine now.
**Fix.** A module user is composed only for a module declaring an own secret named `broker`; the
consumers are derived as above. The written accounts file is unchanged, since the users dropped never
had a password. What the line names from now on is the real gap: a module that can read an account and
has not been issued one. Checked by the broker package's tests: no user for a module without an
account, and its consumer still made.
## Answers to the report's questions
- *Should a bus user be composed for a module that declares no `broker` secret?* No.
- *Is a `broker` secret ever correctly made by the generic generator?* No; issue 203 already refuses it.
- *Should a module that speaks on the bus be refused when it declares no `broker` secret?* Not while the
runtime carries it; left for a machine without one, where no case exists today.
## Resolved — 2026-10-04
Live on the control node the same day. The first push after the controller restarted named no users
without a credential, and the number of modules with a durable consumer was unchanged.
@@ -0,0 +1,60 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 233 — A host without its package manager's configuration refuses the declaration that would restore it
## What was observed
2026-10-04, on a workstation. The package manager's module writes that tool's main configuration
file whole, keeping the file it found. A declaration that no longer named the module
([issue 234](../234-a-later-declaration-left-out-four-assigned-modules-and-a-machine-applied-it/00-report.md))
reached a host older than the fix that gives a written-over file back its original. The host
removed the file, and the kept original stayed where the host keeps such files.
From the next declaration on, the host refused every declaration whole:
```
refused a declaration: this is the arch host and pacman does not answer here. Either this machine
is not Arch, or its package database is broken: pacman exited 1: error: config file … could not be
read: No such file or directory
```
The machine stayed in that state for an hour and a half. It retried every five minutes and was
refused every time. Two declarations it refused would have repaired it:
- the current one, which names the package manager's module again and so writes the file;
- the one carrying the newer host, which gives a kept original back.
Neither could be applied. A person restored the kept original by hand, and the next push applied.
`status` showed the machine as `refused` with the error above. That is correct, but nothing said
that the mesh's own tools could no longer reach it.
## Why it matters beyond this instance
The host checks that the package manager answers before it applies anything. That check is right for
a machine that is not what the mesh thinks it is. Here the check depends on a file that the mesh's own
modules own and can remove. Once that file is gone:
- every module's change waits behind it, including the host's own upgrade;
- the only way back is a person on the machine;
- so a fault that one declaration caused, and a later declaration would fix, cannot be undone
through the mesh.
The same shape holds for anything the host probes before applying. If what it probes is one of
the mesh's own resources, one bad declaration can wedge the machine.
## Open questions
1. Should a probe that fails refuse the declaration whole? Or should it fail only the resources that
need the tool, and apply the rest (files, units, the host's own upgrade)? The rest may include
the very resource that restores the tool.
2. Should the host refuse to remove a file a seat's holder needs to answer, or warn before it does?
3. Is a machine that refuses every declaration for longer than one apply a fault the mesh raises by
itself ([issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md))? Today it
appears only to someone who asks for `status`.
@@ -0,0 +1,67 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 234 — A later declaration left out four assigned modules, and a machine applied it
## What was observed
2026-10-04, on a workstation. Four modules (the package manager's, sudo, localization and the
container runtime's) had been assigned and pushed, and the host was applying them in one long apply,
about eleven minutes on a loaded machine. Seven more declarations arrived while it worked. When the
apply ended, the host logged six times `set aside a declaration: a newer one arrived with it` and
applied the one it kept.
The host chooses by the sequence a declaration carries
([issue 107](../107-a-declaration-carries-no-order/00-report.md)), so the one it kept was the latest
the controller had sent. That declaration did not name the four modules. Over the next eight minutes
the host undeclared all of them:
```
removed sudo.operator (…)
removed pacman.config (…)
forgotten pacman.package (pacman)
removed localization.locale (…)
removed docker.prune-service (…)
```
Throughout, the controller's assignments named the four modules on that machine, and they still do:
`plan` for the machine lists them. Nobody had unassigned them.
What happened around it:
- A catalogue asked the controller to catch up twice in the same minutes, and every builder rebuilt
the same catalogue commit.
- The controller daemon restarted three minutes after the stale declaration was applied.
- Another session was working on the mesh and pushing at the same time.
The controller logs no send. Which process sent the stale declaration, and from what view, cannot be
read back from anything the mesh keeps.
## Why it matters beyond this instance
A declaration is the mesh's word on what a machine should be, and the host applies it in full,
including removing what it does not name. A stale one is not harmless: here it removed four modules,
and on a host older than the fix for written-over files it deleted four system files outright
([issue 233](../233-a-host-without-its-package-managers-configuration-refuses-the-declaration-that-would-restore-it/00-report.md)).
[Issue 204](../204-a-controller-handover-re-sent-every-node-a-stale-declaration/00-report.md) was this
family once before and is marked resolved. Ordering by sequence (issue 107) protects against an old
declaration arriving late. It does not protect against a declaration that is new in sequence but
composed from an old view.
## Open questions
1. Where can a declaration be composed with fewer modules than the assignments hold? Candidates:
- a send from a process with an out-of-date view;
- composition while a module's build is being replaced;
- a plan sending what it composed when it was made.
2. Should every send be recorded with its sequence and its sender, so `status` can show what a machine
was last told and by whom?
3. Should a declaration that removes something carry a check the host can verify? For example, the
assignment generation it was composed from, so a host refuses one older than the generation it has
already applied.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 235 — An assignment that cannot be composed is recorded anyway, and the node is one push from leaving the mesh
## What was observed
2026-10-04, assigning thirteen desktop modules to a workstation in one `assign`. The controller
answered `ok: false`. Its output first confirmed every module (`<node> is assigned <module>`, thirteen
times), then repeated the same reason eleven times:
```
<node> is not counted as on the network: it does not resolve: these assignments cannot be applied:
- i3status-rust and pacman both declare the package "pacman-contrib"
<node> is left out of the rest of the mesh: it does not resolve: …
```
and ended with the node's needs from the control node (artifact store, internal CA, package registry)
reported as unmet "because they are not both on the private network".
The thirteen assignments stayed recorded. `plan` for the node failed until one module was unassigned
by hand. Had anything pushed in that window — a person, a plan rolling out a rebuilt module, another
session — the mesh would have composed every node without this one on the private network, and this
node without its own declaration.
## Why it matters beyond this instance
[ADR 0207](../../02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
says `assign` judges the modules together and refuses what cannot be met. It judges seats. A
collision found only at composition — two modules declaring one package, one file, one port — is
reported after the assignment is written, as if it were a warning. The result is the most dangerous
state a node can be in: assigned, unresolvable, and silently dropped by the next push of anything.
The reason was also hard to read: one fault, said eleven times, with three follow-on complaints that
point at the network instead of at the collision.
## Open questions
1. Should `assign` compose the node with the new assignments before writing them, and refuse the
whole call when composition fails? The node would then never become unresolvable through `assign`.
2. When a node does not resolve for any reason, should a push of other nodes keep its last composed
state in theirs rather than drop it from the private network?
3. Should one composition fault be reported once, with the follow-on unmet needs folded under it?
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 236 — The catalogue check passes a manifest the host refuses
## What was observed
2026-10-04. A login-manager module passed `mesh-controller module check`, was registered, built and
assigned. The first push of its declaration was refused whole by the host:
```
refused a declaration: this declaration is refused, and none of it was applied:
- resource "lemurs.service": a service that omits state leaves the unit's lifecycle to the
machine, and boot and takes-over are both its lifecycle
```
The rule is the host's declaration validation. The controller's check never applies it, so a manifest
can pass every check the catalogue has and still fail on the first machine that receives it.
## Why it matters beyond this instance
A refusal is whole, so one module with this fault blocks every other change for that node until the
module is fixed, rebuilt and pushed again. The fault is static: it is in the manifest, and could be
found before a module is merged. Today it is found by assigning the module to a live machine.
## Open questions
1. Should the host's declaration validation be importable, so that `module check` (and registration)
runs it over each resource the manifest declares?
2. Or should the controller validate the composed declaration before it sends it, and refuse to send
one the host would refuse?
3. Which other host-side rules are not visible to the catalogue check today?
@@ -0,0 +1,40 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 237 — `assign` says a seat is held and that the node does not resolve for lack of it, in one answer
## What was observed
2026-10-04. A workstation already ran a module that contributes to `node-hotkeys`, and the module
holding that seat was assigned to it. The answer said, in order:
```
<node> is assigned triggerhappy
and asus-zephyrus-g14 on <node> now has node-hotkeys held
run `push <node>` to send it
<node> is left out of the rest of the mesh: it does not resolve: these assignments cannot be applied:
- asus-zephyrus-g14 on <node> depends on node-hotkeys, which nothing on <node> holds (novox/hq ADR 0207) — assign one that holds it: triggerhappy
```
Both statements cannot be true. `plan` for the node, run straight after, resolved: it held
`node-hotkeys`. A push applied all 300 resources.
## Why it matters beyond this instance
The last lines of an answer are the ones a person and an agent act on. Here they say the node is cut
off from the mesh and name the remedy as the assignment just made. A person would assign it again,
or stop the rollout. An agent following the instruction loops. The tail of `assign` comes from a
second view of the mesh, and that view was not the one the assignment had just changed.
## Open questions
1. Where does the "left out of the rest of the mesh" judgement read the node's assignments from, and
why did it miss the one just recorded: a cached resolution, a read before the write committed, or
the catalogue as registered before the contributing module's new version?
2. Should an answer that contradicts itself be impossible by construction, with every line of it
derived from one resolution taken after the write?
@@ -0,0 +1,40 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 238 — The mesh banned its own operator's address for four weeks
## What was observed
2026-10-04. An agent working for the operator on a workstation polled the forge's ssh port in a loop,
about forty connections in ten minutes, waiting for a branch. The intrusion-prevention holder on the
control node banned the operator's home uplink address in the forge's jail for a day, and then in
`recidive` for four weeks.
From then on, nothing in the operator's home could reach the control node on the banned ports: not
the workstations, not the laptop. The `unban` verb of `node-intrusion-prevention` lifted it, once the
address was found in a ban list of over 400 entries.
## Why it matters beyond this instance
[ADR 0186](../../02-DECISIONS/0186-a-ban-list-never-holds-a-neighbour.md) says a ban list never holds a
neighbour. The operator's home address is the one address the mesh can be sure belongs to it. Every
machine behind it is a node, and the operator reaches the mesh from it. Yet nothing told the jails
so. A ban there locks the mesh out of itself, for longer than any repair takes, and the remedy needs
a path that does not go through the banned address.
The output channel being researched
([research 028](../../01-RESEARCH/028-the-meshs-output-channel/00-overview.md)) would not have said
anything either: a ban is not reported as an event.
## Open questions
1. Which addresses are the mesh's own? The public uplink of every node, as each node reports it, and
the operator's known addresses. Should every jail's ignore list carry them, derived rather than
configured?
2. Should a ban of an address any node reports as its own be refused, or at least emitted as an event
the output channel carries?
@@ -0,0 +1,58 @@
# 238 — Diagnosis
## 2026-10-04, from the control node
**What the forge refused, from the operator's uplink, in the ban's last minute** (the forge's ssh log):
```
14:05:28 Invalid user jochen from <uplink> port 38564
14:05:31 Accepted publickey for git from <uplink> port 45156 (the laptop's key)
14:05:44 Invalid user jochen from <uplink> port 52074
14:05:59 Invalid user jochen from <uplink> port 37978
```
Three refusals in thirty-one seconds — `maxretry = 3` — and the `gitea` jail banned the uplink at
14:06:00; `recidive` counted it the same second. Between the refusals the same key logged in as `git`:
the agent's own git operations were fine, and the refusals were ssh commands that named no user.
**Why they named the wrong user.** On the laptop and the workstation, `ssh -G <forge's public name>`
resolves to the operator's account and port 22: nothing in the ssh configuration the mesh writes
(`ssh-client`, to-be 29) names the forge. A bare `ssh <forge>` — or a git URL without `git@` — presents
the login name, which the forge does not have.
**Why the operator's uplink is bannable at all.** The jails' `ignoreip` (the `fail2ban` module's
`jail.local`) holds loopback, the mesh's private range and every private range (ADR 0186). A node at
home reaches the control node from the home's public address, which is none of those. Nothing the
mesh knows puts it there.
## What would have stopped it
1. **The forge in the ssh configuration the mesh writes**: a `Host` block for the forge's public and
internal names with `User git` and the forge's ssh port. It cannot be written today: the `git`
provision serves the forge's http port only, and the controller translates a served `port` to the
machine's published port but no other key (`ServedOn`), so an `ssh-port` would reach a consumer as
the container's 22, not the machine's 222.
2. **The mesh's own public addresses in every jail's ignore list.** Two sources: each node's public
egress as the hub sees it (the tunnel's peer endpoints — every node at home shows the uplink there),
or an operator setting naming them. The first is derived and stays true when the uplink changes;
the second is a value somebody must remember to edit.
## The operator's call (2026-10-05): fix the logins, not the jail
Remedy 2 is rejected. **A machine of the mesh should not fail logins, and when it does, the ban is
the jail working.** Exempting the operator's address would hide exactly the failures worth seeing —
here a misconfigured ssh client, and in the same week a mail server that had lost every account
(issue 241). Both bans were genuinely failed logins; fail2ban cannot tell why a login failed, and it
should not have to.
What stays:
1. **Remedy 1** — the forge in the ssh configuration the mesh writes, so no machine of the mesh
presents the wrong user. This is the fix.
2. **A broken server must not look like bad clients.** When a service refuses *every* login — the
mail server after its database was emptied — its own health check should fail and say so before
the jail has banned its users. That belongs to the service's module, not to the jail.
## Status
Not located further: remedy 1 needs a served port the controller translates by listen.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 239 — A module name is taken over by another repository, and nothing refuses it
## What was observed
2026-10-04. Five of the photo app's six public names stopped answering on the control node — the API,
two client sites and two aliases — while the sixth answered with a different program. Nothing failed:
the controller composed, the host applied, every check passed.
Two definitions held the module name `photos`:
| | the app's own repository | the catalogue's `modules/photos` |
|---|---|---|
| built from | `photos.git`, branch `nox-mesh`, until 2026-09-28 | from 2026-10-04 04:14 |
| containers | server, admin, two client sites | server, an admin client |
| routes | six | one |
Rebuilding `photos` "to `main`" for [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)
(see [issue 227](../227-the-photo-apps-admin-client-asks-for-the-port-the-proxy-holds/00-report.md)) built the
catalogue's main, not the repository the module had been built from. The controller recorded the new
source, the module moved, the next push replaced the app with the stub, and two sites' containers and
five routes went with it. The stub had sat in the catalogue since 2026-09-03 without ever being the
running definition.
Restored the same day by building from `photos.git` again and removing the stub (the app repository's pull request and
mesh-catalog#278).
## Why it is an issue and not an incident
**A module's source is a fact the mesh records, and any build may overwrite it.** The build history
showed the switch plainly — `photos.git at nox-mesh` on one line, `mesh-catalog.git at main` on the
next — and nothing asked whether a module built from one repository should now come from another.
"The last build wins" is the rule in practice; it is written nowhere, and it lets a stale or unrelated
definition replace a working one silently.
## Open questions
1. Should a build whose source differs from the module's recorded source be refused unless it says so
explicitly (a `--move-source`, or the operator's confirmation)?
2. Should the catalogue's check refuse a module whose name another registered repository already
defines?
@@ -0,0 +1,45 @@
---
status: resolved
opened: 2026-10-04
located-in: [mesh-controller cmd/mesh-controller/main.go, mesh-controller cmd/mesh-builder, mesh-controller internal/link/build.go]
fixed-by: mesh-controller#47
amended-design:
---
# 240 — A dry-run build is recorded, and what it built is applied
## What was observed
2026-10-04, restoring the photo app ([issue 239](../239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md)).
`mesh-controller build <repository> --ref <unmerged branch> --dry-run` was run to prove the branch built
before it was merged. The command's help says *"build and print the manifest, recording nothing"*.
Afterwards:
- `builds photos` listed the dry run as a build — `photos.git at <the unmerged branch>`, with its three
images — beside the real ones.
- On the control node, the module's state directory was rewritten at a time between the dry run and
the merged build: the secret files and environment files the branch's definition declares, which the
definition then running did not.
The content happened to equal what was merged a few minutes later, so nothing broke. Had the branch
been rejected in review, its definition would already have been on the machine.
## Why it matters
A dry run is how a change is proven before a person approves it. If it records the build and the mesh
acts on it, review becomes a formality: the unreviewed definition reaches a machine first.
## Open questions
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it?
## Resolution (2026-10-05)
A dry run is marked on the request (`DryRun`), the builder echoes the mark on its outcome, and the
controller's daemon sets a marked outcome aside: no record, no registration, no plan, nothing a push
could send (mesh-controller#47, with a test that the daemon takes a dry run in with no store at all).
Left open as a follow-up: the catalogue module also hears build outcomes and records their edges; it
should skip a dry run too.
@@ -0,0 +1,89 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply]
fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22
amended-design:
---
# 241 — One unreadable grants file dropped every database on the control node
## What was observed
2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every
connection to the seven databases it manages, dropped them, and seconds later created them again,
empty: the forge, the mail server's admin, the identity provider, the file-sync service, the
analytics service, the catalogue and the licence manager. postgres's own log shows it — connections
"terminating due to administrator command", then clients told the database "seems to have just been
dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42.
Nothing failed loudly. Each application kept running against an empty database: the mail server
refused every login (`relation "user" does not exist`), and those refused logins got the operator's
home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed;
the file-sync service answered 500.
Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's
177 GB of objects.
## Why
**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's
contributions file every five seconds. `readContributions` returned an empty list — without a log
line — when the file could not be read, was not JSON, or was for another requirement. The reconcile
pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was
`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied
its consumers, was enough. The next read was fine, so every consumer came back — to an empty database.
**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged
since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In
the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day
while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and
an agent session on the laptop was inspecting this provisioner's process from 17:52.
**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the
database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the
mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted
the site. Any of them would have lost its data to the same misread.
## Resolution
- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or
unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says
why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made.
A test reproduces the incident and fails on the old harness.
- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and
keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key
and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the
analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a
database out of service is an operator's tool (`postgres_retire_database`), and it renames to
`<name>_deleted_<date>` — nothing in the mesh drops a database.
- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only
while it holds exactly what the host wrote; otherwise it is moved aside to `<path>.removed-<time>`.
Directories with content were already kept.
## Recovery
No database had a backup newer than the migration; issue 242 is the gap. Each was restored from the
newest copy on the control node, by the same path: copy the source, start a throwaway postgres of its
version with no network, dump the one database, restore it beside the live one owned by the consumer's
role, check counts, stop the application, rename the empty live database aside
(`…_deleted_20261005`, kept) and the restored one into place, start, verify as a user.
| application | restored to | notes |
|---|---|---|
| forge | 2026-09-22 | git data complete and current; three repositories created since were adopted |
| mail | 2026-09-25 | all accounts |
| identity provider | 2026-09-26 | both realms |
| file-sync service | 2026-09-25 (a MariaDB dump, converted with the application's own `db:convert-type`) | objects intact; index entries without an object were previews, trash and stock sample files |
| analytics, catalogue | 2026-09-24 | |
| licence manager | none | schema re-created by its preparation step; licences re-adopted from the nodes |
The forge's restore had its own consequences: the forge watcher's admin account and the build agents'
registry accounts were created after the backup and were recreated (the first by the module's own
bootstrap command, the second by its provisioner), and every pull request and issue since 2026-09-22
is gone from the forge's records — the branches remain.
## How it is checked
The SDK test above; the postgres provider's test that withdrawal and retirement issue no `DROP`; the
host's test that a file holding more than the host wrote is moved aside, never deleted.
@@ -0,0 +1,73 @@
---
status: located
opened: 2026-10-05
located-in: [mesh-controller internal/catalogue/seats.go, mesh-catalog modules/restic]
fixed-by: mesh-controller#49, mesh-catalog#49, mesh-catalog#54, mesh-catalog#55, mesh-catalog#56, mesh-media-catalog#1
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
---
# 242 — The mesh has no backups
## What was observed
When seven databases were dropped on 2026-10-04 (issue 241), the newest copy of any of them was a
leftover of the migration: data directories and one dump, nine to twelve days old, found by searching
the control node's disk. Nothing in the mesh takes a backup, no record says what should be backed up,
and nothing would have said so until the day one was needed.
What survived did so by accident: git history because every repository is also cloned on the
operator's machines; the file-sync service's files because they live in the object store, which
nothing dropped; the mailboxes because they are files outside the database.
## Scope, set by the operator (2026-10-05)
**Mistakes, not disasters.** Backups protect against a dropped database, a deleted bucket, a bad
migration — not against a dead disk or a lost site; that loss is accepted. So no off-site copy is
needed, and question 4's "spread across the mesh" and "encrypted to whom" fall away. Worked out in
research 030; proposed as ADR 0214.
## Questions this must answer
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
buckets, the mail spool, the forge's repositories and data, the vault and the controller's own
records, each module's state directories? Is it declared by the module that owns the data, the way
a module declares its listens and its jails?
2. **How often, and how long kept** — a daily schedule? Rotation: how many daily, weekly, monthly?
3. **Full or incremental** — full dumps for databases, deltas (snapshots, deduplicating archives) for
large object stores and file trees?
4. **Where** — never only on the machine whose disk it copies. Spread across the mesh (the home server
holding the control node's, and the other way round), an external target, or both? Encrypted to
whom?
5. **Who runs it** — a seat (`mesh-backup`?) whose holder schedules and stores, with the verbs a person
needs: what was backed up and when, restore one database beside the live one?
6. **How it is proven** — a backup never restored is a hope. A scheduled restore of the newest copy into
a throwaway instance, compared with the live one, and a failure that reaches the operator.
## Why it matters beyond this incident
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.
## Where it stands (2026-10-05)
Backups run (ADR 0214, to-be 43): the control node keeps nightly restore points of its eight stores
and services, the home server of its databases and its media apps' libraries and cover art; each on
its own machine, on its larger filesystem. The first runs were tried one module, then two, then a
whole node, each proven by a restore beside the live data. Not yet built, which is why this stays
open: the weekly test restore into a throwaway instance, the 48-hour status line, and failures
reaching the operator rather than the holder's log.
What the rollout taught, for the next module that takes contributions:
- **A contribution makes the contributing module depend on the seat.** Merging the stores'
contributions before a holder was assigned left every node's plan unresolvable — the whole mesh,
not only the machines running a store — until the holder was assigned. Assign the holder in the
same step as the merge.
- **A brand-new module is not built by the push that adds it**; its first build is asked for by hand.
- **The holder must look as the account that can see.** Its first run called a store's dumps missing
because it checked as the runtime's account, which cannot see inside the store's own directory.
- **A kept single file is not a directory.** restic restores a snapshot's subfolder, not a file;
the first restore of a module keeping its settings file refused it.
- **An update kills a running backup** — the holder's restic is its child. A hand-off to the new
version, the run living on as its own unit under the machine's service manager, is the idea to
take forward.
@@ -0,0 +1,44 @@
---
status: located
opened: 2026-10-05
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
fixed-by:
amended-design:
---
# 243 — A rebuilt licence store silences every node, and nothing says so
## What was observed
The licence manager's store lost its tables in the incident of
[issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md). When
they were made again, the manager adopted the remaining licence from a node's report, as
[ADR 0206](../../02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md)
intends. It bound every node again and rotated the licence on schedule. Its log showed nothing wrong.
On the machines, the morning after:
| machine | generation it last applied | generation the manager gave it | what it did |
|---|---|---|---|
| workstation A | 16 | lower ones, then 17 | ignored three bindings; applied the fourth |
| home server | 11 | 10 | ignored it, kept a token issued before the loss |
| workstation B | 12 | 9 | ignored it, kept a token issued before the loss |
| control node | 12 | 12 | in step, by chance |
- **A login waited three minutes.** The operator logged in to a second account on workstation A. The
manager adopted the login and moved the machine to it at once (ADR 0209). The machine went on
reporting its login as waiting. The manager adopted the same login again three more times, a vendor
refresh each time, about a minute apart. Only when the count passed 16 did the machine take it.
- **Two machines would have lost the agent when their tokens expired.** Neither had taken the rotation
made since the loss. Each would have stopped working, with nothing anywhere saying why.
## Why it matters
The manager counted generations from one again, because its counter is a database sequence and was
made again with the tables. A machine applies a binding only when its generation is greater than the
one it applied last. So a store that is rebuilt, the very event a mesh must survive, silences every
machine that applied a higher number. And the manager cannot see it: it publishes, the publish
succeeds, and the machine discards the binding without a word.
The generation guards against something that cannot happen. The bindings state keeps only the latest
value per machine, so an older binding cannot arrive after a newer one.
@@ -0,0 +1,33 @@
# 243 — Diagnosis
## 2026-10-05
**Trail.**
1. The operator reported that a login on workstation A was "not going well".
2. The machine's log showed it reporting a login as waiting, four times, while it still held the old
licence's binding at generation 16.
3. The manager's log showed it adopting that login four times, and the store's tables missing from
02:09 until it adopted the remaining licence again at 02:22.
4. The bindings the manager listed carried lower generations than the ones each machine reported:
10 against 11 on one machine, 9 against 12 on another.
5. A refresh of the remaining licence, the manager's own verb, gave every binding a generation
above the old numbers. All three of its machines applied it within the second.
**Located in two places.**
- **The agent module, on each machine.** It applies a binding only when its generation is greater than
the one applied, a rule that assumes the counter never goes back. It now applies any binding that
differs from the one applied, in licence or generation. The state holds only the latest value per
machine, so nothing older can arrive. Checked by the module's tests: a lower generation after a
rebuild is applied, and an equal one is still skipped.
- **The licence manager.** It numbers bindings from a database sequence, which starts again with a
rebuilt store. Before it considers the machines' reports, it now moves the sequence past every
generation a machine reports having applied, and never moves it back. Without this, a binding could
be given exactly the number a machine already applied, and that one would still be skipped. Checked
by the manager's tests, and the statement was tried on a database: a fresh sequence asked to pass
16 gives 17 next, and asking it to pass 5 afterwards leaves it where it was.
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
same second it was published. The licences themselves were sound. The remaining licence refreshed
on every attempt, and the second account's login was adopted from its first report.
@@ -0,0 +1,66 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-tools, mesh-controller]
fixed-by:
amended-design:
---
# 244 — A verb whose published schema is empty cannot be called through the console, and says only that it needs the argument it will not take
## What was observed
Through the console, 2026-10-05, composing one machine's declaration before pushing it:
```
mesh_describe mesh-controller.plan
→ {"properties": {}, "type": "object"} ... it takes nothing
mesh_call mesh-controller.plan {}
→ plan failed: plan needs "node" ... it takes something
mesh_call mesh-controller.plan {"node": "novox"}
→ plan failed: plan needs "node" ... and not that
```
`mesh-controller.node` describes the same way and behaves the same way. Both are listed in the
node's own instructions as the mesh's verbs, so they are the first thing a session reaches for.
The verb is not broken where it runs: the same `plan novox --json` through the machine's login
shell answers in full. What cannot be done is calling it **through the console**, which the
node's instructions say is the only way to the mesh.
## Why
Not diagnosed past the symptom, and the symptom is specific enough to place it: the published
schema declares no properties, and an argument that is not in the schema does not reach the verb.
So the verb reports a missing argument that the console has dropped, every time, whatever is
sent. The two halves are each defensible — publish a schema, honour it — and together they make
a verb that can only refuse.
## Why it matters beyond this instance
**A tool that cannot be called is worse than one that is absent**, because it is listed. It
appears in `mesh_overview`, it describes itself, and the only thing it will say is that it wants
something it will not accept. A reader has no way to tell that from their own mistake, and the
obvious next move — pass the argument it names — is the one that does not work.
It also marks what the console does not check. Nothing compares a verb's published schema against
the arguments the verb actually requires, so a verb can be published in a shape that makes it
uncallable and nothing says so. That is the shape of an unenforced rule: the schema is believed,
and it is wrong.
## What a fix has to settle
- Where the schema for a seat's verbs comes from, and why these two publish an empty one while
the verbs beside them (`status`, `builds`, `push`) take their arguments and answer.
- Whether the console should refuse to publish a verb whose schema cannot satisfy it, rather than
listing one that can only fail.
- **How it is checked:** every verb the console lists is callable with the arguments its own
schema describes — a test that calls each with its schema's required set and asserts the answer
is not "needs" an argument the schema does not have.
## Worked around, for now
Through `<node>/node-login-shell.execute`, running the controller's own command line on the
machine. That is the path the console exists to replace, so it is a workaround and not an answer.
@@ -0,0 +1,39 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-media-catalog modules/plex]
fixed-by: mesh-media-catalog#2
amended-design:
---
# 245 — A media server's previews were reached through a link its container never mounted
## What was observed
On the home server the media server's scrub previews — 423 GB, about 57 000 preview files, one per
video — had not grown in seven months: the newest was from the day its preview folder was moved off
the server's own disk onto the large storage pool. The server's log repeated, live, that it could not
create a directory under its preview folder.
## Why
The move left the preview folder as a **symbolic link** inside the server's configuration directory,
pointing at the pool. The container mounts the configuration directory and the media libraries, and
not the pool's path, so inside the container the link pointed at nothing: the server could neither
show the previews it had nor make new ones, and said so only in its own log. Nothing the mesh reports
showed it — the container ran, and answered.
It is the failure the mesh's rule against symbolic links exists for: a link resolves differently in
every place that reads it, and a container is such a place.
## Resolution
The previews are a directory of the module, mounted at the server's preview path; where that
directory lives on a machine is the machine's placement setting, here the pool path the files were
already in. The link was removed at the cutover, the container recreated, the preview folders visible
inside it again and the log's errors gone. The previews are not backed up, by the operator's choice.
## How it is checked
The catalogue check that every mount is declared passes over the module. On the machine: the preview
folder seen from inside the container lists the same folders as the pool path.
@@ -0,0 +1,45 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller (settings), mesh-controller internal/inventory (SetSettings)]
fixed-by:
amended-design:
---
# 246 — Setting a module's settings replaces the whole layer, and nothing shows it first
## What was observed
An operator's agent set one placement — where a media server's preview folder lives on one machine —
with `settings set <module> {"places": {…}} --node <machine>`. The command answered that the setting
was recorded. The machine's plan then mounted the server's configuration from an empty default
directory: the node's layer had held the placements of three other directories, eight media
accesses, a public exposure, four endpoints and the account's ids, and every one of them was gone.
Caught before any push, by reading the plan. The previous layer was read back from the controller
database's nightly dump — the backups that issue 242 asked for, a few hours old.
## Why
The layer is a statement of the whole, by design (the inventory's `SetSettings`: "replacing rather
than merging … removing a key is done by leaving it out"). That design is sound; what is missing
around it is everything that makes it safe to use:
- **There is no way to read a layer.** `settings` has `set` and `clear`, no `show`; the console verb
likewise. To change one key, a person must already know every other key in the layer.
- **There is no history.** The row is updated in place; the previous values exist nowhere but a
database backup.
- **The answer does not say what was dropped.** "places" was reported as set; the six keys removed
were not mentioned.
## What would fix it
1. A way to read a layer — `settings show <module> [--node]` and the same on the verb.
2. `set` answers with what changed: keys added, changed and **removed**, so dropping one is never
silent. A removal could even require saying so.
3. The previous value kept: a settings history row per change, so an undo needs no backup.
## Status
Open. Until it is fixed: read the layer (from the store, read-only) before setting it, and compare
the node's plan before and after.