Author SHA1 Message Date
jschoubben 2a501af0eb Merge pull request 'Issues 239–242: a name taken over, a dry run recorded, every database dropped, no backups' (#83) from issues/239-242 into main 2026-10-05 00:25:50 +00:00
jschoubben bd7fc40099 Issue 239: name no installation 2026-10-05 02:24:27 +02:00
jschoubben 93d4d29b72 Issues 241, 242: one unreadable grants file dropped every database, and the mesh has no backups
241: the SDK harness read a failed read as "no consumer" and withdrew all seven databases on the
control node; withdrawal destroyed data in seven providers. Fixed in mesh-sdk 0.1.10, mesh-catalog#44
and mesh-host#22; the recovery and what it lost are recorded. 242: nothing backs anything up.
2026-10-05 02:24:11 +02:00
jschoubben 33c10aa85a Issues 239, 240, and 238's diagnosis: a module name taken over, a dry run recorded, the operator banned
239: two repositories defined photos; a rebuild to the catalogue's main replaced the app with a stub
and nothing refused it. 240: a dry-run build was recorded and its definition reached the machine.
238: the forge refused three ssh logins as the operator's account in 31 seconds; nothing in the mesh's
ssh configuration names the forge, and no jail ignores the mesh's own public addresses.
2026-10-04 17:31:29 +02:00
10 changed files with 259 additions and 423 deletions
@@ -1,70 +0,0 @@
---
status: active
initiated: 2026-10-04
touches:
- 03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md
- 03-DESIGN/00-as-is/15-the-agent-and-its-licences.md
- 02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
- 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
became: []
---
# 029 — The agent configured through its module
## What
Everything about the operator's coding agent that can be configured is configured through the agent
module's tools, and reaches the machines as **one plugin the mesh serves**:
- subagents, skills, slash commands, hooks, output styles and tool servers;
- the agent's settings and its instructions.
Each item is registered once, from any machine, and goes to one machine, several, or all of them,
including a machine that joins later. The mechanism is the one the module already uses for tool
servers: the registration is kept in the module's state on the bus, and each machine's instance writes
what applies to it. The operator chose the plugin route on the day this effort opened.
**Three scopes** (the operator's direction, the same day):
- **mesh:** in the mesh's plugin and the managed files, on every machine;
- **node:** the same places, rendered for one machine or a list of them;
- **home:** placed in the operator account's own agent directory on a machine, beside what the
person writes there by hand.
Instructions follow the same scopes: the mesh's piece, the node's piece, then further customisation per
machine. See [03](03-options.md).
## Why
The vendor gives the agent a machine-wide directory for its settings, its tool servers and one
instruction file, and **nothing machine-wide for skills, subagents, commands or hooks**. Those exist
only in a home or a project. So today they are copied into each home by hand, and they drift and go
stale. [01](01-what-is-configured-today.md) measures that on four machines.
A plugin is the vendor's own unit for carrying all of those at once. A machine-wide setting can name
a marketplace and enable a plugin from it. If the module serves the plugin and its own managed settings
enable it, the mesh gets a machine-wide place for everything the vendor left home-only, and the home
stays the person's ([ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md)).
## What it touches
- **The agent module's design** ([to-be 36](../../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md)):
its managed directory, its state, and its tools.
- **Module state** ([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)):
where registrations are kept, and which file content fits in a bucket.
- **Contributions** ([ADR 0210](../../02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)):
a module other than the agent's, say the forge's, wanting the agent to have a skill for it. That is a
contribution to the agent's seat, not a file it writes.
- **The managed settings key the operator sets** ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)),
which a settings tool would write rather than a hand-composed settings layer.
## Documents
- [01 — What is configured today](01-what-is-configured-today.md): the evidence.
- [02 — What the vendor allows](02-what-the-vendor-allows.md): plugins, marketplaces and managed
settings, as documented, with sources.
- [03 — Options](03-options.md): where each kind of item goes, how it is registered and stored, and
the questions a decision has to answer.
- [04 — What was confirmed](04-what-was-confirmed.md): the checks, tried on one workstation.
@@ -1,51 +0,0 @@
# 01 — What is configured today
Measured on 2026-10-04 on four machines that run the agent module: two workstations, the control node
and a home server. The figures count what sits in each operator account's agent directory in its home,
outside the module's managed directory.
## What sits in the homes
| what | workstation A | workstation B | control node | home server |
|---|---|---|---|---|
| rule files (`rules/`) | 4 | 2 | 2 | 1 |
| skills of the person's own (`skills/`, beside the vendor's synced ones) | 6 | 2 | 2 | 2 |
| subagents (`agents/`) | 0 | 1 | 0 | 0 |
| slash commands (`commands/`) | 0 | 0 | 0 | 0 |
| plugin marketplaces known | 1 | 2 | 1 | 1 |
## What that shows
- **Six files the design says the operator removes are still on every machine.** To-be 36 §1 lists the
predecessor's rule files and skills and leaves their removal to the operator, "once, on each
workstation". On all four machines, the two predecessor skills are present, byte-identical to each
other:
- one that switches licences through tools that no longer exist;
- one that names the predecessor's forge.
So a session can still load a skill whose every instruction fails.
- **One instruction, three versions.** The predecessor's node-identity rule file is on three machines,
with three different contents. It was written per machine and then left alone.
- **Two instruction sets that contradict each other, loaded together.** On a workstation, one session
reads two sets of instructions:
- the module's managed instruction file says to search the mesh's records first;
- the predecessor's rule files in the home say to search the predecessor's knowledge base first,
through tools that are no longer served.
Both are loaded, and neither says the other is stale.
- **A subagent exists on one machine only.** A reviewer for module definitions was written on one
workstation. The other three machines cannot use it, and nothing says it exists.
- **Settings are per home, and so per machine.** The agent's auto-mode environment, the rules that
decide which actions the agent may take unasked, is written in one home's settings file. It
describes another organisation's cloud, and it answers for this mesh's forge only through a list of
trusted domains. When the agent refused a merge the operator had approved, the only lawful fix was
a managed key ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)).
The agent could not change its own settings, and nothing else in the mesh could either.
## What already works the way this effort wants
Tool servers. A server registered through the module's register tool is kept in the module's state
on the bus, keyed `all.<server>` or `<node>.<server>`. Every instance watches that state and writes what
applies to it into the managed tool-server file, and a machine that joins later takes it at its first
start. Today one server is registered there, for one workstation. That is the shape this effort extends
to everything else.
@@ -1,106 +0,0 @@
# 02 — What the vendor allows
Read from the vendor's documentation on 2026-10-04; the agent installed on the machines measured in
[01](01-what-is-configured-today.md) was a 2.1 release. Each fact names the page it came from. Where the
documentation is silent, this says so. A fact that a design rests on is to be confirmed on one machine
before it is built on (see [03](03-options.md), *What to confirm first*).
## What a plugin can carry
A plugin is a directory with a manifest in `.claude-plugin/plugin.json` and, beside it, any of:
- skills, slash commands and subagents;
- hooks;
- tool servers (`.mcp.json`) and language servers;
- output styles, workflows, themes and monitors;
- a `bin/` directory;
- a `settings.json`.
Its components are namespaced by the plugin's name, so a subagent `reviewer` in a plugin `mesh` is
`mesh:reviewer`, and it never collides with a person's own of the same name.
— *plugins/manifest-reference, plugins/loading (name conflicts)*
**What a plugin cannot carry:**
- **Settings.** Only two keys of a plugin's `settings.json` take effect: the default agent and the
subagent status line. The rest are dropped. — *plugins/manifest-reference, settings*
- **Permission rules.** Not documented as a plugin capability.
- **Instructions.** A `CLAUDE.md` at a plugin's root is not loaded, and the validator warns about it.
Instructions reach a session through skills only. — *plugins/manifest-reference, standard layout*
## Marketplaces, and a marketplace on the machine's own disk
A marketplace is a `marketplace.json` listing plugins and where each comes from. Its sources include:
- a relative path inside the marketplace;
- a forge repository, a git URL or a subdirectory of one;
- a package from a registry;
- an archive over HTTPS;
- the output of a command.
**A marketplace can be a directory on the machine.** Its plugins with relative paths are **loaded in
place**, not copied into the cache. An edit takes effect at the next session start, or at
`/reload-plugins` in a running session, and the plugin's version need not change.
— *plugins/marketplace-reference (marketplace sources), plugins/loading (in-place and copied plugins)*
A plugin from any other source is copied into a cache in the home, under
`plugins/cache/<marketplace>/<plugin>/<version>/`. — *plugins/loading*
## What managed settings do with plugins
These keys work in the machine-wide managed settings file — *plugins/org*:
| key | what it does |
|---|---|
| `extraKnownMarketplaces` | registers a marketplace on every session of the machine |
| `enabledPlugins` | `true` installs and enables a plugin; `false` blocks and hides it at every scope. The managed value outranks every other scope |
| `strictKnownMarketplaces`, `blockedMarketplaces` | allow-list or block-list of marketplace sources |
| `strictPluginOnlyCustomization` | refuses skills, subagents, hooks and tool servers that come from neither a plugin nor managed settings |
| `allowManagedHooksOnly` | runs only the hooks from managed sources |
| `disableSideloadFlags` | blocks loading a plugin from the command line |
| `syncClaudeAiPlugins` | stops plugins synced from the vendor's web account |
**Installed without anyone being asked.** Once the settings reach a machine, the marketplace is
registered and the plugins installed at the next session start. A non-interactive run installs them in
the background. Managed plugins do not wait for the workspace trust prompt. — *plugins/org*
## What the managed settings file honours besides
`permissions` (with its default mode and the switch that disables bypassing it), `autoMode`, `hooks`,
`env`, `model`, `statusLine`, `outputStyle`, `apiKeyHelper`, and the managed-only switches for permission
rules, hooks and tool servers. — *managed-settings*
That `autoMode` is honoured from the managed file is documented. That it changes what the agent
refuses on these machines is still to be seen ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)
left that open).
## Tool servers: the exclusive file wins over a plugin's
When the managed tool-server file is present, as the module writes it, it is **exclusive**: only its
servers load. The vendor's web connectors load too when a managed key allows them. **A plugin's
`.mcp.json` servers are blocked.** — *managed-mcp (exclusive control)*
So a tool server registered through the module stays in the managed tool-server file. Putting it in
the plugin would silently stop it loading.
## Variables inside a plugin
- `${CLAUDE_PLUGIN_ROOT}`: the plugin's directory.
- `${CLAUDE_PLUGIN_DATA}`: a directory that survives updates.
- `${CLAUDE_PROJECT_DIR}`: the project's root.
These resolve in hook commands, tool and language server configuration, and the content of skills,
subagents and commands. A plugin's declared options (`userConfig`) can be marked sensitive; the agent
asks the person for them and stores them itself. — *plugins/manifest-reference (environment variables)*
## Reload
A running session does not see a changed plugin until `/reload-plugins` or a new session.
`/reload-plugins` reloads skills, subagents, hooks and servers. It does not restart monitors.
— *plugins/loading*
## Not documented
- a machine-wide directory for bare skills, subagents or commands. Only a plugin enabled by managed
settings puts them machine-wide;
- permission rules or instructions carried by a plugin.
@@ -1,159 +0,0 @@
# 03 — Options
The route is chosen: a plugin the mesh serves. What is left open is where each kind of item goes, how it
is registered and kept, and what the module does about what it finds in the homes.
## Scopes (the operator's direction, 2026-10-04)
The plugin is not the only place the module manages. **The agent's configuration is managed at three
scopes, and each item is registered at one of them:**
| scope | where it lands | reaches |
|---|---|---|
| **mesh** | the mesh's plugin, and the mesh's part of the managed files | every machine running the agent, including one that joins later |
| **node** | the same plugin and managed files, as rendered on that machine | one machine, or a list of them |
| **home** | the operator account's own agent directory on a machine (`~/.claude`) | that account on that machine |
Each machine renders its own plugin from the registrations that apply to it, so a node-scoped skill sits
in the same `mesh` plugin as a mesh-scoped one, on that machine only. The home scope places an item
where the person's own items live, without the plugin's prefix, as if written there by hand. The
difference is that the mesh knows it placed the item and can change or remove it.
**Instructions follow the same scopes.** The agent reads the managed instruction file first, then the
home's instruction file and its rule files, then the project's. These are concatenated, not overridden:
a later file does not cancel an earlier one, which is how the contradiction measured in
[01](01-what-is-configured-today.md) came about.
- **The mesh's piece** sits in the managed instruction file and is the same on every machine: how a
session on this mesh works, and the conventions.
- **The node's piece** sits in the same file, rendered per machine: its role, and instruction sections
registered for it.
- **Further customisation per machine** sits in the home: a rule file the module places, at the home
scope, beside whatever the person writes there by hand.
**What the home scope needs from [ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md).**
That ADR already lets the mesh own what it places in a home and hold the rest as found. So the module
owns each home item it placed, path by path, recorded in its state. It never writes, renames or
removes an item it did not place. A home item with the same name as one the person made is refused
at registration, never overwritten.
## Where each kind of item goes
[02](02-what-the-vendor-allows.md) puts a hard limit on the plugin: it carries skills, subagents, commands,
hooks, output styles and language servers, but no settings, no permission rules, no instructions, and
no tool server the exclusive managed file does not list. So there are four places, not one:
| kind | goes to | why there |
|---|---|---|
| skills, subagents, slash commands, output styles | **the mesh's plugin** | the only machine-wide place the vendor has for them |
| hooks | **the mesh's plugin**, with the scripts beside them | a hook's script can live in the plugin and be named through `${CLAUDE_PLUGIN_ROOT}`. A hook in the managed settings would need its script placed somewhere else |
| tool servers | **the managed tool-server file**, as today | the exclusive file blocks a plugin's servers |
| settings and permission rules | **the managed settings file**, beside the mesh's keys ([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)) | a plugin's settings are dropped |
| instructions | **the managed instruction file**, in sections | a plugin's instruction file is not loaded |
The plugin is reached through two managed keys the module already owns the file for:
`extraKnownMarketplaces`, naming a marketplace directory the module writes, and `enabledPlugins`, set to
`true` for the mesh's plugin. Neither is the operator's to set. Like the attribution key, they are the
mesh's keys and outrank whatever the operator sets.
### Option A — one plugin
Everything the mesh serves is in one plugin, `mesh`, so every invocation reads `mesh:<name>`. That is
simple, and the name says where an item came from.
### Option B — a plugin per source
One plugin for what the operator registers, and one for what other modules contribute (below). An item
then says in its name whether a person or a module definition put it there. But the operator has two
prefixes to remember, and an item has two owners to ask about.
*Leaning:* A. Where an item came from belongs in the module's list tool, not in the item's name.
## Who registers an item
- **The operator, through the module's tools**, from any machine, for one, several or all of them. The
pattern is the tool-server register tool's, extended to every kind:
- `claude_code_<kind>_list`, `_register`, `_unregister` for skills, subagents, commands, hooks, output
styles and instruction sections;
- `claude_code_settings_get` / `_set` and `claude_code_permission_allow` / `_deny` / `_ask` /
`_remove` for the managed settings.
- **Another module, through the agent's seat** ([ADR 0210](../../02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)).
The forge's module wanting the agent to know how pull requests are made here contributes a skill. It
declares the contribution in its definition, and the controller renders it to the agent's holder on
each machine where both run. That depends on the agent module holding a seat; today it holds none.
The two meet in the one plugin. A contribution and a registration with the same name are refused at
registration, and the list tool shows the owner of each.
## Where a registration is kept
The tool-server registrations live in a key-value bucket the module declares
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)),
keyed `all.<name>` or `<node>.<name>`. Skills differ: a skill is a folder, and it can carry scripts
and reference files beside its main file. Every message on the bus is limited to about a megabyte.
1. **One value per item.** The item and its files go in one value, refused above a limit well under
the bus's. It is simple, it fits the module's existing state, and every item measured in
[01](01-what-is-configured-today.md) takes 20 KB or less on disk,
the vendor's synced skills aside. But a skill with a large reference file cannot
be registered at all.
2. **The bus's object store for files, the bucket for the item.** Large files are stored in pieces and
the item names them. Nothing in the mesh uses the object store yet, so ADR 0201 would need
extending.
3. **A repository on the forge.** The plugin is built from a repository, and registering an item is a
commit. This is reviewable and versioned. But a register tool would have to write to the forge,
and the forge would sit on the path to every machine.
*Leaning:* 1 now, with the limit stated and checked at registration. 2 when an item outgrows it. 3
mixes the operator's configuration into the code review cycle, which it does not need.
## Where the plugin is written
The module owns the managed directory, so the marketplace goes under it, written whole by the
module's code:
- the marketplace file;
- one plugin directory beside it.
It is loaded in place, so a change takes effect at the next session, or at `/reload-plugins` in a
running one. Nothing is copied into the home.
## What the module does about what it did not place
[01](01-what-is-configured-today.md) found stale predecessor files on every machine. The module did not
place those, so they are held as found (ADR 0182). It can:
- **report** them: a status tool lists the home's skills, subagents, commands and rule files, says which
the mesh placed, and names those that duplicate a mesh item or call tools no longer served;
- **import** one on request: `claude_code_<kind>_import` takes an item from one machine's home and
registers it at a scope the operator chooses. A skill written by hand on one workstation becomes a
mesh, node or home item in one call. Removing the original stays the person's act.
`strictPluginOnlyCustomization` would make home items stop loading altogether, and home-scoped items
with them. That is the operator's choice to make through the managed settings, not a default of the
module.
## What to confirm first, on one workstation
1. A directory marketplace named in the managed settings, with its plugin enabled there, loads with
no prompt, in place, in an interactive session and in a non-interactive one.
2. The plugin's skills, subagents and commands are offered under `mesh:`, beside the home's own
items, with no collision.
3. A hook in the plugin runs, with its script found through `${CLAUDE_PLUGIN_ROOT}`.
4. The exclusive tool-server file still loads the console, and a server in the plugin does not load,
as documented.
5. `autoMode` in the managed settings changes what the agent refuses (ADR 0213's open point).
6. The account can read the marketplace in the managed directory, which root owns.
7. The managed instruction file and a home rule file the module placed are both loaded, in that order.
## Questions a decision has to answer
- One plugin or one per source (leaning: one).
- How a registration is kept, and the size limit (leaning: one value per item, with a stated limit).
- Whether the managed settings are set through tools writing the module's state, or through the
controller's settings layer as ADR 0213 has it. If both, which one wins on the same key.
- Whether the agent module holds a seat, so that other modules can contribute to it.
- The three scopes, and the home scope's ownership rule: the module owns exactly the home paths it
placed, recorded in its state, and refuses a name the person already uses.
- Whether settings take the same three scopes. The home's settings file is the person's own, so it is
left out unless the operator chooses otherwise.
@@ -1,37 +0,0 @@
# 04 — What was confirmed
On 2026-10-04, on one workstation running the agent's 2.1 release, the checks [03](03-options.md)
listed were tried with a probe. The probe was a directory marketplace holding one plugin named
`mesh`, which carried:
- a skill, a subagent and a slash command;
- a session-start hook running a script in the plugin;
- a tool server in the plugin's own `.mcp.json`.
The managed settings were set through the agent module's `managed_settings`
([ADR 0213](../../02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)),
on that machine's layer only, and sent by a push. The module rendered them into the managed file
within seconds, without a restart.
| # | check | result |
|---|---|---|
| 1 | a marketplace named in the managed settings, with its plugin enabled there, loads with no prompt, in place | **confirmed.** A non-interactive session registered the marketplace and enabled the plugin at start, with nothing asked. The plugin was not copied into the home's plugin cache and is not listed among installed plugins: it is read where it lies |
| — | a change to the plugin needs no version bump | **confirmed.** A skill added to the plugin's directory after the first session was offered by the next one |
| 2 | the plugin's items are offered under its name, beside the home's | **confirmed.** `mesh:probe-skill`, `mesh:probe-agent` and the command `/mesh:probe`. A collision with a home item of the same name was not tried |
| 3 | a hook in the plugin runs, its script found through `${CLAUDE_PLUGIN_ROOT}` | **confirmed.** The session-start hook ran its script. The vendor's validator asks for the placeholder to be quoted |
| 4 | the exclusive tool-server file still loads the console, and a plugin's server does not | **confirmed.** The session started the console and the registered servers, and never the plugin's server |
| 5 | `autoMode` in the managed settings changes what the agent refuses | **very likely.** With a probe rule forbidding one harmless read-only command, a session that was already running had that command refused moments after the rule was rendered, though the refusal gave no reason. It also suggests the rule reached a running session without a restart. A clean check needs a session whose only difference is the rule; the agent may not start one in its own auto mode, so it is left to the operator |
| 6 | the account can read the managed directory, which root owns | **confirmed** by what already runs: every session reads the managed instruction file from that directory |
| 7 | the managed instruction file and the home's are both loaded, managed first | **confirmed** by what already runs: a session lists the managed instruction file first, then the home's instruction file, then each of the home's rule files |
## What this changes in the options
- The plugin route works as documented, with no file copied into the home. The module's managed
directory can hold the marketplace.
- **A tool server stays in the managed tool-server file.** Check 4 closes that.
- **Undoing a managed setting is not the agent's to do.** When the probe was over, the agent tried to
clear its own machine's settings layer, and its own auto mode refused that as self-modification.
Setting it had been allowed only because it made the agent stricter. So the tools that set the
agent's settings and permissions are tools the operator calls, and the agent calling them for
itself is refused by the vendor's own guard. A design must not assume an agent can tidy up after
itself.
@@ -0,0 +1,43 @@
# 238 — Diagnosis
## 2026-10-04, from the control node
**What the forge refused, from the operator's uplink, in the ban's last minute** (the forge's ssh log):
```
14:05:28 Invalid user jochen from <uplink> port 38564
14:05:31 Accepted publickey for git from <uplink> port 45156 (the laptop's key)
14:05:44 Invalid user jochen from <uplink> port 52074
14:05:59 Invalid user jochen from <uplink> port 37978
```
Three refusals in thirty-one seconds — `maxretry = 3` — and the `gitea` jail banned the uplink at
14:06:00; `recidive` counted it the same second. Between the refusals the same key logged in as `git`:
the agent's own git operations were fine, and the refusals were ssh commands that named no user.
**Why they named the wrong user.** On the laptop and the workstation, `ssh -G <forge's public name>`
resolves to the operator's account and port 22: nothing in the ssh configuration the mesh writes
(`ssh-client`, to-be 29) names the forge. A bare `ssh <forge>` — or a git URL without `git@` — presents
the login name, which the forge does not have.
**Why the operator's uplink is bannable at all.** The jails' `ignoreip` (the `fail2ban` module's
`jail.local`) holds loopback, the mesh's private range and every private range (ADR 0186). A node at
home reaches the control node from the home's public address, which is none of those. Nothing the
mesh knows puts it there.
## What would have stopped it
1. **The forge in the ssh configuration the mesh writes**: a `Host` block for the forge's public and
internal names with `User git` and the forge's ssh port. It cannot be written today: the `git`
provision serves the forge's http port only, and the controller translates a served `port` to the
machine's published port but no other key (`ServedOn`), so an `ssh-port` would reach a consumer as
the container's 22, not the machine's 222.
2. **The mesh's own public addresses in every jail's ignore list.** Two sources: each node's public
egress as the hub sees it (the tunnel's peer endpoints — every node at home shows the uplink there),
or an operator setting naming them. The first is derived and stays true when the uplink changes;
the second is a value somebody must remember to edit.
## Status
Not located further: both remedies are design choices — (1) a served port the controller translates
by listen, (2) the open question 1 of the report.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 239 — A module name is taken over by another repository, and nothing refuses it
## What was observed
2026-10-04. Five of the photo app's six public names stopped answering on the control node — the API,
two client sites and two aliases — while the sixth answered with a different program. Nothing failed:
the controller composed, the host applied, every check passed.
Two definitions held the module name `photos`:
| | the app's own repository | the catalogue's `modules/photos` |
|---|---|---|
| built from | `photos.git`, branch `nox-mesh`, until 2026-09-28 | from 2026-10-04 04:14 |
| containers | server, admin, two client sites | server, an admin client |
| routes | six | one |
Rebuilding `photos` "to `main`" for [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)
(see [issue 227](../227-the-photo-apps-admin-client-asks-for-the-port-the-proxy-holds/00-report.md)) built the
catalogue's main, not the repository the module had been built from. The controller recorded the new
source, the module moved, the next push replaced the app with the stub, and two sites' containers and
five routes went with it. The stub had sat in the catalogue since 2026-09-03 without ever being the
running definition.
Restored the same day by building from `photos.git` again and removing the stub (the app repository's pull request and
mesh-catalog#278).
## Why it is an issue and not an incident
**A module's source is a fact the mesh records, and any build may overwrite it.** The build history
showed the switch plainly — `photos.git at nox-mesh` on one line, `mesh-catalog.git at main` on the
next — and nothing asked whether a module built from one repository should now come from another.
"The last build wins" is the rule in practice; it is written nowhere, and it lets a stale or unrelated
definition replace a working one silently.
## Open questions
1. Should a build whose source differs from the module's recorded source be refused unless it says so
explicitly (a `--move-source`, or the operator's confirmation)?
2. Should the catalogue's check refuse a module whose name another registered repository already
defines?
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 240 — A dry-run build is recorded, and what it built is applied
## What was observed
2026-10-04, restoring the photo app ([issue 239](../239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md)).
`mesh-controller build <repository> --ref <unmerged branch> --dry-run` was run to prove the branch built
before it was merged. The command's help says *"build and print the manifest, recording nothing"*.
Afterwards:
- `builds photos` listed the dry run as a build — `photos.git at <the unmerged branch>`, with its three
images — beside the real ones.
- On the control node, the module's state directory was rewritten at a time between the dry run and
the merged build: the secret files and environment files the branch's definition declares, which the
definition then running did not.
The content happened to equal what was merged a few minutes later, so nothing broke. Had the branch
been rejected in review, its definition would already have been on the machine.
## Why it matters
A dry run is how a change is proven before a person approves it. If it records the build and the mesh
acts on it, review becomes a formality: the unreviewed definition reaches a machine first.
## Open questions
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it?
@@ -0,0 +1,89 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply]
fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22
amended-design:
---
# 241 — One unreadable grants file dropped every database on the control node
## What was observed
2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every
connection to the seven databases it manages, dropped them, and seconds later created them again,
empty: the forge, the mail server's admin, the identity provider, the file-sync service, the
analytics service, the catalogue and the licence manager. postgres's own log shows it — connections
"terminating due to administrator command", then clients told the database "seems to have just been
dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42.
Nothing failed loudly. Each application kept running against an empty database: the mail server
refused every login (`relation "user" does not exist`), and those refused logins got the operator's
home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed;
the file-sync service answered 500.
Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's
177 GB of objects.
## Why
**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's
contributions file every five seconds. `readContributions` returned an empty list — without a log
line — when the file could not be read, was not JSON, or was for another requirement. The reconcile
pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was
`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied
its consumers, was enough. The next read was fine, so every consumer came back — to an empty database.
**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged
since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In
the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day
while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and
an agent session on the laptop was inspecting this provisioner's process from 17:52.
**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the
database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the
mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted
the site. Any of them would have lost its data to the same misread.
## Resolution
- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or
unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says
why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made.
A test reproduces the incident and fails on the old harness.
- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and
keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key
and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the
analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a
database out of service is an operator's tool (`postgres_retire_database`), and it renames to
`<name>_deleted_<date>` — nothing in the mesh drops a database.
- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only
while it holds exactly what the host wrote; otherwise it is moved aside to `<path>.removed-<time>`.
Directories with content were already kept.
## Recovery
No database had a backup newer than the migration; issue 242 is the gap. Each was restored from the
newest copy on the control node, by the same path: copy the source, start a throwaway postgres of its
version with no network, dump the one database, restore it beside the live one owned by the consumer's
role, check counts, stop the application, rename the empty live database aside
(`…_deleted_20261005`, kept) and the restored one into place, start, verify as a user.
| application | restored to | notes |
|---|---|---|
| forge | 2026-09-22 | git data complete and current; three repositories created since were adopted |
| mail | 2026-09-25 | all accounts |
| identity provider | 2026-09-26 | both realms |
| file-sync service | 2026-09-25 (a MariaDB dump, converted with the application's own `db:convert-type`) | objects intact; index entries without an object were previews, trash and stock sample files |
| analytics, catalogue | 2026-09-24 | |
| licence manager | none | schema re-created by its preparation step; licences re-adopted from the nodes |
The forge's restore had its own consequences: the forge watcher's admin account and the build agents'
registry accounts were created after the backup and were recreated (the first by the module's own
bootstrap command, the second by its provisioner), and every pull request and issue since 2026-09-22
is gone from the forge's records — the branches remain.
## How it is checked
The SDK test above; the postgres provider's test that withdrawal and retirement issue no `DROP`; the
host's test that a file holding more than the host wrote is moved aside, never deleted.
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 242 — The mesh has no backups
## What was observed
When seven databases were dropped on 2026-10-04 (issue 241), the newest copy of any of them was a
leftover of the migration: data directories and one dump, nine to twelve days old, found by searching
the control node's disk. Nothing in the mesh takes a backup, no record says what should be backed up,
and nothing would have said so until the day one was needed.
What survived did so by accident: git history because every repository is also cloned on the
operator's machines; the file-sync service's files because they live in the object store, which
nothing dropped; the mailboxes because they are files outside the database.
## Questions this must answer
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
buckets, the mail spool, the forge's repositories and data, the vault and the controller's own
records, each module's state directories? Is it declared by the module that owns the data, the way
a module declares its listens and its jails?
2. **How often, and how long kept** — a daily schedule? Rotation: how many daily, weekly, monthly?
3. **Full or incremental** — full dumps for databases, deltas (snapshots, deduplicating archives) for
large object stores and file trees?
4. **Where** — never only on the machine whose disk it copies. Spread across the mesh (the home server
holding the control node's, and the other way round), an external target, or both? Encrypted to
whom?
5. **Who runs it** — a seat (`mesh-backup`?) whose holder schedules and stores, with the verbs a person
needs: what was backed up and when, restore one database beside the live one?
6. **How it is proven** — a backup never restored is a hope. A scheduled restore of the newest copy into
a throwaway instance, compared with the live one, and a failure that reaches the operator.
## Why it matters beyond this incident
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.