Author SHA1 Message Date
mesh-admin b59d479992 Merge pull request 'ADR 0216: the agent's configuration is registered through its module, at three scopes, and served as one plugin' (#89) from decision/the-agent-configured-through-its-module into main 2026-10-05 09:50:44 +00:00
jochen f291d113c8 ADR 0216: the agent's configuration is registered through its module, at three scopes, and served as one plugin
Graduates research 029 and amends design 36 (section 8): skills, subagents,
commands, hooks and output styles in one nox-mesh plugin; servers, settings
and instructions in the managed files; mesh, node and home scopes.
2026-10-05 11:45:58 +02:00
jschoubben daf2f6d2b1 Merge pull request 'to-be 43: backups against mistakes (graduates research 030)' (#90) from design/43-backups-against-mistakes into main 2026-10-05 09:35:22 +00:00
mesh-admin 4924b26b3c Merge pull request 'ADR 0215: the machine's message bus is a node seat, and it is never restarted live' (#91) from decision/0215-the-message-bus-is-a-node-seat into main 2026-10-05 09:33:36 +00:00
jochen d088f8ec2f ADR 0215: the machine's message bus is a node seat, and it is never restarted live 2026-10-05 11:33:20 +02:00
jschoubben 431375d16c to-be 43: backups against mistakes, declared by modules, kept on the machine
Graduates research 030 into a design for ADR 0214, which merged without one and left the cycle
check failing on main.
2026-10-05 11:31:19 +02:00
jschoubben e51d6f1d86 Merge pull request 'Research 030 + proposed ADR 0214: backups against our own mistakes' (#88) from research/030-backups-against-our-own-mistakes into main 2026-10-05 09:30:35 +00:00
jschoubben a906cd8e67 ADR 0214: accepted by the operator 2026-10-05 11:30:33 +02:00
jschoubben 3d0ce3e6dd Research 030 and proposed ADR 0214: backups guard against mistakes and stay on the machine
The operator scoped backups to mistakes, not disasters (issue 242). Issue 238: fix the failed logins
at their source instead of exempting the operator's address.
2026-10-05 10:07:36 +02:00
13 changed files with 641 additions and 8 deletions
@@ -1,5 +1,5 @@
---
status: active
status: graduated
initiated: 2026-10-04
touches:
- 03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md
@@ -8,7 +8,9 @@ touches:
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md
- 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
became: []
became:
- 02-DECISIONS/0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md
- 03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md
---
# 029 — The agent configured through its module
@@ -0,0 +1,29 @@
---
status: graduated
became: [03-DESIGN/01-to-be/43-backups-against-mistakes.md, 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md]
initiated: 2026-10-05
touches: [04-ISSUES/242-the-mesh-has-no-backups, 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md, 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md, 03-DESIGN/01-to-be/31-a-module-declares-its-fail2ban-jail.md]
---
# 030 — Backups against our own mistakes
**What.** How the mesh keeps restore points of its data: what is copied, how often, how long it is
kept, full or incremental, where it lives, who runs it and how it is proven to restore.
**Why.** Issue 242: nothing in the mesh backs anything up, and when one misread file dropped every
database on the control node (issue 241), the newest copies were migration leftovers nine to twelve
days old, found by searching a disk.
**The scope is set by the operator, and it is narrow on purpose: mistakes, not disasters.** A
backup here protects against what a person, an agent or the mesh itself does wrong — a dropped
database, a deleted bucket, a bad migration, a file overwritten — and not against a disk dying or a
building burning. Losing data to a disaster is an accepted risk. That removes the off-site copy, the
cross-site transfer and the second key-holder from the problem, and leaves the part that would have
saved the night of issue 241.
**What it touches.** The store providers (each knows how to dump its own store consistently), the
host's scheduled steps (ADR 0053, built), the node-wide composition pattern a module's jail already
uses (to-be 31), and the operator's output channel (research 028) for saying a backup failed.
Documents: [01 — what the mesh holds](01-what-the-mesh-holds.md),
[02 — options and a proposal](02-options-and-proposal.md).
@@ -0,0 +1,55 @@
# 01 — What the mesh holds, measured 2026-10-05
Four machines: the control node (hosted, holds every public service), the home server (media, home
automation, a large ZFS pool), a workstation and a laptop. Sizes are apparent sizes, rounded.
## Nothing backs anything up
On every machine: no backup tool other than `rsync` and `pg_dump` is installed, no systemd timer and
no cron line mentions a backup, dump or snapshot. Every live data directory is on ext4 except the
home server's pool (ZFS), so a filesystem snapshot is available only there.
## The control node — the data that cannot be recreated
| what | size | how it changes |
|---|---|---|
| object store (file-sync service's files, photos) | 183 GB | slowly; files added, rarely rewritten |
| forge (repositories, attachments, its database) | 7 GB | daily |
| MS SQL Server databases | 5 GB | daily |
| mail (mailboxes; accounts in postgres) | 2 GB | continuously |
| postgres (forge, mail admin, identity, file-sync index, analytics, catalogue, licence manager) | ~2 GB | continuously |
| file-sync service's own directory, website, analytics | ~4 GB | slowly |
| MongoDB | 0.4 GB | daily |
| the mesh's own records (controller, vault, module state) | ~1.5 GB | continuously |
Recreatable and not worth copying: container images (210 GB), the artifact registry (40 GB — every
artifact is rebuilt from git), a 115 GB speed-test bucket and a 7 GB pre-migration object-store copy.
Free space: 833 GB on the filesystem holding the data, 2.9 TB on a second one.
## The home server
MS SQL Server 80 GB, postgres and a self-hosted backend platform ~3 GB, chat server 6 GB, home
automation, network controller and time-series data each under 2 GB, and the media services'
libraries (tens of GB, mostly cover art and metadata they re-fetch). The 89 TB media library is
replaceable by its nature and out of scope. Free: 31 TB on the pool, 453 GB on the system disk.
## The workstation and the laptop
The workstation has 142 GB under its services directory and 31 GB of container volumes; the laptop
3 GB. Mostly development; what among it is data nobody can regenerate is for each module to say.
## Between the sites
Control node to home server ~285 Mbit/s, home server to control node ~19 Mbit/s. Irrelevant now that
backups stay on the machine whose data they hold, recorded because it is why an off-site copy would
have been expensive.
## What issue 241 says about the requirement
- The mistake was noticed within hours. A restore point a day old would have lost a day.
- The restore had to go *beside* the live database, not over it, and that worked well.
- A copy of a live postgres data directory needed a throwaway server of the right version to read;
a logical dump would have restored directly.
- The data that survived was the data outside the dropped stores. A backup that lives inside the
store it protects — a database's own snapshot table, a bucket's own versions — dies with a drop.
@@ -0,0 +1,74 @@
# 02 — Options and a proposal
## Who decides what is backed up
1. **A central list** on the backup holder. Rejected: it is the attentiveness rule ADR 0030
rejected — a store added and not listed is silently unprotected.
2. **Each module declares its own data, a node-wide holder composes them.** The pattern of to-be 31
(a module declares its jail; the mesh composes them per node). A store provider declares *how*
to take a consistent copy (a dump command), a module with plain files declares *which* paths. The
holder composes every declaration on the node into one schedule. **Proposed.**
The data a module keeps in a database it gets from a provider is backed up by the provider, which
dumps every database it serves — so a consumer declares nothing, and a new consumer is covered the
day it is provisioned.
## Full or incremental
- **Databases: a full logical dump every time** (`pg_dump -Fc`, MS SQL `BACKUP DATABASE`,
`mongodump`). A dump restores with the store's own tool into a database beside the live one —
issue 241's recovery without the throwaway server — and is consistent, which a copy of a live data
directory is not.
- **Everything goes into one deduplicating repository per node** (restic or borg). Each night is a
complete restore point, yet only changed chunks cost space: the object store's 183 GB is copied
once, then each night adds what changed. This removes the full-vs-delta trade-off rather than
choosing a side.
Considered for the object store alone: the object store's own versioning with a lifecycle rule.
Rejected as the only copy — it lives inside the store, and a removed bucket or data directory takes
its versions with it.
## Where
On the same machine, outside every data directory the mesh manages, on a second filesystem where the
machine has one (the control node does). Not off-site: the scope is mistakes. The repository is
encrypted anyway (both tools require it); its key is a mesh secret (ADR 0085), so a person can
restore without the holder.
## How often, how long
- **Nightly**, at a quiet hour, as a scheduled step (ADR 0053).
- **On demand before a risky act** — a migration, a retirement, an operator's experiment — through a
verb; the act's own tooling can call it.
- **Kept: 14 daily, 8 weekly, 6 monthly.** A mistake is usually noticed within days, sometimes weeks
(a deleted file nobody opens). Six months bounds the space a slowly-noticed mistake needs.
Estimated cost on the control node: ~200 GB for the first night, a few GB a night after; well within
the second filesystem's 2.9 TB.
## Who runs it
A node seat, **`node-backup`**, held on every machine that has data by one module (named for the tool
it wraps). It receives the declarations, runs them, keeps the repository, and offers the verbs a
person needs:
- what is backed up here, and the last good night of each;
- take a backup now;
- restore one item **beside** the live one — a database to `<name>_restore`, a path to
`<path>.restored-<date>` — never over it. Swapping it in stays a person's act, as in issue 241.
## How it is proven
- Every run checks its own result; a failed or skipped night goes to the operator's output channel
(research 028), not only a log.
- Weekly: the repository's integrity check, and one database restored from the newest dump into a
throwaway instance and counted against the live one.
- A machine with data and no successful backup in 48 hours is a problem the mesh's status shows.
## Open questions
- restic or borg — both fit; restic is a single binary with no server, which suits a module.
- Whether the workstation and laptop take part at all, or only once a module there declares data.
- The mail spool is files and the forge has a dump command of its own; whether the forge's
repositories are worth backing up at all when every clone is a copy (issue 241 says the forge's
*database* is the part with no other copy).
@@ -0,0 +1,66 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md
---
# 214. Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data
## Context
Nothing in the mesh backed anything up (issue 242). When a misread file dropped every database on the
control node (issue 241), recovery took a night and used copies nine to twelve days old.
The operator sets the scope: backups exist for **mistakes** — a person's, an agent's, the mesh's own
— not for disasters. Losing data to a dead disk or a lost site is accepted. Research 030 measured
what the machines hold and found no backup tooling anywhere.
## Considered Options
1. **A central list of what to back up.** Rejected — whatever is not listed is unprotected, silently.
2. **Each store's own mechanisms** (bucket versioning, database snapshots). Rejected as the only copy:
they live inside what they protect, and a drop takes them with it.
3. **Off-site copies.** Out of scope by the operator's decision; recorded so the absence is a choice.
4. **Modules declare, a node seat composes, the copy stays on the machine.** Adopted.
## Decision
**A module declares the data it owns; a node seat, `node-backup`, composes every declaration on the
machine and keeps nightly restore points there.** A store provider declares how to dump each database
it serves, so a consumer of a store declares nothing. A module with files declares their paths.
**Databases are dumped in full, logically, every night; everything lands in one encrypted,
deduplicating repository per machine,** so every night is a complete restore point and only what
changed costs space.
**Kept: 14 daily, 8 weekly, 6 monthly.** A backup is also taken on demand before a risky act.
**The repository is on the machine, outside every directory the mesh manages,** on a second
filesystem where there is one. Its key is a mesh secret.
**A restore goes beside the live data, never over it.** Swapping it in is a person's act.
**A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours
shows in the mesh's status.
## Consequences
- Adding a store provider means declaring its dump; the catalogue check can refuse a store provider
that declares none.
- The control node's first backup is ~200 GB, then a few GB a night.
- A dead disk or a lost machine still loses its data, by choice.
## How it is checked
The catalogue check refuses a module providing a store seat without a backup declaration. The holder's
weekly restore test restores one dump into a throwaway instance and compares counts. The mesh's
status lists every machine whose last good backup is older than 48 hours.
## References
- [04-ISSUES/242](../04-ISSUES/242-the-mesh-has-no-backups/00-report.md), [04-ISSUES/241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
- [01-RESEARCH/030](../01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md)
- ADR 0030 (data outlives its declaration), ADR 0053 (scheduled steps), ADR 0085 (a secret is a provision)
@@ -0,0 +1,74 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md
---
# 215. The machine's message bus is a node seat, and it is never restarted live
## Context
Every machine runs a D-Bus system bus, and the workstations a session bus per login. The service
manager, logind, the network manager, the Bluetooth stack, the GPU switcher, the keyring, the
desktop portal and the power module's sleep lock all speak on it. Nothing in the mesh owned it.
On 2026-10-04 a full upgrade on a workstation restarted the system bus in the middle of the upgrade.
From then on logins hung, sshd answered nothing, and the machine's host stopped reporting, until a
person rebooted it at its keyboard. The mesh had no record of what the bus is, no view of it, and no
rule about when it may restart.
## Considered Options
1. **Leave the bus to the distribution.** Rejected: the outage above is what that gives, and nothing
would ever say the bus is unwell.
2. **Make it part of the service manager's holder.** Rejected: the bus is a separate program with its
own policy, its own clients and its own failure. A machine can have a healthy service manager and
a wedged bus, which is exactly what happened.
3. **A node seat held by a `dbus` module.** Chosen.
## Decision
**1. `node-message-bus` is a node seat in the mesh's own set.** Every machine has one. The first
holder is a module named `dbus`, which owns the bus implementation's package and its system
service, and serves tools to look at both buses.
**2. The bus is never restarted live.** The holder declares the bus running and enabled, and never
restarts or reloads it on any change. A new version of the bus takes effect at the machine's next
boot. A module's change that needs the bus to pick up a policy uses the bus's own reload of policy
files, which keeps every connection, never a restart.
**3. Curated events, never traffic.** The holder publishes on the mesh's bus only what matters about
the machine's bus:
- the bus's health (up, stalled, restarted);
- a well-known system service appearing on the bus or leaving it;
- a policy denial.
The bus's traffic, which carries secrets, notification text and the clipboard, never leaves the
machine. The holder's tools let a person watch it, bounded in time, on request.
**4. The seat receives nothing yet.** Packages ship their own D-Bus policy and service files, and no
module writes one of its own today. When one does, it is a contribution to this seat (ADR 0210, ADR
0212), and the seat lists the kind then.
## Consequences
- Phase 1 of [to-be 42](../03-DESIGN/01-to-be/42-the-machines-modules-in-order.md) gains `dbus` on
every machine.
- A full upgrade that brings a new bus no longer breaks a running machine through the mesh. The
distribution's own upgrade still restarts it, so the rule is enforced only for what the mesh
does. The holder's check says whether the running bus is older than the installed package, which
is the sign that a reboot is due.
- **What got harder:** a fix to the bus itself waits for a reboot. That is the price of never taking
every login on the machine down with it.
## How it is checked
| Rule | Checked by |
|---|---|
| `node-message-bus` is a node seat of the mesh's own set | the seat table's tests |
| The bus's service is declared running and enabled, with no restart or reload trigger | the dbus module's manifest test |
| No traffic is published, only the curated events | the dbus module's tests over its event code |
| A bus older than its installed package is said | the dbus module's check, over a recorded answer |
@@ -0,0 +1,167 @@
---
topic: what runs on it
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
---
# 216. The agent's configuration is registered through its module, at three scopes, and served as one plugin
## Context
The operator wants everything about the coding agent that can be configured to be configured through
the agent module's tools ([research 029](../01-RESEARCH/029-the-agent-configured-through-its-module/00-overview.md)).
That covers subagents, skills, slash commands, hooks, output styles, settings, permissions,
instructions and tool servers. Each is registered once, from any machine, for one machine, several or
all of them.
The agent module already manages three files in the agent's machine-wide managed directory: the
settings, the tool servers and one instruction file ([to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md)).
Tool servers registered through its tools are kept in its state on the bus and reach every machine
they apply to. The vendor has **no machine-wide place for skills, subagents, commands or hooks**; they
live only in a home or a project. Measured on four machines
([029/01](../01-RESEARCH/029-the-agent-configured-through-its-module/01-what-is-configured-today.md)),
what was copied there by hand had drifted and gone stale:
- two skills of a retired system were on all four machines;
- one rule file existed in three versions;
- two contradicting instruction sets were loaded into the same session;
- a subagent existed on one machine only.
The vendor's **plugin** carries skills, subagents, commands, hooks and output styles. A machine-wide
setting can name a marketplace and enable a plugin from it
([029/02](../01-RESEARCH/029-the-agent-configured-through-its-module/02-what-the-vendor-allows.md)).
Tried on one workstation ([029/04](../01-RESEARCH/029-the-agent-configured-through-its-module/04-what-was-confirmed.md)):
- a plugin in a directory marketplace, enabled by the managed settings, loads in every session with no
prompt, read in place, and an edit reaches the next session with no version change;
- its items are offered under the plugin's name;
- its hooks run;
- its tool servers are blocked by the exclusive managed tool-server file.
A plugin cannot carry settings, permission rules or instructions.
The operator also set the scopes. A skill may be meant for:
- every machine;
- one machine;
- one machine's own account, as if written there by hand.
The instruction file the same: the mesh's piece, the node's piece, and further customisation per machine.
## Considered Options
1. **Copy everything into each home**, owned by the mesh path by path. Rejected as the *only* place:
the home is the person's ([ADR 0182](0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md)),
a mesh item there is indistinguishable from the person's by name, and it fills the directory where
the drift was measured. Kept as one scope of three (below).
2. **A plugin per source:** one for the operator's registrations, one for what other modules
contribute. Rejected: two prefixes to remember for one agent. Where an item came from belongs in
the module's list, not in its name.
3. **Registrations in a repository on the forge**, the plugin built from it. Rejected for now:
registering would be a commit, and the forge would sit on the path to every machine. Configuration
would enter the code review cycle, which it does not need.
4. **One plugin, `nox-mesh`, plus the managed files the module already writes, at three scopes, all
registered through the module's tools and kept in its state on the bus.** Chosen.
## Decision
Option 4.
**1. What goes where.** Each kind of item goes to the one place the vendor honours for it:
| kind | place |
|---|---|
| skills, subagents, slash commands, hooks, output styles | the plugin `nox-mesh` |
| tool servers | the managed tool-server file, as today: the exclusive file blocks a plugin's servers |
| settings and permission rules | the managed settings file: a plugin's settings are dropped |
| instructions | the managed instruction file, in sections: a plugin's instruction file is not loaded |
The plugin is named `nox-mesh` (the operator's choice): the name its items carry in every session, and
not one a person's own plugin is likely to take. It lives in a marketplace directory inside the module's
managed directory, written whole by the module's code. The managed settings name that marketplace and enable the plugin. Those two keys
are the mesh's, laid last like the attribution key (ADR 0213), and no setting replaces them.
**2. Three scopes.** Every registration names one:
- **mesh:** every machine running the agent, including one that joins later.
- **node:** one machine, or a list of them. Rendered into the same plugin and managed files, on those
machines only.
- **home:** the operator account's own agent directory on one machine. The item is placed where the
person's own items live, without the plugin's prefix.
Settings and permission rules take the mesh and node scopes only. The home's settings file stays the
person's.
**3. Instructions follow the scopes.** The managed instruction file holds, in order:
1. the mesh's piece, the same everywhere;
2. the node's piece: its role, and the sections registered for it.
Further customisation per machine is a rule file placed at the home scope. The vendor concatenates
these and does not override, so the module's status tool names any section that contradicts another,
or that calls a tool the mesh no longer serves.
**4. Registered through tools, kept on the bus.** For each kind, the module serves `list`, `register`
and `unregister` tools; for settings, tools that read and set them at a scope. Each registration is a
key in the module's state ([ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)):
- an item and its files are one value, refused above **256 KiB**, well under the bus's message limit;
- every instance watches the state and renders what applies to its machine.
The settings registered this way are laid over the `managed_settings` layer of ADR 0213. In order:
ADR 0213's setting, then the mesh scope, then the node scope, then the mesh's own keys.
**5. The home scope owns only what it placed.** The module records each home path it placed, in its
state. It writes, changes and removes only those. It refuses to register a name the person already
uses there, rather than overwrite it.
**6. What the module did not place, it reports and can import.**
- A status tool lists the home's items and says which the mesh placed. It also names any that
duplicate a mesh item or call tools no longer served.
- An import tool registers an item found in one machine's home at a scope the operator chooses.
- Removing the original stays the person's act.
- Whether home items load at all stays the operator's choice, through a vendor setting in the managed
settings.
**7. Changing the agent's own settings is the operator's act.** The vendor's guard refuses an agent that
loosens its own settings. A settings or permission tool is called on the operator's word, and the
module does not try to get around that refusal.
## Consequences
- One registration puts a skill, a subagent or a rule on every machine, on some, or in one account. A
machine that joins takes the mesh and node items at its first start. Nothing is copied by hand.
- The plugin's items are named `nox-mesh:<name>`, and a person's own items keep their names. Nothing the
mesh adds can shadow them.
- A change reaches the next session on each machine, or a running one at its next plugin reload.
- **What got harder:**
- an item larger than 256 KiB cannot be registered until the state can hold files in pieces;
- the module's state now holds file content, not only small records;
- the stale files already in the homes stay until the person removes them. The module names them;
it does not remove them.
- **Not decided here:** another module contributing a skill or a subagent to the agent through a seat
([ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)).
The agent module holds no seat yet. When one is decided, contributions land in the same plugin.
## How it is checked
| Rule | Checked by |
|---|---|
| each kind lands in its one place | the module's render test: a registered skill, subagent, command, hook and output style appear in the plugin; a tool server in the managed tool-server file; a setting in the managed settings file; an instruction section in the managed instruction file |
| scopes | the same test, for one machine of two: a mesh item on both, a node item on one, a home item only in that machine's home |
| the marketplace keys are the mesh's | the render test: a setting naming either key is overridden |
| the home scope owns only what it placed | the module's test: a name the person already uses is refused; unregistering removes only the placed path |
| the size limit | the module's test: an item above 256 KiB is refused at registration |
| live | a skill registered at the mesh scope is offered as `nox-mesh:<name>` in a new session on each machine |
## References
- [Research 029](../01-RESEARCH/029-the-agent-configured-through-its-module/00-overview.md) — the evidence, the vendor's rules, and what was confirmed
- [ADR 0213](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md) — the managed settings setting this lays over
- [ADR 0182](0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md) — what the mesh may do inside a home
- [ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md) — module state on the bus
- [to-be 36](../03-DESIGN/01-to-be/36-the-operators-agent-on-a-machine.md) — the design this amends
+3
View File
@@ -313,6 +313,9 @@ python3 00-META/checks/index.py fail if stale
- **0209** — [A login on a node moves that node to the account it logged in to; an API key is added from any node, sealed](0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md)
- **0211** — [A machine's power is a node seat, its moments take contributions, and its states are events](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)
- **0213** — [The operator sets the agent's managed settings through the agent module, under the mesh's own keys](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
- **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)
- **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)
### How it is built
@@ -2,8 +2,9 @@
layer: to-be
status: in-progress
code: [mesh-catalog modules/claude-code]
updated: 2026-10-04
updated: 2026-10-05
decisions:
- 02-DECISIONS/0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md
- 02-DECISIONS/0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md
- 02-DECISIONS/0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md
- 02-DECISIONS/0206-a-node-reports-the-anthropic-grant-it-holds-and-the-licence-manager-adopts-a-licence-by-refreshing-it.md
@@ -41,7 +42,8 @@ The agent reads a machine-wide, administrator-owned configuration directory unde
by the vendor: a managed settings file that outranks every user and project setting; a key in it that
adds HTTP tool servers *beside* a person's own without blocking them; and a managed instruction file every
session reads before the user's and the project's. The agent has **no** machine-wide directory for
rules, skills, slash commands or hooks; those exist only under a home or a project.
rules, skills, slash commands or hooks; those exist only under a home or a project — or in a **plugin**
the managed settings enable, which is how the mesh puts them on every machine (§8).
So the mesh's part of the agent's configuration lives there, **owned whole by the module**, and the home
is left alone. What the predecessor shipped as two rule files and two skills folds into the managed
@@ -205,6 +207,60 @@ answer is a package repository for this ecosystem as a seat
and trusted by every node's package manager; not built, and not this module's to build. The vendor's own
installer is rejected: it puts a self-updating binary under the person's home, invisible to the mesh.
## 8. The agent's configuration, registered at three scopes
Everything about the agent that can be configured is registered through this module's tools, once, from
any machine, and kept in the module's state on the bus
([ADR 0216](../../02-DECISIONS/0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)). Every instance watches that state and writes what applies to its
machine. A machine that joins later takes it at its first start.
**What goes where.** The vendor honours each kind of item in one place only, so the module writes four:
- skills, subagents, slash commands, hooks and output styles go into **one plugin named `nox-mesh`**. It
sits in a marketplace directory inside the managed directory, written whole by the module and read in
place by the agent. Its items are offered as `nox-mesh:<name>`, so nothing the mesh adds shadows a
person's own item;
- tool servers go into the managed tool-server file, as in §4. The exclusive file would block a
plugin's servers;
- settings and permission rules go into the managed settings file, as in §2;
- instructions go into the managed instruction file, as sections (§3).
The managed settings name the marketplace and enable the plugin. Those two keys are the mesh's, laid
last with the attribution key, and no setting replaces them.
**Three scopes.** Every registration names one:
- **mesh:** every machine running the agent;
- **node:** one machine or a list of them, rendered into the same plugin and files there only;
- **home:** the operator account's own agent directory on one machine, where the item sits as if
written there by hand.
Settings take the first two scopes only. In the managed settings file they are laid in this order:
the operator's `managed_settings` setting (§2), then the mesh scope, then the node scope, then the
mesh's own keys.
**Instructions follow the scopes.** The managed instruction file holds the mesh's piece, then the node's
piece: its role, and the sections registered for it. Further customisation per machine is a rule file
placed at the home scope. The agent concatenates these and does not override, so the status tool names
a section that contradicts another, or that calls a tool the mesh no longer serves.
**The home scope owns only what it placed.** The module records each home path it placed and touches
only those (ADR 0182). It refuses to register a name the person already uses there.
**The tools.**
- For each kind: list, register and unregister. Each register takes a scope, and a list says where
each item came from.
- For settings and permission rules: read, and set at a scope.
- A **status tool** lists the home's own items beside the mesh's and names the stale ones.
- An **import tool** registers an item found in one machine's home at a scope the operator chooses.
An item and its files are one value in the state, refused above 256 KiB.
**Changing the agent's own settings is the operator's act.** The vendor refuses an agent that loosens its
own settings, and the module does not route around that refusal. A settings tool is called on the
operator's word.
## How it is checked
| Check | Defends |
@@ -215,6 +271,9 @@ installer is rejected: it puts a self-updating binary under the person's home, i
| a switch asked of the seat through the console changes the licence and the token on the node; no tool answer and no log line holds a token | ADR 0183 |
| the API-key binding writes nothing under the home and the agent authenticates through the helper | ADR 0183 |
| the module's test: keys set in `managed_settings` (an auto-mode allow list, a permissions list) appear in the rendered managed settings file, a setting naming the attribution, the connectors key or a key-helper is overridden, and a key-helper appears only for an API-key binding | ADR 0213 |
| the module's render test: a registered skill, subagent, command, hook and output style land in the `nox-mesh` plugin; a tool server, a setting and an instruction section in their managed files; for one machine of two, a mesh item on both, a node item on one, a home item only in that home; a setting naming the marketplace keys is overridden | ADR 0216 |
| the module's test: a home name the person already uses is refused, unregistering removes only the placed path, and an item above 256 KiB is refused | ADR 0216, ADR 0182 |
| live: a skill registered at the mesh scope is offered as `nox-mesh:<name>` in a new session on each machine | ADR 0216 |
| the console's provision resolves by co-location; a machine without the console refuses the module by name | ADR 0027, ADR 0152 |
| a new session on the assigned workstation lists the console's five tools under `mesh` ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)) and answers "which node am I" from the instruction file | the exit of the build |
@@ -4,6 +4,7 @@ status: in-progress
code: [mesh-catalog, mesh-controller, mesh-host]
updated: 2026-10-04
decisions:
- 02-DECISIONS/0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md
- 02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md
- 02-DECISIONS/0040-what-a-module-is.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
@@ -83,7 +84,7 @@ until ADRs 0165 and 0166 are accepted.
- `power` holds `node-power` ([ADR 0211](../../02-DECISIONS/0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)). Other modules contribute code for
its moments (after boot, before sleep, after waking, before shutdown, on mains, on battery), and it
publishes the machine's power states on the bus.
- `dbus` holds the message bus. Modules shipping D-Bus policies or services contribute them to it.
- `dbus` holds `node-message-bus` ([ADR 0215](../../02-DECISIONS/0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)). Modules shipping D-Bus policies or services contribute them to it.
It shares curated events, never raw traffic. An upgrade never restarts the bus live: its package
waits for a reboot. A live restart in the middle of a full upgrade took down a workstation's
logins on the day this was written.
@@ -0,0 +1,81 @@
---
layer: to-be
status: designed
code: []
updated: 2026-10-05
decisions:
- 02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0085-a-secret-is-a-provision.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 43 — Backups against mistakes: a module declares its data, the node keeps restore points
**Every machine with data keeps a restore point of it for every night of the last two weeks, every
week of the last two months and every month of the last half year, on the machine itself.** It is
there for the day a person, an agent or the mesh does something wrong — drops a database, empties a
bucket, runs a bad migration — and not for the day a disk dies (ADR 0214).
## The shape
- **A node seat, `node-backup`,** held on each machine by one module, named for the tool it wraps.
It keeps one encrypted, deduplicating repository on the machine and runs a nightly scheduled step
(ADR 0053). The repository's key is a secret provisioned to the holder (ADR 0085), held in the
vault, so a person can open the repository without the holder running.
- **A module declares its data in its manifest,** naming no node and no absolute path (ADR 0112),
in one of two forms:
- **a dump** — for a store provider: the command that writes a consistent, logical copy of each
database it serves, run by the provider in its own container, its output handed to the holder.
The provider covers every consumer it provisions, so a module that only *uses* a database
declares nothing.
- **paths** — for files a module keeps itself (mailboxes, the forge's attachments, a service's
state directory), named through the module's own directory references.
- **The mesh composes the declarations per node,** as it composes jails (to-be 31) and filters: the
holder receives, as contributions, exactly the data of the modules assigned to its machine. A
module assigned is covered the next night; a module unassigned stops being backed up, and its
restore points age out by the rotation, never at once.
## A night
Each declared dump runs and writes a full logical copy; each declared path is read as it stands. All
of it goes into the repository as one snapshot, tagged by module. The repository keeps only chunks it
has not seen, so the object store's first night costs its full size and later nights cost what
changed. Then the rotation prunes to 14 daily, 8 weekly and 6 monthly snapshots. A dump that fails
fails the night for that module only; the others are still taken.
## The verbs
On the seat, for a person or an agent:
- **what is backed up here** — each module, what it declared, its last good night and its size;
- **take one now** — for one module or all, before a risky act; a migration or a database's retirement
calls it first;
- **restore** — one module's database or path, from a named night, **beside** the live one: a database
as `<name>_restore` owned by the consumer's role, a path as `<path>.restored-<date>`. Swapping it
in stays a person's act. Nothing restores over live data.
## Where the repository lives
On the machine, outside every directory the mesh manages, on a filesystem other than the live data's
where the machine has one. The holder's module names the place as a machine setting, never a path in
a manifest. Not off-site: a lost machine loses its backups with its data, by the operator's choice.
## Proving it
- A night that fails, or does not run, reaches the operator's output channel, naming the module.
- Weekly, the holder checks the repository's integrity and restores the newest dump of one database,
in rotation, into a throwaway instance with no network, comparing table row counts with the live
database.
- The mesh's status lists every machine whose last good night is older than 48 hours.
## Not in this design
- Copies off the machine, and encryption to anyone but the mesh's own vault.
- The media library and anything else a module declares no data for.
- Recreatable things: container images, the artifact registry, caches.
## How it is checked
The catalogue check refuses a module that provides a store seat and declares no dump. The weekly
restore test above, and the 48-hour status line, are the running checks.
@@ -37,7 +37,22 @@ mesh knows puts it there.
or an operator setting naming them. The first is derived and stays true when the uplink changes;
the second is a value somebody must remember to edit.
## The operator's call (2026-10-05): fix the logins, not the jail
Remedy 2 is rejected. **A machine of the mesh should not fail logins, and when it does, the ban is
the jail working.** Exempting the operator's address would hide exactly the failures worth seeing —
here a misconfigured ssh client, and in the same week a mail server that had lost every account
(issue 241). Both bans were genuinely failed logins; fail2ban cannot tell why a login failed, and it
should not have to.
What stays:
1. **Remedy 1** — the forge in the ssh configuration the mesh writes, so no machine of the mesh
presents the wrong user. This is the fix.
2. **A broken server must not look like bad clients.** When a service refuses *every* login — the
mail server after its database was emptied — its own health check should fail and say so before
the jail has banned its users. That belongs to the service's module, not to the jail.
## Status
Not located further: both remedies are design choices — (1) a served port the controller translates
by listen, (2) the open question 1 of the report.
Not located further: remedy 1 needs a served port the controller translates by listen.
@@ -3,7 +3,7 @@ status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
---
# 242 — The mesh has no backups
@@ -19,6 +19,13 @@ What survived did so by accident: git history because every repository is also c
operator's machines; the file-sync service's files because they live in the object store, which
nothing dropped; the mailboxes because they are files outside the database.
## Scope, set by the operator (2026-10-05)
**Mistakes, not disasters.** Backups protect against a dropped database, a deleted bucket, a bad
migration — not against a dead disk or a lost site; that loss is accepted. So no off-site copy is
needed, and question 4's "spread across the mesh" and "encrypted to whom" fall away. Worked out in
research 030; proposed as ADR 0214.
## Questions this must answer
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's