Research 030 and proposed ADR 0214: backups guard against mistakes and stay on the machine

The operator scoped backups to mistakes, not disasters (issue 242). Issue 238: fix the failed logins
at their source instead of exempting the operator's address.
This commit is contained in:
2026-10-05 10:07:36 +02:00
parent 2a501af0eb
commit 3d0ce3e6dd
7 changed files with 248 additions and 2 deletions
@@ -0,0 +1,28 @@
---
status: active
initiated: 2026-10-05
touches: [04-ISSUES/242-the-mesh-has-no-backups, 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md, 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md, 03-DESIGN/01-to-be/31-a-module-declares-its-fail2ban-jail.md]
---
# 030 — Backups against our own mistakes
**What.** How the mesh keeps restore points of its data: what is copied, how often, how long it is
kept, full or incremental, where it lives, who runs it and how it is proven to restore.
**Why.** Issue 242: nothing in the mesh backs anything up, and when one misread file dropped every
database on the control node (issue 241), the newest copies were migration leftovers nine to twelve
days old, found by searching a disk.
**The scope is set by the operator, and it is narrow on purpose: mistakes, not disasters.** A
backup here protects against what a person, an agent or the mesh itself does wrong — a dropped
database, a deleted bucket, a bad migration, a file overwritten — and not against a disk dying or a
building burning. Losing data to a disaster is an accepted risk. That removes the off-site copy, the
cross-site transfer and the second key-holder from the problem, and leaves the part that would have
saved the night of issue 241.
**What it touches.** The store providers (each knows how to dump its own store consistently), the
host's scheduled steps (ADR 0053, built), the node-wide composition pattern a module's jail already
uses (to-be 31), and the operator's output channel (research 028) for saying a backup failed.
Documents: [01 — what the mesh holds](01-what-the-mesh-holds.md),
[02 — options and a proposal](02-options-and-proposal.md).
@@ -0,0 +1,55 @@
# 01 — What the mesh holds, measured 2026-10-05
Four machines: the control node (hosted, holds every public service), the home server (media, home
automation, a large ZFS pool), a workstation and a laptop. Sizes are apparent sizes, rounded.
## Nothing backs anything up
On every machine: no backup tool other than `rsync` and `pg_dump` is installed, no systemd timer and
no cron line mentions a backup, dump or snapshot. Every live data directory is on ext4 except the
home server's pool (ZFS), so a filesystem snapshot is available only there.
## The control node — the data that cannot be recreated
| what | size | how it changes |
|---|---|---|
| object store (file-sync service's files, photos) | 183 GB | slowly; files added, rarely rewritten |
| forge (repositories, attachments, its database) | 7 GB | daily |
| MS SQL Server databases | 5 GB | daily |
| mail (mailboxes; accounts in postgres) | 2 GB | continuously |
| postgres (forge, mail admin, identity, file-sync index, analytics, catalogue, licence manager) | ~2 GB | continuously |
| file-sync service's own directory, website, analytics | ~4 GB | slowly |
| MongoDB | 0.4 GB | daily |
| the mesh's own records (controller, vault, module state) | ~1.5 GB | continuously |
Recreatable and not worth copying: container images (210 GB), the artifact registry (40 GB — every
artifact is rebuilt from git), a 115 GB speed-test bucket and a 7 GB pre-migration object-store copy.
Free space: 833 GB on the filesystem holding the data, 2.9 TB on a second one.
## The home server
MS SQL Server 80 GB, postgres and a self-hosted backend platform ~3 GB, chat server 6 GB, home
automation, network controller and time-series data each under 2 GB, and the media services'
libraries (tens of GB, mostly cover art and metadata they re-fetch). The 89 TB media library is
replaceable by its nature and out of scope. Free: 31 TB on the pool, 453 GB on the system disk.
## The workstation and the laptop
The workstation has 142 GB under its services directory and 31 GB of container volumes; the laptop
3 GB. Mostly development; what among it is data nobody can regenerate is for each module to say.
## Between the sites
Control node to home server ~285 Mbit/s, home server to control node ~19 Mbit/s. Irrelevant now that
backups stay on the machine whose data they hold, recorded because it is why an off-site copy would
have been expensive.
## What issue 241 says about the requirement
- The mistake was noticed within hours. A restore point a day old would have lost a day.
- The restore had to go *beside* the live database, not over it, and that worked well.
- A copy of a live postgres data directory needed a throwaway server of the right version to read;
a logical dump would have restored directly.
- The data that survived was the data outside the dropped stores. A backup that lives inside the
store it protects — a database's own snapshot table, a bucket's own versions — dies with a drop.
@@ -0,0 +1,74 @@
# 02 — Options and a proposal
## Who decides what is backed up
1. **A central list** on the backup holder. Rejected: it is the attentiveness rule ADR 0030
rejected — a store added and not listed is silently unprotected.
2. **Each module declares its own data, a node-wide holder composes them.** The pattern of to-be 31
(a module declares its jail; the mesh composes them per node). A store provider declares *how*
to take a consistent copy (a dump command), a module with plain files declares *which* paths. The
holder composes every declaration on the node into one schedule. **Proposed.**
The data a module keeps in a database it gets from a provider is backed up by the provider, which
dumps every database it serves — so a consumer declares nothing, and a new consumer is covered the
day it is provisioned.
## Full or incremental
- **Databases: a full logical dump every time** (`pg_dump -Fc`, MS SQL `BACKUP DATABASE`,
`mongodump`). A dump restores with the store's own tool into a database beside the live one —
issue 241's recovery without the throwaway server — and is consistent, which a copy of a live data
directory is not.
- **Everything goes into one deduplicating repository per node** (restic or borg). Each night is a
complete restore point, yet only changed chunks cost space: the object store's 183 GB is copied
once, then each night adds what changed. This removes the full-vs-delta trade-off rather than
choosing a side.
Considered for the object store alone: the object store's own versioning with a lifecycle rule.
Rejected as the only copy — it lives inside the store, and a removed bucket or data directory takes
its versions with it.
## Where
On the same machine, outside every data directory the mesh manages, on a second filesystem where the
machine has one (the control node does). Not off-site: the scope is mistakes. The repository is
encrypted anyway (both tools require it); its key is a mesh secret (ADR 0085), so a person can
restore without the holder.
## How often, how long
- **Nightly**, at a quiet hour, as a scheduled step (ADR 0053).
- **On demand before a risky act** — a migration, a retirement, an operator's experiment — through a
verb; the act's own tooling can call it.
- **Kept: 14 daily, 8 weekly, 6 monthly.** A mistake is usually noticed within days, sometimes weeks
(a deleted file nobody opens). Six months bounds the space a slowly-noticed mistake needs.
Estimated cost on the control node: ~200 GB for the first night, a few GB a night after; well within
the second filesystem's 2.9 TB.
## Who runs it
A node seat, **`node-backup`**, held on every machine that has data by one module (named for the tool
it wraps). It receives the declarations, runs them, keeps the repository, and offers the verbs a
person needs:
- what is backed up here, and the last good night of each;
- take a backup now;
- restore one item **beside** the live one — a database to `<name>_restore`, a path to
`<path>.restored-<date>` — never over it. Swapping it in stays a person's act, as in issue 241.
## How it is proven
- Every run checks its own result; a failed or skipped night goes to the operator's output channel
(research 028), not only a log.
- Weekly: the repository's integrity check, and one database restored from the newest dump into a
throwaway instance and counted against the live one.
- A machine with data and no successful backup in 48 hours is a problem the mesh's status shows.
## Open questions
- restic or borg — both fit; restic is a single binary with no server, which suits a module.
- Whether the workstation and laptop take part at all, or only once a module there declares data.
- The mail spool is files and the forge has a dump command of its own; whether the forge's
repositories are worth backing up at all when every clone is a copy (issue 241 says the forge's
*database* is the part with no other copy).
@@ -0,0 +1,66 @@
---
topic: what runs on it
status: proposed
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md
---
# 214. Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data
## Context
Nothing in the mesh backed anything up (issue 242). When a misread file dropped every database on the
control node (issue 241), recovery took a night and used copies nine to twelve days old.
The operator sets the scope: backups exist for **mistakes** — a person's, an agent's, the mesh's own
— not for disasters. Losing data to a dead disk or a lost site is accepted. Research 030 measured
what the machines hold and found no backup tooling anywhere.
## Considered Options
1. **A central list of what to back up.** Rejected — whatever is not listed is unprotected, silently.
2. **Each store's own mechanisms** (bucket versioning, database snapshots). Rejected as the only copy:
they live inside what they protect, and a drop takes them with it.
3. **Off-site copies.** Out of scope by the operator's decision; recorded so the absence is a choice.
4. **Modules declare, a node seat composes, the copy stays on the machine.** Adopted.
## Decision
**A module declares the data it owns; a node seat, `node-backup`, composes every declaration on the
machine and keeps nightly restore points there.** A store provider declares how to dump each database
it serves, so a consumer of a store declares nothing. A module with files declares their paths.
**Databases are dumped in full, logically, every night; everything lands in one encrypted,
deduplicating repository per machine,** so every night is a complete restore point and only what
changed costs space.
**Kept: 14 daily, 8 weekly, 6 monthly.** A backup is also taken on demand before a risky act.
**The repository is on the machine, outside every directory the mesh manages,** on a second
filesystem where there is one. Its key is a mesh secret.
**A restore goes beside the live data, never over it.** Swapping it in is a person's act.
**A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours
shows in the mesh's status.
## Consequences
- Adding a store provider means declaring its dump; the catalogue check can refuse a store provider
that declares none.
- The control node's first backup is ~200 GB, then a few GB a night.
- A dead disk or a lost machine still loses its data, by choice.
## How it is checked
The catalogue check refuses a module providing a store seat without a backup declaration. The holder's
weekly restore test restores one dump into a throwaway instance and compares counts. The mesh's
status lists every machine whose last good backup is older than 48 hours.
## References
- [04-ISSUES/242](../04-ISSUES/242-the-mesh-has-no-backups/00-report.md), [04-ISSUES/241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
- [01-RESEARCH/030](../01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md)
- ADR 0030 (data outlives its declaration), ADR 0053 (scheduled steps), ADR 0085 (a secret is a provision)
+1
View File
@@ -313,6 +313,7 @@ python3 00-META/checks/index.py fail if stale
- **0209** — [A login on a node moves that node to the account it logged in to; an API key is added from any node, sealed](0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md)
- **0211** — [A machine's power is a node seat, its moments take contributions, and its states are events](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)
- **0213** — [The operator sets the agent's managed settings through the agent module, under the mesh's own keys](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md) *(proposed)*
### How it is built
@@ -37,7 +37,22 @@ mesh knows puts it there.
or an operator setting naming them. The first is derived and stays true when the uplink changes;
the second is a value somebody must remember to edit.
## The operator's call (2026-10-05): fix the logins, not the jail
Remedy 2 is rejected. **A machine of the mesh should not fail logins, and when it does, the ban is
the jail working.** Exempting the operator's address would hide exactly the failures worth seeing —
here a misconfigured ssh client, and in the same week a mail server that had lost every account
(issue 241). Both bans were genuinely failed logins; fail2ban cannot tell why a login failed, and it
should not have to.
What stays:
1. **Remedy 1** — the forge in the ssh configuration the mesh writes, so no machine of the mesh
presents the wrong user. This is the fix.
2. **A broken server must not look like bad clients.** When a service refuses *every* login — the
mail server after its database was emptied — its own health check should fail and say so before
the jail has banned its users. That belongs to the service's module, not to the jail.
## Status
Not located further: both remedies are design choices — (1) a served port the controller translates
by listen, (2) the open question 1 of the report.
Not located further: remedy 1 needs a served port the controller translates by listen.
@@ -19,6 +19,13 @@ What survived did so by accident: git history because every repository is also c
operator's machines; the file-sync service's files because they live in the object store, which
nothing dropped; the mailboxes because they are files outside the database.
## Scope, set by the operator (2026-10-05)
**Mistakes, not disasters.** Backups protect against a dropped database, a deleted bucket, a bad
migration — not against a dead disk or a lost site; that loss is accepted. So no off-site copy is
needed, and questions 4's "spread across the mesh" and "encrypted to whom" fall away. Worked out in
research 030; proposed as ADR 0214.
## Questions this must answer
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's