diff --git a/01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md b/01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md new file mode 100644 index 0000000..9a2a9a6 --- /dev/null +++ b/01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md @@ -0,0 +1,28 @@ +--- +status: active +initiated: 2026-10-05 +touches: [04-ISSUES/242-the-mesh-has-no-backups, 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md, 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md, 03-DESIGN/01-to-be/31-a-module-declares-its-fail2ban-jail.md] +--- + +# 030 — Backups against our own mistakes + +**What.** How the mesh keeps restore points of its data: what is copied, how often, how long it is +kept, full or incremental, where it lives, who runs it and how it is proven to restore. + +**Why.** Issue 242: nothing in the mesh backs anything up, and when one misread file dropped every +database on the control node (issue 241), the newest copies were migration leftovers nine to twelve +days old, found by searching a disk. + +**The scope is set by the operator, and it is narrow on purpose: mistakes, not disasters.** A +backup here protects against what a person, an agent or the mesh itself does wrong — a dropped +database, a deleted bucket, a bad migration, a file overwritten — and not against a disk dying or a +building burning. Losing data to a disaster is an accepted risk. That removes the off-site copy, the +cross-site transfer and the second key-holder from the problem, and leaves the part that would have +saved the night of issue 241. + +**What it touches.** The store providers (each knows how to dump its own store consistently), the +host's scheduled steps (ADR 0053, built), the node-wide composition pattern a module's jail already +uses (to-be 31), and the operator's output channel (research 028) for saying a backup failed. + +Documents: [01 — what the mesh holds](01-what-the-mesh-holds.md), +[02 — options and a proposal](02-options-and-proposal.md). diff --git a/01-RESEARCH/030-backups-against-our-own-mistakes/01-what-the-mesh-holds.md b/01-RESEARCH/030-backups-against-our-own-mistakes/01-what-the-mesh-holds.md new file mode 100644 index 0000000..2c787e1 --- /dev/null +++ b/01-RESEARCH/030-backups-against-our-own-mistakes/01-what-the-mesh-holds.md @@ -0,0 +1,55 @@ +# 01 — What the mesh holds, measured 2026-10-05 + +Four machines: the control node (hosted, holds every public service), the home server (media, home +automation, a large ZFS pool), a workstation and a laptop. Sizes are apparent sizes, rounded. + +## Nothing backs anything up + +On every machine: no backup tool other than `rsync` and `pg_dump` is installed, no systemd timer and +no cron line mentions a backup, dump or snapshot. Every live data directory is on ext4 except the +home server's pool (ZFS), so a filesystem snapshot is available only there. + +## The control node — the data that cannot be recreated + +| what | size | how it changes | +|---|---|---| +| object store (file-sync service's files, photos) | 183 GB | slowly; files added, rarely rewritten | +| forge (repositories, attachments, its database) | 7 GB | daily | +| MS SQL Server databases | 5 GB | daily | +| mail (mailboxes; accounts in postgres) | 2 GB | continuously | +| postgres (forge, mail admin, identity, file-sync index, analytics, catalogue, licence manager) | ~2 GB | continuously | +| file-sync service's own directory, website, analytics | ~4 GB | slowly | +| MongoDB | 0.4 GB | daily | +| the mesh's own records (controller, vault, module state) | ~1.5 GB | continuously | + +Recreatable and not worth copying: container images (210 GB), the artifact registry (40 GB — every +artifact is rebuilt from git), a 115 GB speed-test bucket and a 7 GB pre-migration object-store copy. + +Free space: 833 GB on the filesystem holding the data, 2.9 TB on a second one. + +## The home server + +MS SQL Server 80 GB, postgres and a self-hosted backend platform ~3 GB, chat server 6 GB, home +automation, network controller and time-series data each under 2 GB, and the media services' +libraries (tens of GB, mostly cover art and metadata they re-fetch). The 89 TB media library is +replaceable by its nature and out of scope. Free: 31 TB on the pool, 453 GB on the system disk. + +## The workstation and the laptop + +The workstation has 142 GB under its services directory and 31 GB of container volumes; the laptop +3 GB. Mostly development; what among it is data nobody can regenerate is for each module to say. + +## Between the sites + +Control node to home server ~285 Mbit/s, home server to control node ~19 Mbit/s. Irrelevant now that +backups stay on the machine whose data they hold, recorded because it is why an off-site copy would +have been expensive. + +## What issue 241 says about the requirement + +- The mistake was noticed within hours. A restore point a day old would have lost a day. +- The restore had to go *beside* the live database, not over it, and that worked well. +- A copy of a live postgres data directory needed a throwaway server of the right version to read; + a logical dump would have restored directly. +- The data that survived was the data outside the dropped stores. A backup that lives inside the + store it protects — a database's own snapshot table, a bucket's own versions — dies with a drop. diff --git a/01-RESEARCH/030-backups-against-our-own-mistakes/02-options-and-proposal.md b/01-RESEARCH/030-backups-against-our-own-mistakes/02-options-and-proposal.md new file mode 100644 index 0000000..0a40a01 --- /dev/null +++ b/01-RESEARCH/030-backups-against-our-own-mistakes/02-options-and-proposal.md @@ -0,0 +1,74 @@ +# 02 — Options and a proposal + +## Who decides what is backed up + +1. **A central list** on the backup holder. Rejected: it is the attentiveness rule ADR 0030 + rejected — a store added and not listed is silently unprotected. +2. **Each module declares its own data, a node-wide holder composes them.** The pattern of to-be 31 + (a module declares its jail; the mesh composes them per node). A store provider declares *how* + to take a consistent copy (a dump command), a module with plain files declares *which* paths. The + holder composes every declaration on the node into one schedule. **Proposed.** + +The data a module keeps in a database it gets from a provider is backed up by the provider, which +dumps every database it serves — so a consumer declares nothing, and a new consumer is covered the +day it is provisioned. + +## Full or incremental + +- **Databases: a full logical dump every time** (`pg_dump -Fc`, MS SQL `BACKUP DATABASE`, + `mongodump`). A dump restores with the store's own tool into a database beside the live one — + issue 241's recovery without the throwaway server — and is consistent, which a copy of a live data + directory is not. +- **Everything goes into one deduplicating repository per node** (restic or borg). Each night is a + complete restore point, yet only changed chunks cost space: the object store's 183 GB is copied + once, then each night adds what changed. This removes the full-vs-delta trade-off rather than + choosing a side. + +Considered for the object store alone: the object store's own versioning with a lifecycle rule. +Rejected as the only copy — it lives inside the store, and a removed bucket or data directory takes +its versions with it. + +## Where + +On the same machine, outside every data directory the mesh manages, on a second filesystem where the +machine has one (the control node does). Not off-site: the scope is mistakes. The repository is +encrypted anyway (both tools require it); its key is a mesh secret (ADR 0085), so a person can +restore without the holder. + +## How often, how long + +- **Nightly**, at a quiet hour, as a scheduled step (ADR 0053). +- **On demand before a risky act** — a migration, a retirement, an operator's experiment — through a + verb; the act's own tooling can call it. +- **Kept: 14 daily, 8 weekly, 6 monthly.** A mistake is usually noticed within days, sometimes weeks + (a deleted file nobody opens). Six months bounds the space a slowly-noticed mistake needs. + +Estimated cost on the control node: ~200 GB for the first night, a few GB a night after; well within +the second filesystem's 2.9 TB. + +## Who runs it + +A node seat, **`node-backup`**, held on every machine that has data by one module (named for the tool +it wraps). It receives the declarations, runs them, keeps the repository, and offers the verbs a +person needs: + +- what is backed up here, and the last good night of each; +- take a backup now; +- restore one item **beside** the live one — a database to `_restore`, a path to + `.restored-` — never over it. Swapping it in stays a person's act, as in issue 241. + +## How it is proven + +- Every run checks its own result; a failed or skipped night goes to the operator's output channel + (research 028), not only a log. +- Weekly: the repository's integrity check, and one database restored from the newest dump into a + throwaway instance and counted against the live one. +- A machine with data and no successful backup in 48 hours is a problem the mesh's status shows. + +## Open questions + +- restic or borg — both fit; restic is a single binary with no server, which suits a module. +- Whether the workstation and laptop take part at all, or only once a module there declares data. +- The mail spool is files and the forge has a dump command of its own; whether the forge's + repositories are worth backing up at all when every clone is a copy (issue 241 says the forge's + *database* is the part with no other copy). diff --git a/02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md b/02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md new file mode 100644 index 0000000..68d9ca1 --- /dev/null +++ b/02-DECISIONS/0214-backups-guard-against-mistakes-and-stay-on-the-machine.md @@ -0,0 +1,66 @@ +--- +topic: what runs on it +status: accepted +date: 2026-10-05 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md +--- + +# 214. Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data + +## Context + +Nothing in the mesh backed anything up (issue 242). When a misread file dropped every database on the +control node (issue 241), recovery took a night and used copies nine to twelve days old. + +The operator sets the scope: backups exist for **mistakes** — a person's, an agent's, the mesh's own +— not for disasters. Losing data to a dead disk or a lost site is accepted. Research 030 measured +what the machines hold and found no backup tooling anywhere. + +## Considered Options + +1. **A central list of what to back up.** Rejected — whatever is not listed is unprotected, silently. +2. **Each store's own mechanisms** (bucket versioning, database snapshots). Rejected as the only copy: + they live inside what they protect, and a drop takes them with it. +3. **Off-site copies.** Out of scope by the operator's decision; recorded so the absence is a choice. +4. **Modules declare, a node seat composes, the copy stays on the machine.** Adopted. + +## Decision + +**A module declares the data it owns; a node seat, `node-backup`, composes every declaration on the +machine and keeps nightly restore points there.** A store provider declares how to dump each database +it serves, so a consumer of a store declares nothing. A module with files declares their paths. + +**Databases are dumped in full, logically, every night; everything lands in one encrypted, +deduplicating repository per machine,** so every night is a complete restore point and only what +changed costs space. + +**Kept: 14 daily, 8 weekly, 6 monthly.** A backup is also taken on demand before a risky act. + +**The repository is on the machine, outside every directory the mesh manages,** on a second +filesystem where there is one. Its key is a mesh secret. + +**A restore goes beside the live data, never over it.** Swapping it in is a person's act. + +**A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours +shows in the mesh's status. + +## Consequences + +- Adding a store provider means declaring its dump; the catalogue check can refuse a store provider + that declares none. +- The control node's first backup is ~200 GB, then a few GB a night. +- A dead disk or a lost machine still loses its data, by choice. + +## How it is checked + +The catalogue check refuses a module providing a store seat without a backup declaration. The holder's +weekly restore test restores one dump into a throwaway instance and compares counts. The mesh's +status lists every machine whose last good backup is older than 48 hours. + +## References + +- [04-ISSUES/242](../04-ISSUES/242-the-mesh-has-no-backups/00-report.md), [04-ISSUES/241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md) +- [01-RESEARCH/030](../01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md) +- ADR 0030 (data outlives its declaration), ADR 0053 (scheduled steps), ADR 0085 (a secret is a provision) diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 0ebf373..39eb267 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -313,6 +313,7 @@ python3 00-META/checks/index.py fail if stale - **0209** — [A login on a node moves that node to the account it logged in to; an API key is added from any node, sealed](0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md) - **0211** — [A machine's power is a node seat, its moments take contributions, and its states are events](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md) - **0213** — [The operator sets the agent's managed settings through the agent module, under the mesh's own keys](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md) +- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md) ### How it is built diff --git a/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md b/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md index 285ecd4..88445f5 100644 --- a/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md +++ b/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md @@ -37,7 +37,22 @@ mesh knows puts it there. or an operator setting naming them. The first is derived and stays true when the uplink changes; the second is a value somebody must remember to edit. +## The operator's call (2026-10-05): fix the logins, not the jail + +Remedy 2 is rejected. **A machine of the mesh should not fail logins, and when it does, the ban is +the jail working.** Exempting the operator's address would hide exactly the failures worth seeing — +here a misconfigured ssh client, and in the same week a mail server that had lost every account +(issue 241). Both bans were genuinely failed logins; fail2ban cannot tell why a login failed, and it +should not have to. + +What stays: + +1. **Remedy 1** — the forge in the ssh configuration the mesh writes, so no machine of the mesh + presents the wrong user. This is the fix. +2. **A broken server must not look like bad clients.** When a service refuses *every* login — the + mail server after its database was emptied — its own health check should fail and say so before + the jail has banned its users. That belongs to the service's module, not to the jail. + ## Status -Not located further: both remedies are design choices — (1) a served port the controller translates -by listen, (2) the open question 1 of the report. +Not located further: remedy 1 needs a served port the controller translates by listen. diff --git a/04-ISSUES/242-the-mesh-has-no-backups/00-report.md b/04-ISSUES/242-the-mesh-has-no-backups/00-report.md index 6fdc07b..247096f 100644 --- a/04-ISSUES/242-the-mesh-has-no-backups/00-report.md +++ b/04-ISSUES/242-the-mesh-has-no-backups/00-report.md @@ -19,6 +19,13 @@ What survived did so by accident: git history because every repository is also c operator's machines; the file-sync service's files because they live in the object store, which nothing dropped; the mailboxes because they are files outside the database. +## Scope, set by the operator (2026-10-05) + +**Mistakes, not disasters.** Backups protect against a dropped database, a deleted bucket, a bad +migration — not against a dead disk or a lost site; that loss is accepted. So no off-site copy is +needed, and question 4's "spread across the mesh" and "encrypted to whom" fall away. Worked out in +research 030; proposed as ADR 0214. + ## Questions this must answer 1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's