Research 030 and proposed ADR 0214: backups guard against mistakes and stay on the machine
The operator scoped backups to mistakes, not disasters (issue 242). Issue 238: fix the failed logins at their source instead of exempting the operator's address.
This commit is contained in:
@@ -0,0 +1,28 @@
|
||||
---
|
||||
status: active
|
||||
initiated: 2026-10-05
|
||||
touches: [04-ISSUES/242-the-mesh-has-no-backups, 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md, 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md, 03-DESIGN/01-to-be/31-a-module-declares-its-fail2ban-jail.md]
|
||||
---
|
||||
|
||||
# 030 — Backups against our own mistakes
|
||||
|
||||
**What.** How the mesh keeps restore points of its data: what is copied, how often, how long it is
|
||||
kept, full or incremental, where it lives, who runs it and how it is proven to restore.
|
||||
|
||||
**Why.** Issue 242: nothing in the mesh backs anything up, and when one misread file dropped every
|
||||
database on the control node (issue 241), the newest copies were migration leftovers nine to twelve
|
||||
days old, found by searching a disk.
|
||||
|
||||
**The scope is set by the operator, and it is narrow on purpose: mistakes, not disasters.** A
|
||||
backup here protects against what a person, an agent or the mesh itself does wrong — a dropped
|
||||
database, a deleted bucket, a bad migration, a file overwritten — and not against a disk dying or a
|
||||
building burning. Losing data to a disaster is an accepted risk. That removes the off-site copy, the
|
||||
cross-site transfer and the second key-holder from the problem, and leaves the part that would have
|
||||
saved the night of issue 241.
|
||||
|
||||
**What it touches.** The store providers (each knows how to dump its own store consistently), the
|
||||
host's scheduled steps (ADR 0053, built), the node-wide composition pattern a module's jail already
|
||||
uses (to-be 31), and the operator's output channel (research 028) for saying a backup failed.
|
||||
|
||||
Documents: [01 — what the mesh holds](01-what-the-mesh-holds.md),
|
||||
[02 — options and a proposal](02-options-and-proposal.md).
|
||||
@@ -0,0 +1,55 @@
|
||||
# 01 — What the mesh holds, measured 2026-10-05
|
||||
|
||||
Four machines: the control node (hosted, holds every public service), the home server (media, home
|
||||
automation, a large ZFS pool), a workstation and a laptop. Sizes are apparent sizes, rounded.
|
||||
|
||||
## Nothing backs anything up
|
||||
|
||||
On every machine: no backup tool other than `rsync` and `pg_dump` is installed, no systemd timer and
|
||||
no cron line mentions a backup, dump or snapshot. Every live data directory is on ext4 except the
|
||||
home server's pool (ZFS), so a filesystem snapshot is available only there.
|
||||
|
||||
## The control node — the data that cannot be recreated
|
||||
|
||||
| what | size | how it changes |
|
||||
|---|---|---|
|
||||
| object store (file-sync service's files, photos) | 183 GB | slowly; files added, rarely rewritten |
|
||||
| forge (repositories, attachments, its database) | 7 GB | daily |
|
||||
| MS SQL Server databases | 5 GB | daily |
|
||||
| mail (mailboxes; accounts in postgres) | 2 GB | continuously |
|
||||
| postgres (forge, mail admin, identity, file-sync index, analytics, catalogue, licence manager) | ~2 GB | continuously |
|
||||
| file-sync service's own directory, website, analytics | ~4 GB | slowly |
|
||||
| MongoDB | 0.4 GB | daily |
|
||||
| the mesh's own records (controller, vault, module state) | ~1.5 GB | continuously |
|
||||
|
||||
Recreatable and not worth copying: container images (210 GB), the artifact registry (40 GB — every
|
||||
artifact is rebuilt from git), a 115 GB speed-test bucket and a 7 GB pre-migration object-store copy.
|
||||
|
||||
Free space: 833 GB on the filesystem holding the data, 2.9 TB on a second one.
|
||||
|
||||
## The home server
|
||||
|
||||
MS SQL Server 80 GB, postgres and a self-hosted backend platform ~3 GB, chat server 6 GB, home
|
||||
automation, network controller and time-series data each under 2 GB, and the media services'
|
||||
libraries (tens of GB, mostly cover art and metadata they re-fetch). The 89 TB media library is
|
||||
replaceable by its nature and out of scope. Free: 31 TB on the pool, 453 GB on the system disk.
|
||||
|
||||
## The workstation and the laptop
|
||||
|
||||
The workstation has 142 GB under its services directory and 31 GB of container volumes; the laptop
|
||||
3 GB. Mostly development; what among it is data nobody can regenerate is for each module to say.
|
||||
|
||||
## Between the sites
|
||||
|
||||
Control node to home server ~285 Mbit/s, home server to control node ~19 Mbit/s. Irrelevant now that
|
||||
backups stay on the machine whose data they hold, recorded because it is why an off-site copy would
|
||||
have been expensive.
|
||||
|
||||
## What issue 241 says about the requirement
|
||||
|
||||
- The mistake was noticed within hours. A restore point a day old would have lost a day.
|
||||
- The restore had to go *beside* the live database, not over it, and that worked well.
|
||||
- A copy of a live postgres data directory needed a throwaway server of the right version to read;
|
||||
a logical dump would have restored directly.
|
||||
- The data that survived was the data outside the dropped stores. A backup that lives inside the
|
||||
store it protects — a database's own snapshot table, a bucket's own versions — dies with a drop.
|
||||
@@ -0,0 +1,74 @@
|
||||
# 02 — Options and a proposal
|
||||
|
||||
## Who decides what is backed up
|
||||
|
||||
1. **A central list** on the backup holder. Rejected: it is the attentiveness rule ADR 0030
|
||||
rejected — a store added and not listed is silently unprotected.
|
||||
2. **Each module declares its own data, a node-wide holder composes them.** The pattern of to-be 31
|
||||
(a module declares its jail; the mesh composes them per node). A store provider declares *how*
|
||||
to take a consistent copy (a dump command), a module with plain files declares *which* paths. The
|
||||
holder composes every declaration on the node into one schedule. **Proposed.**
|
||||
|
||||
The data a module keeps in a database it gets from a provider is backed up by the provider, which
|
||||
dumps every database it serves — so a consumer declares nothing, and a new consumer is covered the
|
||||
day it is provisioned.
|
||||
|
||||
## Full or incremental
|
||||
|
||||
- **Databases: a full logical dump every time** (`pg_dump -Fc`, MS SQL `BACKUP DATABASE`,
|
||||
`mongodump`). A dump restores with the store's own tool into a database beside the live one —
|
||||
issue 241's recovery without the throwaway server — and is consistent, which a copy of a live data
|
||||
directory is not.
|
||||
- **Everything goes into one deduplicating repository per node** (restic or borg). Each night is a
|
||||
complete restore point, yet only changed chunks cost space: the object store's 183 GB is copied
|
||||
once, then each night adds what changed. This removes the full-vs-delta trade-off rather than
|
||||
choosing a side.
|
||||
|
||||
Considered for the object store alone: the object store's own versioning with a lifecycle rule.
|
||||
Rejected as the only copy — it lives inside the store, and a removed bucket or data directory takes
|
||||
its versions with it.
|
||||
|
||||
## Where
|
||||
|
||||
On the same machine, outside every data directory the mesh manages, on a second filesystem where the
|
||||
machine has one (the control node does). Not off-site: the scope is mistakes. The repository is
|
||||
encrypted anyway (both tools require it); its key is a mesh secret (ADR 0085), so a person can
|
||||
restore without the holder.
|
||||
|
||||
## How often, how long
|
||||
|
||||
- **Nightly**, at a quiet hour, as a scheduled step (ADR 0053).
|
||||
- **On demand before a risky act** — a migration, a retirement, an operator's experiment — through a
|
||||
verb; the act's own tooling can call it.
|
||||
- **Kept: 14 daily, 8 weekly, 6 monthly.** A mistake is usually noticed within days, sometimes weeks
|
||||
(a deleted file nobody opens). Six months bounds the space a slowly-noticed mistake needs.
|
||||
|
||||
Estimated cost on the control node: ~200 GB for the first night, a few GB a night after; well within
|
||||
the second filesystem's 2.9 TB.
|
||||
|
||||
## Who runs it
|
||||
|
||||
A node seat, **`node-backup`**, held on every machine that has data by one module (named for the tool
|
||||
it wraps). It receives the declarations, runs them, keeps the repository, and offers the verbs a
|
||||
person needs:
|
||||
|
||||
- what is backed up here, and the last good night of each;
|
||||
- take a backup now;
|
||||
- restore one item **beside** the live one — a database to `<name>_restore`, a path to
|
||||
`<path>.restored-<date>` — never over it. Swapping it in stays a person's act, as in issue 241.
|
||||
|
||||
## How it is proven
|
||||
|
||||
- Every run checks its own result; a failed or skipped night goes to the operator's output channel
|
||||
(research 028), not only a log.
|
||||
- Weekly: the repository's integrity check, and one database restored from the newest dump into a
|
||||
throwaway instance and counted against the live one.
|
||||
- A machine with data and no successful backup in 48 hours is a problem the mesh's status shows.
|
||||
|
||||
## Open questions
|
||||
|
||||
- restic or borg — both fit; restic is a single binary with no server, which suits a module.
|
||||
- Whether the workstation and laptop take part at all, or only once a module there declares data.
|
||||
- The mail spool is files and the forge has a dump command of its own; whether the forge's
|
||||
repositories are worth backing up at all when every clone is a copy (issue 241 says the forge's
|
||||
*database* is the part with no other copy).
|
||||
@@ -0,0 +1,66 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: proposed
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md
|
||||
---
|
||||
|
||||
# 214. Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data
|
||||
|
||||
## Context
|
||||
|
||||
Nothing in the mesh backed anything up (issue 242). When a misread file dropped every database on the
|
||||
control node (issue 241), recovery took a night and used copies nine to twelve days old.
|
||||
|
||||
The operator sets the scope: backups exist for **mistakes** — a person's, an agent's, the mesh's own
|
||||
— not for disasters. Losing data to a dead disk or a lost site is accepted. Research 030 measured
|
||||
what the machines hold and found no backup tooling anywhere.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **A central list of what to back up.** Rejected — whatever is not listed is unprotected, silently.
|
||||
2. **Each store's own mechanisms** (bucket versioning, database snapshots). Rejected as the only copy:
|
||||
they live inside what they protect, and a drop takes them with it.
|
||||
3. **Off-site copies.** Out of scope by the operator's decision; recorded so the absence is a choice.
|
||||
4. **Modules declare, a node seat composes, the copy stays on the machine.** Adopted.
|
||||
|
||||
## Decision
|
||||
|
||||
**A module declares the data it owns; a node seat, `node-backup`, composes every declaration on the
|
||||
machine and keeps nightly restore points there.** A store provider declares how to dump each database
|
||||
it serves, so a consumer of a store declares nothing. A module with files declares their paths.
|
||||
|
||||
**Databases are dumped in full, logically, every night; everything lands in one encrypted,
|
||||
deduplicating repository per machine,** so every night is a complete restore point and only what
|
||||
changed costs space.
|
||||
|
||||
**Kept: 14 daily, 8 weekly, 6 monthly.** A backup is also taken on demand before a risky act.
|
||||
|
||||
**The repository is on the machine, outside every directory the mesh manages,** on a second
|
||||
filesystem where there is one. Its key is a mesh secret.
|
||||
|
||||
**A restore goes beside the live data, never over it.** Swapping it in is a person's act.
|
||||
|
||||
**A night that fails reaches the operator,** and a machine with data and no good backup in 48 hours
|
||||
shows in the mesh's status.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Adding a store provider means declaring its dump; the catalogue check can refuse a store provider
|
||||
that declares none.
|
||||
- The control node's first backup is ~200 GB, then a few GB a night.
|
||||
- A dead disk or a lost machine still loses its data, by choice.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The catalogue check refuses a module providing a store seat without a backup declaration. The holder's
|
||||
weekly restore test restores one dump into a throwaway instance and compares counts. The mesh's
|
||||
status lists every machine whose last good backup is older than 48 hours.
|
||||
|
||||
## References
|
||||
|
||||
- [04-ISSUES/242](../04-ISSUES/242-the-mesh-has-no-backups/00-report.md), [04-ISSUES/241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
|
||||
- [01-RESEARCH/030](../01-RESEARCH/030-backups-against-our-own-mistakes/00-overview.md)
|
||||
- ADR 0030 (data outlives its declaration), ADR 0053 (scheduled steps), ADR 0085 (a secret is a provision)
|
||||
@@ -313,6 +313,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0209** — [A login on a node moves that node to the account it logged in to; an API key is added from any node, sealed](0209-a-login-on-a-node-moves-that-node-to-its-account-and-an-api-key-is-added-from-any-node-sealed.md)
|
||||
- **0211** — [A machine's power is a node seat, its moments take contributions, and its states are events](0211-a-machines-power-is-a-node-seat-its-moments-take-contributions-and-its-states-are-events.md)
|
||||
- **0213** — [The operator sets the agent's managed settings through the agent module, under the mesh's own keys](0213-the-operator-sets-the-agents-managed-settings-through-the-agent-module.md)
|
||||
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md) *(proposed)*
|
||||
|
||||
### How it is built
|
||||
|
||||
|
||||
+17
-2
@@ -37,7 +37,22 @@ mesh knows puts it there.
|
||||
or an operator setting naming them. The first is derived and stays true when the uplink changes;
|
||||
the second is a value somebody must remember to edit.
|
||||
|
||||
## The operator's call (2026-10-05): fix the logins, not the jail
|
||||
|
||||
Remedy 2 is rejected. **A machine of the mesh should not fail logins, and when it does, the ban is
|
||||
the jail working.** Exempting the operator's address would hide exactly the failures worth seeing —
|
||||
here a misconfigured ssh client, and in the same week a mail server that had lost every account
|
||||
(issue 241). Both bans were genuinely failed logins; fail2ban cannot tell why a login failed, and it
|
||||
should not have to.
|
||||
|
||||
What stays:
|
||||
|
||||
1. **Remedy 1** — the forge in the ssh configuration the mesh writes, so no machine of the mesh
|
||||
presents the wrong user. This is the fix.
|
||||
2. **A broken server must not look like bad clients.** When a service refuses *every* login — the
|
||||
mail server after its database was emptied — its own health check should fail and say so before
|
||||
the jail has banned its users. That belongs to the service's module, not to the jail.
|
||||
|
||||
## Status
|
||||
|
||||
Not located further: both remedies are design choices — (1) a served port the controller translates
|
||||
by listen, (2) the open question 1 of the report.
|
||||
Not located further: remedy 1 needs a served port the controller translates by listen.
|
||||
|
||||
@@ -19,6 +19,13 @@ What survived did so by accident: git history because every repository is also c
|
||||
operator's machines; the file-sync service's files because they live in the object store, which
|
||||
nothing dropped; the mailboxes because they are files outside the database.
|
||||
|
||||
## Scope, set by the operator (2026-10-05)
|
||||
|
||||
**Mistakes, not disasters.** Backups protect against a dropped database, a deleted bucket, a bad
|
||||
migration — not against a dead disk or a lost site; that loss is accepted. So no off-site copy is
|
||||
needed, and questions 4's "spread across the mesh" and "encrypted to whom" fall away. Worked out in
|
||||
research 030; proposed as ADR 0214.
|
||||
|
||||
## Questions this must answer
|
||||
|
||||
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
|
||||
|
||||
Reference in New Issue
Block a user