Issues 241, 242: one unreadable grants file dropped every database, and the mesh has no backups

241: the SDK harness read a failed read as "no consumer" and withdrew all seven databases on the
control node; withdrawal destroyed data in seven providers. Fixed in mesh-sdk 0.1.10, mesh-catalog#44
and mesh-host#22; the recovery and what it lost are recorded. 242: nothing backs anything up.
This commit is contained in:
2026-10-05 02:24:11 +02:00
parent 33c10aa85a
commit 93d4d29b72
2 changed files with 131 additions and 0 deletions
@@ -0,0 +1,89 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply]
fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22
amended-design:
---
# 241 — One unreadable grants file dropped every database on the control node
## What was observed
2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every
connection to the seven databases it manages, dropped them, and seconds later created them again,
empty: the forge, the mail server's admin, the identity provider, the file-sync service, the
analytics service, the catalogue and the licence manager. postgres's own log shows it — connections
"terminating due to administrator command", then clients told the database "seems to have just been
dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42.
Nothing failed loudly. Each application kept running against an empty database: the mail server
refused every login (`relation "user" does not exist`), and those refused logins got the operator's
home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed;
the file-sync service answered 500.
Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's
177 GB of objects.
## Why
**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's
contributions file every five seconds. `readContributions` returned an empty list — without a log
line — when the file could not be read, was not JSON, or was for another requirement. The reconcile
pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was
`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied
its consumers, was enough. The next read was fine, so every consumer came back — to an empty database.
**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged
since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In
the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day
while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and
an agent session on the laptop was inspecting this provisioner's process from 17:52.
**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the
database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the
mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted
the site. Any of them would have lost its data to the same misread.
## Resolution
- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or
unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says
why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made.
A test reproduces the incident and fails on the old harness.
- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and
keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key
and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the
analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a
database out of service is an operator's tool (`postgres_retire_database`), and it renames to
`<name>_deleted_<date>` — nothing in the mesh drops a database.
- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only
while it holds exactly what the host wrote; otherwise it is moved aside to `<path>.removed-<time>`.
Directories with content were already kept.
## Recovery
No database had a backup newer than the migration; issue 242 is the gap. Each was restored from the
newest copy on the control node, by the same path: copy the source, start a throwaway postgres of its
version with no network, dump the one database, restore it beside the live one owned by the consumer's
role, check counts, stop the application, rename the empty live database aside
(`…_deleted_20261005`, kept) and the restored one into place, start, verify as a user.
| application | restored to | notes |
|---|---|---|
| forge | 2026-09-22 | git data complete and current; three repositories created since were adopted |
| mail | 2026-09-25 | all accounts |
| identity provider | 2026-09-26 | both realms |
| file-sync service | 2026-09-25 (a MariaDB dump, converted with the application's own `db:convert-type`) | objects intact; index entries without an object were previews, trash and stock sample files |
| analytics, catalogue | 2026-09-24 | |
| licence manager | none | schema re-created by its preparation step; licences re-adopted from the nodes |
The forge's restore had its own consequences: the forge watcher's admin account and the build agents'
registry accounts were created after the backup and were recreated (the first by the module's own
bootstrap command, the second by its provisioner), and every pull request and issue since 2026-09-22
is gone from the forge's records — the branches remain.
## How it is checked
The SDK test above; the postgres provider's test that withdrawal and retirement issue no `DROP`; the
host's test that a file holding more than the host wrote is moved aside, never deleted.
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 242 — The mesh has no backups
## What was observed
When seven databases were dropped on 2026-10-04 (issue 241), the newest copy of any of them was a
leftover of the migration: data directories and one dump, nine to twelve days old, found by searching
the control node's disk. Nothing in the mesh takes a backup, no record says what should be backed up,
and nothing would have said so until the day one was needed.
What survived did so by accident: git history because every repository is also cloned on the
operator's machines; the file-sync service's files because they live in the object store, which
nothing dropped; the mailboxes because they are files outside the database.
## Questions this must answer
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
buckets, the mail spool, the forge's repositories and data, the vault and the controller's own
records, each module's state directories? Is it declared by the module that owns the data, the way
a module declares its listens and its jails?
2. **How often, and how long kept** — a daily schedule? Rotation: how many daily, weekly, monthly?
3. **Full or incremental** — full dumps for databases, deltas (snapshots, deduplicating archives) for
large object stores and file trees?
4. **Where** — never only on the machine whose disk it copies. Spread across the mesh (the home server
holding the control node's, and the other way round), an external target, or both? Encrypted to
whom?
5. **Who runs it** — a seat (`mesh-backup`?) whose holder schedules and stores, with the verbs a person
needs: what was backed up and when, restore one database beside the live one?
6. **How it is proven** — a backup never restored is a hope. A scheduled restore of the newest copy into
a throwaway instance, compared with the live one, and a failure that reaches the operator.
## Why it matters beyond this incident
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.