Issues 241, 242: one unreadable grants file dropped every database, and the mesh has no backups
241: the SDK harness read a failed read as "no consumer" and withdrew all seven databases on the control node; withdrawal destroyed data in seven providers. Fixed in mesh-sdk 0.1.10, mesh-catalog#44 and mesh-host#22; the recovery and what it lost are recorded. 242: nothing backs anything up.
This commit is contained in:
+89
@@ -0,0 +1,89 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply]
|
||||
fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 241 — One unreadable grants file dropped every database on the control node
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every
|
||||
connection to the seven databases it manages, dropped them, and seconds later created them again,
|
||||
empty: the forge, the mail server's admin, the identity provider, the file-sync service, the
|
||||
analytics service, the catalogue and the licence manager. postgres's own log shows it — connections
|
||||
"terminating due to administrator command", then clients told the database "seems to have just been
|
||||
dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42.
|
||||
|
||||
Nothing failed loudly. Each application kept running against an empty database: the mail server
|
||||
refused every login (`relation "user" does not exist`), and those refused logins got the operator's
|
||||
home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed;
|
||||
the file-sync service answered 500.
|
||||
|
||||
Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's
|
||||
177 GB of objects.
|
||||
|
||||
## Why
|
||||
|
||||
**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's
|
||||
contributions file every five seconds. `readContributions` returned an empty list — without a log
|
||||
line — when the file could not be read, was not JSON, or was for another requirement. The reconcile
|
||||
pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was
|
||||
`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied
|
||||
its consumers, was enough. The next read was fine, so every consumer came back — to an empty database.
|
||||
|
||||
**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged
|
||||
since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In
|
||||
the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day
|
||||
while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and
|
||||
an agent session on the laptop was inspecting this provisioner's process from 17:52.
|
||||
|
||||
**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the
|
||||
database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the
|
||||
mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted
|
||||
the site. Any of them would have lost its data to the same misread.
|
||||
|
||||
## Resolution
|
||||
|
||||
- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or
|
||||
unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says
|
||||
why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made.
|
||||
A test reproduces the incident and fails on the old harness.
|
||||
- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and
|
||||
keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key
|
||||
and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the
|
||||
analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a
|
||||
database out of service is an operator's tool (`postgres_retire_database`), and it renames to
|
||||
`<name>_deleted_<date>` — nothing in the mesh drops a database.
|
||||
- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only
|
||||
while it holds exactly what the host wrote; otherwise it is moved aside to `<path>.removed-<time>`.
|
||||
Directories with content were already kept.
|
||||
|
||||
## Recovery
|
||||
|
||||
No database had a backup newer than the migration; issue 242 is the gap. Each was restored from the
|
||||
newest copy on the control node, by the same path: copy the source, start a throwaway postgres of its
|
||||
version with no network, dump the one database, restore it beside the live one owned by the consumer's
|
||||
role, check counts, stop the application, rename the empty live database aside
|
||||
(`…_deleted_20261005`, kept) and the restored one into place, start, verify as a user.
|
||||
|
||||
| application | restored to | notes |
|
||||
|---|---|---|
|
||||
| forge | 2026-09-22 | git data complete and current; three repositories created since were adopted |
|
||||
| mail | 2026-09-25 | all accounts |
|
||||
| identity provider | 2026-09-26 | both realms |
|
||||
| file-sync service | 2026-09-25 (a MariaDB dump, converted with the application's own `db:convert-type`) | objects intact; index entries without an object were previews, trash and stock sample files |
|
||||
| analytics, catalogue | 2026-09-24 | |
|
||||
| licence manager | none | schema re-created by its preparation step; licences re-adopted from the nodes |
|
||||
|
||||
The forge's restore had its own consequences: the forge watcher's admin account and the build agents'
|
||||
registry accounts were created after the backup and were recreated (the first by the module's own
|
||||
bootstrap command, the second by its provisioner), and every pull request and issue since 2026-09-22
|
||||
is gone from the forge's records — the branches remain.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The SDK test above; the postgres provider's test that withdrawal and retirement issue no `DROP`; the
|
||||
host's test that a file holding more than the host wrote is moved aside, never deleted.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 242 — The mesh has no backups
|
||||
|
||||
## What was observed
|
||||
|
||||
When seven databases were dropped on 2026-10-04 (issue 241), the newest copy of any of them was a
|
||||
leftover of the migration: data directories and one dump, nine to twelve days old, found by searching
|
||||
the control node's disk. Nothing in the mesh takes a backup, no record says what should be backed up,
|
||||
and nothing would have said so until the day one was needed.
|
||||
|
||||
What survived did so by accident: git history because every repository is also cloned on the
|
||||
operator's machines; the file-sync service's files because they live in the object store, which
|
||||
nothing dropped; the mailboxes because they are files outside the database.
|
||||
|
||||
## Questions this must answer
|
||||
|
||||
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
|
||||
buckets, the mail spool, the forge's repositories and data, the vault and the controller's own
|
||||
records, each module's state directories? Is it declared by the module that owns the data, the way
|
||||
a module declares its listens and its jails?
|
||||
2. **How often, and how long kept** — a daily schedule? Rotation: how many daily, weekly, monthly?
|
||||
3. **Full or incremental** — full dumps for databases, deltas (snapshots, deduplicating archives) for
|
||||
large object stores and file trees?
|
||||
4. **Where** — never only on the machine whose disk it copies. Spread across the mesh (the home server
|
||||
holding the control node's, and the other way round), an external target, or both? Encrypted to
|
||||
whom?
|
||||
5. **Who runs it** — a seat (`mesh-backup`?) whose holder schedules and stores, with the verbs a person
|
||||
needs: what was backed up and when, restore one database beside the live one?
|
||||
6. **How it is proven** — a backup never restored is a hope. A scheduled restore of the newest copy into
|
||||
a throwaway instance, compared with the live one, and a failure that reaches the operator.
|
||||
|
||||
## Why it matters beyond this incident
|
||||
|
||||
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
|
||||
held on another machine, it would have been a ten-minute restore of yesterday.
|
||||
Reference in New Issue
Block a user