From 93d4d29b72268e0a5d676d3358fa7230b5890758 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 5 Oct 2026 02:24:11 +0200 Subject: [PATCH] Issues 241, 242: one unreadable grants file dropped every database, and the mesh has no backups 241: the SDK harness read a failed read as "no consumer" and withdrew all seven databases on the control node; withdrawal destroyed data in seven providers. Fixed in mesh-sdk 0.1.10, mesh-catalog#44 and mesh-host#22; the recovery and what it lost are recorded. 242: nothing backs anything up. --- .../00-report.md | 89 +++++++++++++++++++ .../242-the-mesh-has-no-backups/00-report.md | 42 +++++++++ 2 files changed, 131 insertions(+) create mode 100644 04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md create mode 100644 04-ISSUES/242-the-mesh-has-no-backups/00-report.md diff --git a/04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md b/04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md new file mode 100644 index 0000000..3440623 --- /dev/null +++ b/04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md @@ -0,0 +1,89 @@ +--- +status: resolved +opened: 2026-10-05 +located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply] +fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22 +amended-design: +--- + +# 241 — One unreadable grants file dropped every database on the control node + +## What was observed + +2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every +connection to the seven databases it manages, dropped them, and seconds later created them again, +empty: the forge, the mail server's admin, the identity provider, the file-sync service, the +analytics service, the catalogue and the licence manager. postgres's own log shows it — connections +"terminating due to administrator command", then clients told the database "seems to have just been +dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42. + +Nothing failed loudly. Each application kept running against an empty database: the mail server +refused every login (`relation "user" does not exist`), and those refused logins got the operator's +home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed; +the file-sync service answered 500. + +Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's +177 GB of objects. + +## Why + +**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's +contributions file every five seconds. `readContributions` returned an empty list — without a log +line — when the file could not be read, was not JSON, or was for another requirement. The reconcile +pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was +`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied +its consumers, was enough. The next read was fine, so every consumer came back — to an empty database. + +**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged +since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In +the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day +while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and +an agent session on the laptop was inspecting this provisioner's process from 17:52. + +**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the +database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the +mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted +the site. Any of them would have lost its data to the same misread. + +## Resolution + +- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or + unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says + why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made. + A test reproduces the incident and fails on the old harness. +- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and + keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key + and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the + analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a + database out of service is an operator's tool (`postgres_retire_database`), and it renames to + `_deleted_` — nothing in the mesh drops a database. +- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only + while it holds exactly what the host wrote; otherwise it is moved aside to `.removed-