Author SHA1 Message Date
jschoubben 2a501af0eb Merge pull request 'Issues 239–242: a name taken over, a dry run recorded, every database dropped, no backups' (#83) from issues/239-242 into main 2026-10-05 00:25:50 +00:00
jschoubben bd7fc40099 Issue 239: name no installation 2026-10-05 02:24:27 +02:00
jschoubben 93d4d29b72 Issues 241, 242: one unreadable grants file dropped every database, and the mesh has no backups
241: the SDK harness read a failed read as "no consumer" and withdrew all seven databases on the
control node; withdrawal destroyed data in seven providers. Fixed in mesh-sdk 0.1.10, mesh-catalog#44
and mesh-host#22; the recovery and what it lost are recorded. 242: nothing backs anything up.
2026-10-05 02:24:11 +02:00
jschoubben 33c10aa85a Issues 239, 240, and 238's diagnosis: a module name taken over, a dry run recorded, the operator banned
239: two repositories defined photos; a rebuild to the catalogue's main replaced the app with a stub
and nothing refused it. 240: a dry-run build was recorded and its definition reached the machine.
238: the forge refused three ssh logins as the operator's account in 31 seconds; nothing in the mesh's
ssh configuration names the forge, and no jail ignores the mesh's own public addresses.
2026-10-04 17:31:29 +02:00
5 changed files with 259 additions and 0 deletions
@@ -0,0 +1,43 @@
# 238 — Diagnosis
## 2026-10-04, from the control node
**What the forge refused, from the operator's uplink, in the ban's last minute** (the forge's ssh log):
```
14:05:28 Invalid user jochen from <uplink> port 38564
14:05:31 Accepted publickey for git from <uplink> port 45156 (the laptop's key)
14:05:44 Invalid user jochen from <uplink> port 52074
14:05:59 Invalid user jochen from <uplink> port 37978
```
Three refusals in thirty-one seconds — `maxretry = 3` — and the `gitea` jail banned the uplink at
14:06:00; `recidive` counted it the same second. Between the refusals the same key logged in as `git`:
the agent's own git operations were fine, and the refusals were ssh commands that named no user.
**Why they named the wrong user.** On the laptop and the workstation, `ssh -G <forge's public name>`
resolves to the operator's account and port 22: nothing in the ssh configuration the mesh writes
(`ssh-client`, to-be 29) names the forge. A bare `ssh <forge>` — or a git URL without `git@` — presents
the login name, which the forge does not have.
**Why the operator's uplink is bannable at all.** The jails' `ignoreip` (the `fail2ban` module's
`jail.local`) holds loopback, the mesh's private range and every private range (ADR 0186). A node at
home reaches the control node from the home's public address, which is none of those. Nothing the
mesh knows puts it there.
## What would have stopped it
1. **The forge in the ssh configuration the mesh writes**: a `Host` block for the forge's public and
internal names with `User git` and the forge's ssh port. It cannot be written today: the `git`
provision serves the forge's http port only, and the controller translates a served `port` to the
machine's published port but no other key (`ServedOn`), so an `ssh-port` would reach a consumer as
the container's 22, not the machine's 222.
2. **The mesh's own public addresses in every jail's ignore list.** Two sources: each node's public
egress as the hub sees it (the tunnel's peer endpoints — every node at home shows the uplink there),
or an operator setting naming them. The first is derived and stays true when the uplink changes;
the second is a value somebody must remember to edit.
## Status
Not located further: both remedies are design choices — (1) a served port the controller translates
by listen, (2) the open question 1 of the report.
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 239 — A module name is taken over by another repository, and nothing refuses it
## What was observed
2026-10-04. Five of the photo app's six public names stopped answering on the control node — the API,
two client sites and two aliases — while the sixth answered with a different program. Nothing failed:
the controller composed, the host applied, every check passed.
Two definitions held the module name `photos`:
| | the app's own repository | the catalogue's `modules/photos` |
|---|---|---|
| built from | `photos.git`, branch `nox-mesh`, until 2026-09-28 | from 2026-10-04 04:14 |
| containers | server, admin, two client sites | server, an admin client |
| routes | six | one |
Rebuilding `photos` "to `main`" for [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)
(see [issue 227](../227-the-photo-apps-admin-client-asks-for-the-port-the-proxy-holds/00-report.md)) built the
catalogue's main, not the repository the module had been built from. The controller recorded the new
source, the module moved, the next push replaced the app with the stub, and two sites' containers and
five routes went with it. The stub had sat in the catalogue since 2026-09-03 without ever being the
running definition.
Restored the same day by building from `photos.git` again and removing the stub (the app repository's pull request and
mesh-catalog#278).
## Why it is an issue and not an incident
**A module's source is a fact the mesh records, and any build may overwrite it.** The build history
showed the switch plainly — `photos.git at nox-mesh` on one line, `mesh-catalog.git at main` on the
next — and nothing asked whether a module built from one repository should now come from another.
"The last build wins" is the rule in practice; it is written nowhere, and it lets a stale or unrelated
definition replace a working one silently.
## Open questions
1. Should a build whose source differs from the module's recorded source be refused unless it says so
explicitly (a `--move-source`, or the operator's confirmation)?
2. Should the catalogue's check refuse a module whose name another registered repository already
defines?
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 240 — A dry-run build is recorded, and what it built is applied
## What was observed
2026-10-04, restoring the photo app ([issue 239](../239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md)).
`mesh-controller build <repository> --ref <unmerged branch> --dry-run` was run to prove the branch built
before it was merged. The command's help says *"build and print the manifest, recording nothing"*.
Afterwards:
- `builds photos` listed the dry run as a build — `photos.git at <the unmerged branch>`, with its three
images — beside the real ones.
- On the control node, the module's state directory was rewritten at a time between the dry run and
the merged build: the secret files and environment files the branch's definition declares, which the
definition then running did not.
The content happened to equal what was merged a few minutes later, so nothing broke. Had the branch
been rejected in review, its definition would already have been on the machine.
## Why it matters
A dry run is how a change is proven before a person approves it. If it records the build and the mesh
acts on it, review becomes a formality: the unreviewed definition reaches a machine first.
## Open questions
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it?
@@ -0,0 +1,89 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply]
fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22
amended-design:
---
# 241 — One unreadable grants file dropped every database on the control node
## What was observed
2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every
connection to the seven databases it manages, dropped them, and seconds later created them again,
empty: the forge, the mail server's admin, the identity provider, the file-sync service, the
analytics service, the catalogue and the licence manager. postgres's own log shows it — connections
"terminating due to administrator command", then clients told the database "seems to have just been
dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42.
Nothing failed loudly. Each application kept running against an empty database: the mail server
refused every login (`relation "user" does not exist`), and those refused logins got the operator's
home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed;
the file-sync service answered 500.
Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's
177 GB of objects.
## Why
**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's
contributions file every five seconds. `readContributions` returned an empty list — without a log
line — when the file could not be read, was not JSON, or was for another requirement. The reconcile
pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was
`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied
its consumers, was enough. The next read was fine, so every consumer came back — to an empty database.
**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged
since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In
the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day
while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and
an agent session on the laptop was inspecting this provisioner's process from 17:52.
**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the
database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the
mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted
the site. Any of them would have lost its data to the same misread.
## Resolution
- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or
unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says
why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made.
A test reproduces the incident and fails on the old harness.
- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and
keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key
and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the
analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a
database out of service is an operator's tool (`postgres_retire_database`), and it renames to
`<name>_deleted_<date>` — nothing in the mesh drops a database.
- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only
while it holds exactly what the host wrote; otherwise it is moved aside to `<path>.removed-<time>`.
Directories with content were already kept.
## Recovery
No database had a backup newer than the migration; issue 242 is the gap. Each was restored from the
newest copy on the control node, by the same path: copy the source, start a throwaway postgres of its
version with no network, dump the one database, restore it beside the live one owned by the consumer's
role, check counts, stop the application, rename the empty live database aside
(`…_deleted_20261005`, kept) and the restored one into place, start, verify as a user.
| application | restored to | notes |
|---|---|---|
| forge | 2026-09-22 | git data complete and current; three repositories created since were adopted |
| mail | 2026-09-25 | all accounts |
| identity provider | 2026-09-26 | both realms |
| file-sync service | 2026-09-25 (a MariaDB dump, converted with the application's own `db:convert-type`) | objects intact; index entries without an object were previews, trash and stock sample files |
| analytics, catalogue | 2026-09-24 | |
| licence manager | none | schema re-created by its preparation step; licences re-adopted from the nodes |
The forge's restore had its own consequences: the forge watcher's admin account and the build agents'
registry accounts were created after the backup and were recreated (the first by the module's own
bootstrap command, the second by its provisioner), and every pull request and issue since 2026-09-22
is gone from the forge's records — the branches remain.
## How it is checked
The SDK test above; the postgres provider's test that withdrawal and retirement issue no `DROP`; the
host's test that a file holding more than the host wrote is moved aside, never deleted.
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 242 — The mesh has no backups
## What was observed
When seven databases were dropped on 2026-10-04 (issue 241), the newest copy of any of them was a
leftover of the migration: data directories and one dump, nine to twelve days old, found by searching
the control node's disk. Nothing in the mesh takes a backup, no record says what should be backed up,
and nothing would have said so until the day one was needed.
What survived did so by accident: git history because every repository is also cloned on the
operator's machines; the file-sync service's files because they live in the object store, which
nothing dropped; the mailboxes because they are files outside the database.
## Questions this must answer
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
buckets, the mail spool, the forge's repositories and data, the vault and the controller's own
records, each module's state directories? Is it declared by the module that owns the data, the way
a module declares its listens and its jails?
2. **How often, and how long kept** — a daily schedule? Rotation: how many daily, weekly, monthly?
3. **Full or incremental** — full dumps for databases, deltas (snapshots, deduplicating archives) for
large object stores and file trees?
4. **Where** — never only on the machine whose disk it copies. Spread across the mesh (the home server
holding the control node's, and the other way round), an external target, or both? Encrypted to
whom?
5. **Who runs it** — a seat (`mesh-backup`?) whose holder schedules and stores, with the verbs a person
needs: what was backed up and when, restore one database beside the live one?
6. **How it is proven** — a backup never restored is a hope. A scheduled restore of the newest copy into
a throwaway instance, compared with the live one, and a failure that reaches the operator.
## Why it matters beyond this incident
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.