Merge pull request 'Issues 239–242: a name taken over, a dry run recorded, every database dropped, no backups' (#83) from issues/239-242 into main
This commit was merged in pull request #83.
This commit is contained in:
+43
@@ -0,0 +1,43 @@
|
||||
# 238 — Diagnosis
|
||||
|
||||
## 2026-10-04, from the control node
|
||||
|
||||
**What the forge refused, from the operator's uplink, in the ban's last minute** (the forge's ssh log):
|
||||
|
||||
```
|
||||
14:05:28 Invalid user jochen from <uplink> port 38564
|
||||
14:05:31 Accepted publickey for git from <uplink> port 45156 (the laptop's key)
|
||||
14:05:44 Invalid user jochen from <uplink> port 52074
|
||||
14:05:59 Invalid user jochen from <uplink> port 37978
|
||||
```
|
||||
|
||||
Three refusals in thirty-one seconds — `maxretry = 3` — and the `gitea` jail banned the uplink at
|
||||
14:06:00; `recidive` counted it the same second. Between the refusals the same key logged in as `git`:
|
||||
the agent's own git operations were fine, and the refusals were ssh commands that named no user.
|
||||
|
||||
**Why they named the wrong user.** On the laptop and the workstation, `ssh -G <forge's public name>`
|
||||
resolves to the operator's account and port 22: nothing in the ssh configuration the mesh writes
|
||||
(`ssh-client`, to-be 29) names the forge. A bare `ssh <forge>` — or a git URL without `git@` — presents
|
||||
the login name, which the forge does not have.
|
||||
|
||||
**Why the operator's uplink is bannable at all.** The jails' `ignoreip` (the `fail2ban` module's
|
||||
`jail.local`) holds loopback, the mesh's private range and every private range (ADR 0186). A node at
|
||||
home reaches the control node from the home's public address, which is none of those. Nothing the
|
||||
mesh knows puts it there.
|
||||
|
||||
## What would have stopped it
|
||||
|
||||
1. **The forge in the ssh configuration the mesh writes**: a `Host` block for the forge's public and
|
||||
internal names with `User git` and the forge's ssh port. It cannot be written today: the `git`
|
||||
provision serves the forge's http port only, and the controller translates a served `port` to the
|
||||
machine's published port but no other key (`ServedOn`), so an `ssh-port` would reach a consumer as
|
||||
the container's 22, not the machine's 222.
|
||||
2. **The mesh's own public addresses in every jail's ignore list.** Two sources: each node's public
|
||||
egress as the hub sees it (the tunnel's peer endpoints — every node at home shows the uplink there),
|
||||
or an operator setting naming them. The first is derived and stays true when the uplink changes;
|
||||
the second is a value somebody must remember to edit.
|
||||
|
||||
## Status
|
||||
|
||||
Not located further: both remedies are design choices — (1) a served port the controller translates
|
||||
by listen, (2) the open question 1 of the report.
|
||||
+48
@@ -0,0 +1,48 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-04
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 239 — A module name is taken over by another repository, and nothing refuses it
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-04. Five of the photo app's six public names stopped answering on the control node — the API,
|
||||
two client sites and two aliases — while the sixth answered with a different program. Nothing failed:
|
||||
the controller composed, the host applied, every check passed.
|
||||
|
||||
Two definitions held the module name `photos`:
|
||||
|
||||
| | the app's own repository | the catalogue's `modules/photos` |
|
||||
|---|---|---|
|
||||
| built from | `photos.git`, branch `nox-mesh`, until 2026-09-28 | from 2026-10-04 04:14 |
|
||||
| containers | server, admin, two client sites | server, an admin client |
|
||||
| routes | six | one |
|
||||
|
||||
Rebuilding `photos` "to `main`" for [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md)
|
||||
(see [issue 227](../227-the-photo-apps-admin-client-asks-for-the-port-the-proxy-holds/00-report.md)) built the
|
||||
catalogue's main, not the repository the module had been built from. The controller recorded the new
|
||||
source, the module moved, the next push replaced the app with the stub, and two sites' containers and
|
||||
five routes went with it. The stub had sat in the catalogue since 2026-09-03 without ever being the
|
||||
running definition.
|
||||
|
||||
Restored the same day by building from `photos.git` again and removing the stub (the app repository's pull request and
|
||||
mesh-catalog#278).
|
||||
|
||||
## Why it is an issue and not an incident
|
||||
|
||||
**A module's source is a fact the mesh records, and any build may overwrite it.** The build history
|
||||
showed the switch plainly — `photos.git at nox-mesh` on one line, `mesh-catalog.git at main` on the
|
||||
next — and nothing asked whether a module built from one repository should now come from another.
|
||||
"The last build wins" is the rule in practice; it is written nowhere, and it lets a stale or unrelated
|
||||
definition replace a working one silently.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Should a build whose source differs from the module's recorded source be refused unless it says so
|
||||
explicitly (a `--move-source`, or the operator's confirmation)?
|
||||
2. Should the catalogue's check refuse a module whose name another registered repository already
|
||||
defines?
|
||||
@@ -0,0 +1,37 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-04
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 240 — A dry-run build is recorded, and what it built is applied
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-04, restoring the photo app ([issue 239](../239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md)).
|
||||
`mesh-controller build <repository> --ref <unmerged branch> --dry-run` was run to prove the branch built
|
||||
before it was merged. The command's help says *"build and print the manifest, recording nothing"*.
|
||||
|
||||
Afterwards:
|
||||
|
||||
- `builds photos` listed the dry run as a build — `photos.git at <the unmerged branch>`, with its three
|
||||
images — beside the real ones.
|
||||
- On the control node, the module's state directory was rewritten at a time between the dry run and
|
||||
the merged build: the secret files and environment files the branch's definition declares, which the
|
||||
definition then running did not.
|
||||
|
||||
The content happened to equal what was merged a few minutes later, so nothing broke. Had the branch
|
||||
been rejected in review, its definition would already have been on the machine.
|
||||
|
||||
## Why it matters
|
||||
|
||||
A dry run is how a change is proven before a person approves it. If it records the build and the mesh
|
||||
acts on it, review becomes a formality: the unreviewed definition reaches a machine first.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
|
||||
any other build's?
|
||||
2. Did the controller roll the dry run out, or did a later push compose from it?
|
||||
+89
@@ -0,0 +1,89 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply]
|
||||
fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 241 — One unreadable grants file dropped every database on the control node
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every
|
||||
connection to the seven databases it manages, dropped them, and seconds later created them again,
|
||||
empty: the forge, the mail server's admin, the identity provider, the file-sync service, the
|
||||
analytics service, the catalogue and the licence manager. postgres's own log shows it — connections
|
||||
"terminating due to administrator command", then clients told the database "seems to have just been
|
||||
dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42.
|
||||
|
||||
Nothing failed loudly. Each application kept running against an empty database: the mail server
|
||||
refused every login (`relation "user" does not exist`), and those refused logins got the operator's
|
||||
home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed;
|
||||
the file-sync service answered 500.
|
||||
|
||||
Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's
|
||||
177 GB of objects.
|
||||
|
||||
## Why
|
||||
|
||||
**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's
|
||||
contributions file every five seconds. `readContributions` returned an empty list — without a log
|
||||
line — when the file could not be read, was not JSON, or was for another requirement. The reconcile
|
||||
pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was
|
||||
`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied
|
||||
its consumers, was enough. The next read was fine, so every consumer came back — to an empty database.
|
||||
|
||||
**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged
|
||||
since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In
|
||||
the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day
|
||||
while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and
|
||||
an agent session on the laptop was inspecting this provisioner's process from 17:52.
|
||||
|
||||
**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the
|
||||
database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the
|
||||
mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted
|
||||
the site. Any of them would have lost its data to the same misread.
|
||||
|
||||
## Resolution
|
||||
|
||||
- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or
|
||||
unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says
|
||||
why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made.
|
||||
A test reproduces the incident and fails on the old harness.
|
||||
- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and
|
||||
keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key
|
||||
and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the
|
||||
analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a
|
||||
database out of service is an operator's tool (`postgres_retire_database`), and it renames to
|
||||
`<name>_deleted_<date>` — nothing in the mesh drops a database.
|
||||
- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only
|
||||
while it holds exactly what the host wrote; otherwise it is moved aside to `<path>.removed-<time>`.
|
||||
Directories with content were already kept.
|
||||
|
||||
## Recovery
|
||||
|
||||
No database had a backup newer than the migration; issue 242 is the gap. Each was restored from the
|
||||
newest copy on the control node, by the same path: copy the source, start a throwaway postgres of its
|
||||
version with no network, dump the one database, restore it beside the live one owned by the consumer's
|
||||
role, check counts, stop the application, rename the empty live database aside
|
||||
(`…_deleted_20261005`, kept) and the restored one into place, start, verify as a user.
|
||||
|
||||
| application | restored to | notes |
|
||||
|---|---|---|
|
||||
| forge | 2026-09-22 | git data complete and current; three repositories created since were adopted |
|
||||
| mail | 2026-09-25 | all accounts |
|
||||
| identity provider | 2026-09-26 | both realms |
|
||||
| file-sync service | 2026-09-25 (a MariaDB dump, converted with the application's own `db:convert-type`) | objects intact; index entries without an object were previews, trash and stock sample files |
|
||||
| analytics, catalogue | 2026-09-24 | |
|
||||
| licence manager | none | schema re-created by its preparation step; licences re-adopted from the nodes |
|
||||
|
||||
The forge's restore had its own consequences: the forge watcher's admin account and the build agents'
|
||||
registry accounts were created after the backup and were recreated (the first by the module's own
|
||||
bootstrap command, the second by its provisioner), and every pull request and issue since 2026-09-22
|
||||
is gone from the forge's records — the branches remain.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The SDK test above; the postgres provider's test that withdrawal and retirement issue no `DROP`; the
|
||||
host's test that a file holding more than the host wrote is moved aside, never deleted.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 242 — The mesh has no backups
|
||||
|
||||
## What was observed
|
||||
|
||||
When seven databases were dropped on 2026-10-04 (issue 241), the newest copy of any of them was a
|
||||
leftover of the migration: data directories and one dump, nine to twelve days old, found by searching
|
||||
the control node's disk. Nothing in the mesh takes a backup, no record says what should be backed up,
|
||||
and nothing would have said so until the day one was needed.
|
||||
|
||||
What survived did so by accident: git history because every repository is also cloned on the
|
||||
operator's machines; the file-sync service's files because they live in the object store, which
|
||||
nothing dropped; the mailboxes because they are files outside the database.
|
||||
|
||||
## Questions this must answer
|
||||
|
||||
1. **What is backed up** — every store's databases (postgres, mssql, mongodb), the object store's
|
||||
buckets, the mail spool, the forge's repositories and data, the vault and the controller's own
|
||||
records, each module's state directories? Is it declared by the module that owns the data, the way
|
||||
a module declares its listens and its jails?
|
||||
2. **How often, and how long kept** — a daily schedule? Rotation: how many daily, weekly, monthly?
|
||||
3. **Full or incremental** — full dumps for databases, deltas (snapshots, deduplicating archives) for
|
||||
large object stores and file trees?
|
||||
4. **Where** — never only on the machine whose disk it copies. Spread across the mesh (the home server
|
||||
holding the control node's, and the other way round), an external target, or both? Encrypted to
|
||||
whom?
|
||||
5. **Who runs it** — a seat (`mesh-backup`?) whose holder schedules and stores, with the verbs a person
|
||||
needs: what was backed up and when, restore one database beside the live one?
|
||||
6. **How it is proven** — a backup never restored is a hope. A scheduled restore of the newest copy into
|
||||
a throwaway instance, compared with the live one, and a failure that reaches the operator.
|
||||
|
||||
## Why it matters beyond this incident
|
||||
|
||||
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
|
||||
held on another machine, it would have been a ten-minute restore of yesterday.
|
||||
Reference in New Issue
Block a user