diff --git a/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md b/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md new file mode 100644 index 0000000..285ecd4 --- /dev/null +++ b/04-ISSUES/238-the-mesh-banned-its-own-operators-address-for-four-weeks/01-diagnosis.md @@ -0,0 +1,43 @@ +# 238 — Diagnosis + +## 2026-10-04, from the control node + +**What the forge refused, from the operator's uplink, in the ban's last minute** (the forge's ssh log): + +``` +14:05:28 Invalid user jochen from port 38564 +14:05:31 Accepted publickey for git from port 45156 (the laptop's key) +14:05:44 Invalid user jochen from port 52074 +14:05:59 Invalid user jochen from port 37978 +``` + +Three refusals in thirty-one seconds — `maxretry = 3` — and the `gitea` jail banned the uplink at +14:06:00; `recidive` counted it the same second. Between the refusals the same key logged in as `git`: +the agent's own git operations were fine, and the refusals were ssh commands that named no user. + +**Why they named the wrong user.** On the laptop and the workstation, `ssh -G ` +resolves to the operator's account and port 22: nothing in the ssh configuration the mesh writes +(`ssh-client`, to-be 29) names the forge. A bare `ssh ` — or a git URL without `git@` — presents +the login name, which the forge does not have. + +**Why the operator's uplink is bannable at all.** The jails' `ignoreip` (the `fail2ban` module's +`jail.local`) holds loopback, the mesh's private range and every private range (ADR 0186). A node at +home reaches the control node from the home's public address, which is none of those. Nothing the +mesh knows puts it there. + +## What would have stopped it + +1. **The forge in the ssh configuration the mesh writes**: a `Host` block for the forge's public and + internal names with `User git` and the forge's ssh port. It cannot be written today: the `git` + provision serves the forge's http port only, and the controller translates a served `port` to the + machine's published port but no other key (`ServedOn`), so an `ssh-port` would reach a consumer as + the container's 22, not the machine's 222. +2. **The mesh's own public addresses in every jail's ignore list.** Two sources: each node's public + egress as the hub sees it (the tunnel's peer endpoints — every node at home shows the uplink there), + or an operator setting naming them. The first is derived and stays true when the uplink changes; + the second is a value somebody must remember to edit. + +## Status + +Not located further: both remedies are design choices — (1) a served port the controller translates +by listen, (2) the open question 1 of the report. diff --git a/04-ISSUES/239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md b/04-ISSUES/239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md new file mode 100644 index 0000000..a53151e --- /dev/null +++ b/04-ISSUES/239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md @@ -0,0 +1,48 @@ +--- +status: open +opened: 2026-10-04 +located-in: [] +fixed-by: +amended-design: +--- + +# 239 — A module name is taken over by another repository, and nothing refuses it + +## What was observed + +2026-10-04. Five of the photo app's six public names stopped answering on the control node — the API, +two client sites and two aliases — while the sixth answered with a different program. Nothing failed: +the controller composed, the host applied, every check passed. + +Two definitions held the module name `photos`: + +| | the app's own repository | the catalogue's `modules/photos` | +|---|---|---| +| built from | `photos.git`, branch `nox-mesh`, until 2026-09-28 | from 2026-10-04 04:14 | +| containers | server, admin, two client sites | server, an admin client | +| routes | six | one | + +Rebuilding `photos` "to `main`" for [ADR 0202](../../02-DECISIONS/0202-a-provider-declares-what-it-derives-for-each-consumer.md) +(see [issue 227](../227-the-photo-apps-admin-client-asks-for-the-port-the-proxy-holds/00-report.md)) built the +catalogue's main, not the repository the module had been built from. The controller recorded the new +source, the module moved, the next push replaced the app with the stub, and two sites' containers and +five routes went with it. The stub had sat in the catalogue since 2026-09-03 without ever being the +running definition. + +Restored the same day by building from `photos.git` again and removing the stub (the app repository's pull request and +mesh-catalog#278). + +## Why it is an issue and not an incident + +**A module's source is a fact the mesh records, and any build may overwrite it.** The build history +showed the switch plainly — `photos.git at nox-mesh` on one line, `mesh-catalog.git at main` on the +next — and nothing asked whether a module built from one repository should now come from another. +"The last build wins" is the rule in practice; it is written nowhere, and it lets a stale or unrelated +definition replace a working one silently. + +## Open questions + +1. Should a build whose source differs from the module's recorded source be refused unless it says so + explicitly (a `--move-source`, or the operator's confirmation)? +2. Should the catalogue's check refuse a module whose name another registered repository already + defines? diff --git a/04-ISSUES/240-a-dry-run-build-is-recorded-and-rolled-out/00-report.md b/04-ISSUES/240-a-dry-run-build-is-recorded-and-rolled-out/00-report.md new file mode 100644 index 0000000..51cc62d --- /dev/null +++ b/04-ISSUES/240-a-dry-run-build-is-recorded-and-rolled-out/00-report.md @@ -0,0 +1,37 @@ +--- +status: open +opened: 2026-10-04 +located-in: [] +fixed-by: +amended-design: +--- + +# 240 — A dry-run build is recorded, and what it built is applied + +## What was observed + +2026-10-04, restoring the photo app ([issue 239](../239-a-module-name-is-taken-over-by-another-repository-and-nothing-refuses/00-report.md)). +`mesh-controller build --ref --dry-run` was run to prove the branch built +before it was merged. The command's help says *"build and print the manifest, recording nothing"*. + +Afterwards: + +- `builds photos` listed the dry run as a build — `photos.git at `, with its three + images — beside the real ones. +- On the control node, the module's state directory was rewritten at a time between the dry run and + the merged build: the secret files and environment files the branch's definition declares, which the + definition then running did not. + +The content happened to equal what was merged a few minutes later, so nothing broke. Had the branch +been rejected in review, its definition would already have been on the machine. + +## Why it matters + +A dry run is how a change is proven before a person approves it. If it records the build and the mesh +acts on it, review becomes a formality: the unreviewed definition reaches a machine first. + +## Open questions + +1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as + any other build's? +2. Did the controller roll the dry run out, or did a later push compose from it? diff --git a/04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md b/04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md new file mode 100644 index 0000000..3440623 --- /dev/null +++ b/04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md @@ -0,0 +1,89 @@ +--- +status: resolved +opened: 2026-10-05 +located-in: [mesh-sdk src/provisioner, mesh-catalog modules/postgres, mesh-catalog modules/mssql, mesh-catalog modules/mongodb, mesh-catalog modules/minio, mesh-catalog modules/mailu, mesh-catalog modules/gitea, mesh-catalog modules/umami, mesh-host internal/apply] +fixed-by: mesh-sdk#7 (0.1.10), mesh-catalog#44, mesh-host#22 +amended-design: +--- + +# 241 — One unreadable grants file dropped every database on the control node + +## What was observed + +2026-10-04 17:56:12 local. On the control node, the postgres provider's provisioner terminated every +connection to the seven databases it manages, dropped them, and seconds later created them again, +empty: the forge, the mail server's admin, the identity provider, the file-sync service, the +analytics service, the catalogue and the licence manager. postgres's own log shows it — connections +"terminating due to administrator command", then clients told the database "seems to have just been +dropped" — and the recreated databases carry new identifiers and creation times of 17:56:40–42. + +Nothing failed loudly. Each application kept running against an empty database: the mail server +refused every login (`relation "user" does not exist`), and those refused logins got the operator's +home address and phone banned by the mail jail and by `recidive`; the forge's git operations failed; +the file-sync service answered 500. + +Files outside postgres were untouched: git repositories, mailboxes, and the file-sync service's +177 GB of objects. + +## Why + +**A failed read was read as "nobody asks".** The SDK's provisioner harness reads the provider's +contributions file every five seconds. `readContributions` returned an empty list — without a log +line — when the file could not be read, was not JSON, or was for another requirement. The reconcile +pass then withdrew every consumer this run had provisioned, and the postgres provider's withdrawal was +`pg_terminate_backend`, `DROP DATABASE`, `DROP ROLE`. One bad read, while the provisioner had applied +its consumers, was enough. The next read was fine, so every consumer came back — to an empty database. + +**Why the read failed at 17:56:12 is not established.** The file and its directory were unchanged +since 12:33 and readable afterwards; the failure left no trace, because that path logged nothing. In +the minutes around it: the node's tool runtime had restarted the provisioner a dozen times that day +while module code moved out of containers, the module's memberships were re-issued at 17:56:04, and +an agent session on the laptop was inspecting this provisioner's process from 17:52. + +**And withdrawal destroyed data in seven providers, not one.** postgres, mssql and mongodb dropped the +database; minio removed the bucket (only an error kept a non-empty one); the mail server deleted the +mailbox; the forge purge-deleted the user and every repository it owned; the analytics service deleted +the site. Any of them would have lost its data to the same misread. + +## Resolution + +- **The harness withdraws nothing it cannot read** (mesh-sdk 0.1.10): an unreadable, unparsable or + unrecognised file, or one without a `given` list, makes the pass apply and remove nothing, and says + why once. Only a file read with a `given` list withdraws. Every removal is logged before it is made. + A test reproduces the incident and fails on the old harness. +- **A withdrawal never destroys a consumer's data** (mesh-catalog#44): postgres locks the role and + keeps the database; mssql disables the login; mongodb strips the user's roles; minio revokes the key + and keeps the bucket; the mail server disables the mailbox; the forge prohibits the login; the + analytics site is kept. Each provider's create already re-enables what withdrawal locks. Taking a + database out of service is an operator's tool (`postgres_retire_database`), and it renames to + `_deleted_` — nothing in the mesh drops a database. +- **The host deletes only what is purely its own** (mesh-host#22): a file it created is removed only + while it holds exactly what the host wrote; otherwise it is moved aside to `.removed-