Files
hq/04-ISSUES/097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md
T
jschoubben c4151e6bc4 ADR 0163 built: the take digest, the minted-secret refusal, the networks setting, settings judged where stored, genesis raising the forge as declared; issues 096, 097, 126 resolved
The record gets its built note; designs 05 and 09 the revisions; 086, 098,
099, 100 and 101 stay located because every machine is converged and the
record's live row — a take read on an adopted machine — has not been run;
090 is built in part, its network difference left for the take to say.
2026-10-01 23:45:57 +02:00

91 lines
5.0 KiB
Markdown

---
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by: mesh-host 63 (former targets removed, strays reported), mesh-controller 201/202 (strays shown)
amended-design:
---
# 097 — A resource whose target changes leaves the old one behind, running
## What was observed
On the control-node, 2026-09-23, four hours after the forge's module was first assigned.
A module resource named its container implicitly: with no name of its own, the host derived one
from the module and the resource's id. The manifest then gained an explicit name, because the
module had to take over a container the predecessor already ran under that name
([issue 090](../090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md)).
The host applied the change by raising a container under the new name — and left the one under the
old name **running**.
The host's own record shows why. It keeps one entry per resource id, and that entry holds the
resource's *current* target:
```
{"id": "gitea.server", "type": "container", "origin": "declared",
"holds": [2999, 2222], "target": "gitea", "applied_at": "..."}
```
There is no entry naming the old container. Rewriting the record on a target change is what
erases the only trace of what the host must now remove, so the old one cannot be found by the
thing that would have removed it.
Every other container the mesh made on this machine is declared; this is the only one that is not.
It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing
stops it, and nothing reports it.
Harmless in this instance by luck — and less harmless than it first looked. **Addendum,
2026-09-24**, after reading it properly rather than glancing at it: the stranded container was
running on the **host's own network**, so it was listening on a port on every interface of the
machine, and it held open connections to the mesh's store. It reached the store through a forwarder
that had been put in front of an old address for an unrelated reason, which is the only thing that
kept the two facts from meeting sooner.
What saved it was that its configuration named its own former database rather than the one the
service now uses, so no data was at risk. Nothing in the mesh arranged that. Had the rename gone
the other way — or had the database name not changed with it — two versions of one service would
have been writing to one database, one of them a version older than the schema.
## Why it matters beyond this instance
The mesh's promise is that a machine runs what it was told and nothing else, and that a module
removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record
that remembers only the current target breaks that for **any** resource whose target moves — and a
file is worse than a container, because a stale configuration file at the old path is still read
by whatever reads that path, silently, with no process to notice running twice.
Renaming is not exotic. It happens exactly when a module is taught to take over something that
already exists, which is every module in a migration.
It also crosses [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md):
on an adopted node the host must distinguish what it wrote from what it found, because what it
found is held and never removed. A resource that changes target turns something the host wrote
into something no record claims — which, on the next machine, is indistinguishable from something
found, and so would be kept for ever on purpose.
## Open questions
- Should the record keep every target a resource has had, and the host remove the ones it no
longer declares — and if so, for how long, given a record is also how the host knows what it may
destroy?
- Should the host refuse a target change outright, requiring the old resource to be removed by a
declaration that still names it before a new one may take the name?
- Does the same hole exist for a resource whose **id** changes while the target stays, and for a
module unassigned between the two declarations?
- What reports this? Nothing on the machine currently answers "what is running here that the mesh
did not ask for", which is the question that would have found this in seconds.
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 5: former targets are removed and strays reported. Building follows,
host first, then the controller's `take`.
## Resolved, 2026-10-02
mesh-host 63: the host's record keeps a resource's former targets, removes a container or file it
wrote under a name the declaration no longer names, never what was found, and reports strays — what
runs on the machine that the mesh neither wrote nor holds. mesh-controller 201 and 202 show strays
on `node show` for an adopted and a converged machine alike; the live mesh reported four on the
control node the evening it rolled.