The record gets its built note; designs 05 and 09 the revisions; 086, 098, 099, 100 and 101 stay located because every machine is converged and the record's live row — a take read on an adopted machine — has not been run; 090 is built in part, its network difference left for the take to say.
91 lines
5.0 KiB
Markdown
91 lines
5.0 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-23
|
|
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
|
|
fixed-by: mesh-host 63 (former targets removed, strays reported), mesh-controller 201/202 (strays shown)
|
|
amended-design:
|
|
---
|
|
|
|
# 097 — A resource whose target changes leaves the old one behind, running
|
|
|
|
## What was observed
|
|
|
|
On the control-node, 2026-09-23, four hours after the forge's module was first assigned.
|
|
|
|
A module resource named its container implicitly: with no name of its own, the host derived one
|
|
from the module and the resource's id. The manifest then gained an explicit name, because the
|
|
module had to take over a container the predecessor already ran under that name
|
|
([issue 090](../090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md)).
|
|
The host applied the change by raising a container under the new name — and left the one under the
|
|
old name **running**.
|
|
|
|
The host's own record shows why. It keeps one entry per resource id, and that entry holds the
|
|
resource's *current* target:
|
|
|
|
```
|
|
{"id": "gitea.server", "type": "container", "origin": "declared",
|
|
"holds": [2999, 2222], "target": "gitea", "applied_at": "..."}
|
|
```
|
|
|
|
There is no entry naming the old container. Rewriting the record on a target change is what
|
|
erases the only trace of what the host must now remove, so the old one cannot be found by the
|
|
thing that would have removed it.
|
|
|
|
Every other container the mesh made on this machine is declared; this is the only one that is not.
|
|
It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing
|
|
stops it, and nothing reports it.
|
|
|
|
Harmless in this instance by luck — and less harmless than it first looked. **Addendum,
|
|
2026-09-24**, after reading it properly rather than glancing at it: the stranded container was
|
|
running on the **host's own network**, so it was listening on a port on every interface of the
|
|
machine, and it held open connections to the mesh's store. It reached the store through a forwarder
|
|
that had been put in front of an old address for an unrelated reason, which is the only thing that
|
|
kept the two facts from meeting sooner.
|
|
|
|
What saved it was that its configuration named its own former database rather than the one the
|
|
service now uses, so no data was at risk. Nothing in the mesh arranged that. Had the rename gone
|
|
the other way — or had the database name not changed with it — two versions of one service would
|
|
have been writing to one database, one of them a version older than the schema.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
The mesh's promise is that a machine runs what it was told and nothing else, and that a module
|
|
removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record
|
|
that remembers only the current target breaks that for **any** resource whose target moves — and a
|
|
file is worse than a container, because a stale configuration file at the old path is still read
|
|
by whatever reads that path, silently, with no process to notice running twice.
|
|
|
|
Renaming is not exotic. It happens exactly when a module is taught to take over something that
|
|
already exists, which is every module in a migration.
|
|
|
|
It also crosses [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md):
|
|
on an adopted node the host must distinguish what it wrote from what it found, because what it
|
|
found is held and never removed. A resource that changes target turns something the host wrote
|
|
into something no record claims — which, on the next machine, is indistinguishable from something
|
|
found, and so would be kept for ever on purpose.
|
|
|
|
## Open questions
|
|
|
|
- Should the record keep every target a resource has had, and the host remove the ones it no
|
|
longer declares — and if so, for how long, given a record is also how the host knows what it may
|
|
destroy?
|
|
- Should the host refuse a target change outright, requiring the old resource to be removed by a
|
|
declaration that still names it before a new one may take the name?
|
|
- Does the same hole exist for a resource whose **id** changes while the target stays, and for a
|
|
module unassigned between the two declarations?
|
|
- What reports this? Nothing on the machine currently answers "what is running here that the mesh
|
|
did not ask for", which is the question that would have found this in seconds.
|
|
|
|
## Decided, 2026-10-01
|
|
|
|
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rule 5: former targets are removed and strays reported. Building follows,
|
|
host first, then the controller's `take`.
|
|
|
|
## Resolved, 2026-10-02
|
|
|
|
mesh-host 63: the host's record keeps a resource's former targets, removes a container or file it
|
|
wrote under a name the declaration no longer names, never what was found, and reports strays — what
|
|
runs on the machine that the mesh neither wrote nor holds. mesh-controller 201 and 202 show strays
|
|
on `node show` for an adopted and a converged machine alike; the live mesh reported four on the
|
|
control node the evening it rolled.
|