Files
hq/04-ISSUES/097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md
T
jschoubben 235b9ea0e5 Issues 096 and 097, and 094 diagnosed: a setting stored where it cannot work, and a resource that changed target
094's cause is one blind spot read from two ends, written up in its diagnosis; the fix
answers the first open question and not the other two, which become 096. 097 was found
looking at what the forge's cutover left running.
2026-09-23 02:20:48 +02:00

3.6 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-23

097 — A resource whose target changes leaves the old one behind, running

What was observed

On the control-node, 2026-09-23, four hours after the forge's module was first assigned.

A module resource named its container implicitly: with no name of its own, the host derived one from the module and the resource's id. The manifest then gained an explicit name, because the module had to take over a container the predecessor already ran under that name (issue 090). The host applied the change by raising a container under the new name — and left the one under the old name running.

The host's own record shows why. It keeps one entry per resource id, and that entry holds the resource's current target:

{"id": "gitea.server", "type": "container", "origin": "declared",
 "holds": [2999, 2222], "target": "gitea", "applied_at": "..."}

There is no entry naming the old container. Rewriting the record on a target change is what erases the only trace of what the host must now remove, so the old one cannot be found by the thing that would have removed it.

Every other container the mesh made on this machine is declared; this is the only one that is not. It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing stops it, and nothing reports it.

Harmless in this instance by luck — the stranded container published no ports and held an anonymous volume rather than the service's data — and that luck is the point. Had the rename gone the other way, two containers of the same service would have run against one data directory, or the old one would have kept the port the new one needed and the new one would have failed to bind.

Why it matters beyond this instance

The mesh's promise is that a machine runs what it was told and nothing else, and that a module removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record that remembers only the current target breaks that for any resource whose target moves — and a file is worse than a container, because a stale configuration file at the old path is still read by whatever reads that path, silently, with no process to notice running twice.

Renaming is not exotic. It happens exactly when a module is taught to take over something that already exists, which is every module in a migration.

It also crosses ADR 0100: on an adopted node the host must distinguish what it wrote from what it found, because what it found is held and never removed. A resource that changes target turns something the host wrote into something no record claims — which, on the next machine, is indistinguishable from something found, and so would be kept for ever on purpose.

Open questions

  • Should the record keep every target a resource has had, and the host remove the ones it no longer declares — and if so, for how long, given a record is also how the host knows what it may destroy?
  • Should the host refuse a target change outright, requiring the old resource to be removed by a declaration that still names it before a new one may take the name?
  • Does the same hole exist for a resource whose id changes while the target stays, and for a module unassigned between the two declarations?
  • What reports this? Nothing on the machine currently answers "what is running here that the mesh did not ask for", which is the question that would have found this in seconds.