Files
hq/04-ISSUES/097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md
jschoubben 133e10a738 Issue 102 resolved, verified on the machine with both forwarders gone; 097's orphan was on the host network
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
2026-09-24 00:02:13 +02:00

4.1 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-23

097 — A resource whose target changes leaves the old one behind, running

What was observed

On the control-node, 2026-09-23, four hours after the forge's module was first assigned.

A module resource named its container implicitly: with no name of its own, the host derived one from the module and the resource's id. The manifest then gained an explicit name, because the module had to take over a container the predecessor already ran under that name (issue 090). The host applied the change by raising a container under the new name — and left the one under the old name running.

The host's own record shows why. It keeps one entry per resource id, and that entry holds the resource's current target:

{"id": "gitea.server", "type": "container", "origin": "declared",
 "holds": [2999, 2222], "target": "gitea", "applied_at": "..."}

There is no entry naming the old container. Rewriting the record on a target change is what erases the only trace of what the host must now remove, so the old one cannot be found by the thing that would have removed it.

Every other container the mesh made on this machine is declared; this is the only one that is not. It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing stops it, and nothing reports it.

Harmless in this instance by luck — and less harmless than it first looked. Addendum, 2026-09-24, after reading it properly rather than glancing at it: the stranded container was running on the host's own network, so it was listening on a port on every interface of the machine, and it held open connections to the mesh's store. It reached the store through a forwarder that had been put in front of an old address for an unrelated reason, which is the only thing that kept the two facts from meeting sooner.

What saved it was that its configuration named its own former database rather than the one the service now uses, so no data was at risk. Nothing in the mesh arranged that. Had the rename gone the other way — or had the database name not changed with it — two versions of one service would have been writing to one database, one of them a version older than the schema.

Why it matters beyond this instance

The mesh's promise is that a machine runs what it was told and nothing else, and that a module removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record that remembers only the current target breaks that for any resource whose target moves — and a file is worse than a container, because a stale configuration file at the old path is still read by whatever reads that path, silently, with no process to notice running twice.

Renaming is not exotic. It happens exactly when a module is taught to take over something that already exists, which is every module in a migration.

It also crosses ADR 0100: on an adopted node the host must distinguish what it wrote from what it found, because what it found is held and never removed. A resource that changes target turns something the host wrote into something no record claims — which, on the next machine, is indistinguishable from something found, and so would be kept for ever on purpose.

Open questions

  • Should the record keep every target a resource has had, and the host remove the ones it no longer declares — and if so, for how long, given a record is also how the host knows what it may destroy?
  • Should the host refuse a target change outright, requiring the old resource to be removed by a declaration that still names it before a new one may take the name?
  • Does the same hole exist for a resource whose id changes while the target stays, and for a module unassigned between the two declarations?
  • What reports this? Nothing on the machine currently answers "what is running here that the mesh did not ask for", which is the question that would have found this in seconds.