diff --git a/04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md b/04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md new file mode 100644 index 0000000..ec8c1c4 --- /dev/null +++ b/04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md @@ -0,0 +1,61 @@ +--- +status: located +opened: 2026-09-24 +located-in: [mesh-host internal/apply] +fixed-by: +amended-design: +--- + +# 109 — A container keeps the address it was made with, and nothing recreates it when the node moves + +## What was observed + +On the control-node, 2026-09-24, minutes after the mesh took over the predecessor's tunnel +([ADR 0105](../../02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)). + +Adopting the tunnel moves the mesh onto the tunnel's range, so the hub's private address changed. +Everything the controller composes followed it: the declaration's hosts entries for every container +said the new address within one push. Nothing on the machine followed. The forge's container had +the **old** address written in its own hosts file — put there when the container was created — and +so lost its database. It reported healthy for as long as its existing connections lasted, then +answered failures, and the public name went down. + +``` +declared: novox.internal -> +in the container: novox.internal -> +``` + +The container was not recreated, and was still running from before the move. Removing it and +letting the mesh make it again fixed it in one step. + +The service beside it was untouched by the same change, which is the part worth keeping: it sits on +a network of its own, so it asks a resolver for the name rather than reading a copy of the answer. +Only a container on the runtime's default network gets the answer baked in. + +## Why it matters beyond this instance + +The mesh's rule is that an address is read where it is used and never recorded +([issue 102](../102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)). +This is that rule broken one level further down: the controller obeys it, and then the runtime +writes the answer into the container at creation, where it stays for the life of the container. + +A node's address changes more often than it looks: adopting a tunnel, moving a machine between +sites, renumbering a range. Each time, every container made before the change keeps pointing at +where the node used to be — and says nothing, because from the mesh's side the declaration is +correct and the node reports it applied. + +It is the same shape as +[issue 103](../103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md), whose fix +made the *content of a file* a container reads part of what the host compares. The hosts entries a +container is made with are the same kind of input, read once at creation, and are not compared. + +## Open questions + +- Should the hosts entries a container is created with be part of what the host compares, so a + changed address recreates it — accepting that a node's address change restarts every container on + it? +- Or should a container never be given the answer at all, and always ask the resolver — which is + what the containers that survived this do? That makes the resolver a dependency of every module + and is the stronger statement; it is also what the mesh's own resolver exists for. +- Either way: what tells an operator that a container is running with an address the node no longer + has? Nothing did.