Files
hq/04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md
T
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00

4.1 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
located 2026-09-24
mesh-host internal/apply

109 — A container keeps the address it was made with, and nothing recreates it when the node moves

What was observed

On the control-node, 2026-09-24, minutes after the mesh took over the predecessor's tunnel (ADR 0105).

Adopting the tunnel moves the mesh onto the tunnel's range, so the hub's private address changed. Everything the controller composes followed it: the declaration's hosts entries for every container said the new address within one push. Nothing on the machine followed. The forge's container had the old address written in its own hosts file — put there when the container was created — and so lost its database. It reported healthy for as long as its existing connections lasted, then answered failures, and the public name went down.

declared:   novox.internal -> <new hub address>
in the container: novox.internal -> <old hub address>

The container was not recreated, and was still running from before the move. Removing it and letting the mesh make it again fixed it in one step.

The service beside it was untouched by the same change, which is the part worth keeping: it sits on a network of its own, so it asks a resolver for the name rather than reading a copy of the answer. Only a container on the runtime's default network gets the answer baked in.

Why it matters beyond this instance

The mesh's rule is that an address is read where it is used and never recorded (issue 102). This is that rule broken one level further down: the controller obeys it, and then the runtime writes the answer into the container at creation, where it stays for the life of the container.

A node's address changes more often than it looks: adopting a tunnel, moving a machine between sites, renumbering a range. Each time, every container made before the change keeps pointing at where the node used to be — and says nothing, because from the mesh's side the declaration is correct and the node reports it applied.

It is the same shape as issue 103, whose fix made the content of a file a container reads part of what the host compares. The hosts entries a container is made with are the same kind of input, read once at creation, and are not compared.

Open questions

  • Should the hosts entries a container is created with be part of what the host compares, so a changed address recreates it — accepting that a node's address change restarts every container on it?
  • Or should a container never be given the answer at all, and always ask the resolver — which is what the containers that survived this do? That makes the resolver a dependency of every module and is the stronger statement; it is also what the mesh's own resolver exists for.
  • Either way: what tells an operator that a container is running with an address the node no longer has? Nothing did.

Answered at the cause (2026-09-30)

This was the first of three arrivals of one fact: a container is given the mesh's names when it is created and never looks again, so a name that moves afterwards is wrong inside it for as long as it runs. It arrived again as issue 135, whose fix made the names comparable — and that fix made the roster part of every container's identity, which arrived as issue 151.

ADR 0148 ends the copying: a container resolves through its machine's resolver at the moment it asks. The shape this record reports then has nowhere to occur. It is gated on issue 110, so until that lands the mesh still copies and still compares.