Files
hq/04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md
T
jschoubben 04c9500b5b Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
2026-09-30 14:56:43 +02:00

86 lines
5.0 KiB
Markdown

---
status: resolved
opened: 2026-09-24
located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
amended-design:
---
# 109 — A container keeps the address it was made with, and nothing recreates it when the node moves
## What was observed
On the control-node, 2026-09-24, minutes after the mesh took over the predecessor's tunnel
([ADR 0105](../../02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md)).
Adopting the tunnel moves the mesh onto the tunnel's range, so the hub's private address changed.
Everything the controller composes followed it: the declaration's hosts entries for every container
said the new address within one push. Nothing on the machine followed. The forge's container had
the **old** address written in its own hosts file — put there when the container was created — and
so lost its database. It reported healthy for as long as its existing connections lasted, then
answered failures, and the public name went down.
```
declared: novox.internal -> <new hub address>
in the container: novox.internal -> <old hub address>
```
The container was not recreated, and was still running from before the move. Removing it and
letting the mesh make it again fixed it in one step.
The service beside it was untouched by the same change, which is the part worth keeping: it sits on
a network of its own, so it asks a resolver for the name rather than reading a copy of the answer.
Only a container on the runtime's default network gets the answer baked in.
## Why it matters beyond this instance
The mesh's rule is that an address is read where it is used and never recorded
([issue 102](../102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)).
This is that rule broken one level further down: the controller obeys it, and then the runtime
writes the answer into the container at creation, where it stays for the life of the container.
A node's address changes more often than it looks: adopting a tunnel, moving a machine between
sites, renumbering a range. Each time, every container made before the change keeps pointing at
where the node used to be — and says nothing, because from the mesh's side the declaration is
correct and the node reports it applied.
It is the same shape as
[issue 103](../103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md), whose fix
made the *content of a file* a container reads part of what the host compares. The hosts entries a
container is made with are the same kind of input, read once at creation, and are not compared.
## Open questions
- Should the hosts entries a container is created with be part of what the host compares, so a
changed address recreates it — accepting that a node's address change restarts every container on
it?
- Or should a container never be given the answer at all, and always ask the resolver — which is
what the containers that survived this do? That makes the resolver a dependency of every module
and is the stronger statement; it is also what the mesh's own resolver exists for.
- Either way: what tells an operator that a container is running with an address the node no longer
has? Nothing did.
## Answered at the cause (2026-09-30)
This was the first of three arrivals of one fact: a container is given the mesh's names when it is
created and never looks again, so a name that moves afterwards is wrong inside it for as long as it
runs. It arrived again as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md),
whose fix made the names comparable — and that fix made the roster part of every container's identity,
which arrived as [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) ends the
copying: a container resolves through its machine's resolver at the moment it asks. The shape this
record reports then has nowhere to occur. It is gated on
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
until that lands the mesh still copies and still compares.
## Resolved (2026-09-30)
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
after its containers were recreated once — the last time a name will do that: the forge's container
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.