ADR 0148: the mesh's names are resolved, not copied into every container

Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).

A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.

Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.

Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
This commit is contained in:
2026-09-30 00:34:30 +02:00
parent 5a3dee9e9e
commit ec42ee0846
5 changed files with 219 additions and 1 deletions
@@ -0,0 +1,159 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 148. The mesh's names are resolved, not copied into every container
## Context
The mesh gives every container it declares the whole roster of mesh names as entries written into
the container's own hosts file at creation
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
looks again.
Three issues are the same fact arriving three times.
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
private address; the declaration followed it within one push and nothing on the machine did. The
forge's container held the old address, lost its database, reported healthy while its existing
connections lasted, and then the public name went down
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
against a database it could no longer find, while the mesh reported the machine as doing what it was
told. Inside it, `novox.internal` was an address that had not existed for five days
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
other containers were current, none of them corrected — each had been recreated for some other
reason and picked up the roster on the way.
135 was fixed by putting the roster into the digest the host compares a container against, so a
container whose names moved is recreated like one whose image moved. **That made the roster part of
every container's identity**, which is the third arrival:
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
took four routine actions; each changed the roster, and each replaced every container on the control
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
was unreachable twice while its own store came back through crash recovery
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
the replaced containers had anything to do with the module being migrated, or with its machine.
The blast radius of a name is now every container that carries the list, which is all of them. The
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
twenty-five — and each would be a full restart of every service on the hub.
## Considered Options
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
A rollback costs another.
**2. Scope each container's entries to the names it actually binds.** A container is given the names
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
new mechanism, and keeps 135's guarantee exactly.
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
can call anything on it** — three cases, same machine, the private network, the public network, and
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
told in advance that it would be wanted, and a person debugging inside a container would find names
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
churn returns the moment a widely-bound name moves — smaller, not gone.
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
## Decision
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
mesh name and no mesh address is written into a container, and none is part of a container's
identity.**
The three consequences that make this worth doing:
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
lookup, in every container, with nothing recreated and nothing restarted.
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
container on another.
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
every asker on the machine, exactly as it answers the machine itself.
**The resolver is a machine-level process, not a container** — one of the modules that is not a
container at all — so a container depending on it is not the circularity it would be if the mesh's
own store had to resolve a name through something the store's own runtime had to start first.
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
move when the mesh's roster does, and the mesh does not know what they mean.
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
### The order this lands in, which is not a preference
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
network asks from an address the converged filter drops, so it has no DNS at all
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
public resolver instead. Both are prerequisites, not related work.
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
creation-time argument, or the resolver's address is back in every container's identity and the
problem has only got smaller.
3. **Then, and only then, the roster leaves the declaration and the digest.**
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
closed by this record, only answered by it.
## How this is checked
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
from a container that was running before the move and has not been touched since, the name answers
with the new address. This is the one 109 and 135 would both have failed.
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
machine is recreated. The apply report on each machine says nothing changed. This is 151.
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
routed name the mesh serves resolves — including names the module never declared a requirement on,
which is the guarantee option 2 would have given up.
- **On every network the runtime offers.** The first three hold for a container on the runtime's
default network as well as one on a declared network, because the default network is the case that
has no DNS today.
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
does not move when the mesh's roster does, and does move when the module's own declared entries do.
## Consequences
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
resolve, which is already true of the machine itself, and is a smaller event than a roster change
destroying and recreating every container on the machine.
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
service names and wildcards under `<node>.internal`, which is why the resolver was built.
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
every asker in the mesh; it reaches them through the resolver rather than by being written into each
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
restarting itself whenever it learns a name.
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
a nameserver would be for. This record accepts that consequence rather than working around it: a
person debugging in a hand-started container resolving the same names as everything else is the
behaviour worth having, and it is what "anything can call anything" means.
## References
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
+1
View File
@@ -174,6 +174,7 @@ python3 00-META/checks/index.py fail if stale
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md) - **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md) - **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md) - **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
### What runs on them, and how it gets there ### What runs on them, and how it gets there
+26 -1
View File
@@ -7,8 +7,9 @@ code:
- mesh-controller internal/identity/authority.go - mesh-controller internal/identity/authority.go
- mesh-host internal/identity/serving.go - mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set) - mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-09-29 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md - 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md - 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md - 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
@@ -289,6 +290,30 @@ hosts file by the runtime. That extends the file decision rather than overturnin
mesh and not chosen by a module: a module that listed the machines would go stale the day one mesh and not chosen by a module: a module that listed the machines would go stale the day one
joins, and a module that did not would be one whose containers cannot reach anything by name. joins, and a module that did not would be one whose containers cannot reach anything by name.
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
and how it will keep working. It no longer describes containers.
Copying the roster into each container made the roster part of each container's identity, so one name
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
registry, the edge and mail on another
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
moves, twice found as a container holding an address that had not existed for days
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
circular is being asked for. This is gated on a container being able to reach the resolver from any of
the runtime's networks, which it cannot today
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)) —
until that lands the mesh keeps copying and keeps comparing, and the order is stated in the record.
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
started by hand resolves the same names as everything else, because the resolver answers the machine,
not a list of containers.
**The boundary, which is deliberate and worth stating:** *declared* containers. A container **The boundary, which is deliberate and worth stating:** *declared* containers. A container
somebody starts by hand is not the mesh's to configure, and reaching into every container on a somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for. machine — declared or not — is what a nameserver in `resolv.conf` would be for.
@@ -51,3 +51,19 @@ knows that is what the rule means.
the runtime's default one? That is a stronger rule and would have prevented 109 as well. the runtime's default one? That is a stronger rule and would have prevented 109 as well.
- What checks it? A converged bed with a container on the default network resolving a mesh name is - What checks it? A converged bed with a container on the default network resolving a mesh name is
the missing assertion; nothing in the resolver's own beds covers the filter. the missing assertion; nothing in the resolver's own beds covers the filter.
## What now depends on this (2026-09-30)
This stopped being a container-DNS inconvenience.
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
one name moving from replacing every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
stale address impossible rather than merely noticed
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
machines also bind the resolver to loopback only, so the runtime hands their containers a public
resolver. Both halves are this issue.
@@ -75,3 +75,20 @@ the list. 152 removed the false reasons; the question below is still open.
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
rather than baked entries, or scope each container's entries to the names it actually binds. rather than baked entries, or scope each container's entries to the names it actually binds.
## Answered (2026-09-30): the first of those two
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
the churn returns whenever a widely-bound name moves.
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
**This record stays open**, because the record answers it and the code does not. Nothing may stop
copying names until a container can reach the resolver from any of the runtime's networks
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
silently, on a live mesh, which is how both were found. The order is in the record.