Issues 102–106 and ADR 0105 from the core migration

Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
This commit is contained in:
2026-09-23 22:50:10 +02:00
parent ee2bdf220c
commit cb2117f1c4
8 changed files with 365 additions and 1 deletions
@@ -0,0 +1,66 @@
---
status: located
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-bootstrap, mesh-controller internal/builder, mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 102 — An address recorded at genesis or at build does not follow the node's ports
## What was observed
The control-node, 2026-09-23, migrating the foundation's store and broker onto the ports the
predecessor served them on — the move [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
describes: *the ports given become that node's settings, and every place that uses them reads them
from there.*
Most places did. When the store was given 6852, the bindings handed to the forge, the analytics
service and the catalogue all said 6852. When the broker was given 5679, its consumers reconnected.
Three places did not, and each took the mesh down in a different way:
1. **The controller's own store address.** A secret written at genesis says `127.0.0.1:5432`. The
store moved; the controller could not reach its inventory; the mesh was headless.
2. **The controller's own broker address.** The same, for `127.0.0.1:5672`. The controller
crash-looped and its control queue filled.
3. **Every image reference the mesh has built.** Each recorded build reads
`<registry-address>:5100/<module>/<artifact>@sha256:…`, and declarations carry that literal. Move the
registry and every fresh pull of a mesh image fails — a new node, a recreate after eviction.
Found by reading before the move; the first two were found by making it.
All three are the same fact: an address was **written down** when the port was decided, rather than
**read** from the node's settings when it is used. The first two live in genesis secrets; the third
lives in stored data, which is worse — it is baked into every build ever recorded.
Both outages were closed by hand with a forwarder on the old address to the new one. Two such
forwarders are holding the control-node's mesh together while this is open. They are not the fix;
they are the shape of the bug, made visible.
## Why it matters beyond this instance
A port that is a setting in nine places and a constant in three is a constant. The migration's
whole method — take the predecessor's ports one service at a time — depends on every reader
following the setting, and the readers that do not are exactly the control plane's own, which is
the worst place for them: the failure is headlessness, and headlessness cannot be repaired through
the mesh.
The image reference case will bite any mesh that changes its registry's port, or moves its registry
to another node, or ever has two registries. It also means an image is recorded by *where it was
pushed* rather than *what it is*: the digest is the identity, the address is a route to it, and the
mesh stores them as one string.
This is the same hole [issue 085](../085-the-packages-port-given-at-genesis-is-not-a-setting/00-report.md)
found for the packages port, fixed there for that one reader. It is not one reader; it is a
category.
## Open questions
- Should the controller read its own store and broker addresses from the node's settings at start,
the way it composes them for every other module — and re-read them when they change, since it is
the thing that changes them?
- Should a recorded build store the artifact's **digest and path** only, with the registry address
composed into the declaration from the node's current settings — so a reference is assembled
where it is used, never stored?
- Is there a way to find the remaining constants mechanically — every place a foundation port
number appears as a literal — rather than one outage at a time?
@@ -0,0 +1,45 @@
---
status: located
opened: 2026-09-23
located-in: [mesh-host internal/apply]
fixed-by:
amended-design:
---
# 103 — A container is not recreated when a file it reads changes
## What was observed
The control-node, 2026-09-23. The store was given the predecessor's port. The host rewrote the
forge's and the analytics service's environment files with the new port — correctly — and left
both containers running with the old one in their environment. Both lost their database. Both
stayed "up" and healthy-looking for the twenty minutes it took to notice, then answered 502.
A restart did not help: a container reads its `env-file` when it is **created**, not when it
starts, so `docker restart` handed both containers the same stale environment. Only removing them
and letting the host recreate them fixed it.
The host decides whether a container needs recreating by comparing a hash of its declared spec.
The spec names the env file's *path*; the file's *content* is not part of it. So a change that
alters everything the process will see alters nothing the host compares.
## Why it matters beyond this instance
Every module with an `env-file` — which is most of them — has a configuration the host writes and a
container that reads it once. Any change to that configuration that the host applies without
recreating the container is applied to the disk and not to the service. The mesh then reports the
node as running what it was told, because the file is right; only the process is wrong.
The ports move is the obvious trigger, and it is the migration's whole method. But a rotated
credential, a re-provisioned database, a changed binding — anything the host substitutes into a
file a container reads — has the same shape.
## Open questions
- Should the spec hash cover the **content** of every file the container mounts or reads, so a
changed file recreates it — accepting that every such change is a restart of the service?
- Or should the host recreate on content change only for `env-file` and mounted secrets, and leave
bind-mounted data alone — a file the service reads at start versus a directory it reads while
running?
- How does the host report the difference between "the file is right" and "the process has read
it"? Today it cannot, and that is what made this silent.
@@ -0,0 +1,51 @@
---
status: located
opened: 2026-09-23
located-in: [mesh-host cmd/mesh-host]
fixed-by:
amended-design:
---
# 104 — `reconcile` applies the declaration the host was installed with, and refuses nothing
## What was observed
The control-node, 2026-09-23, adopted, twelve modules assigned, two services already cut over.
Chasing why an assignment had not finished, an operator ran the host's own `reconcile` command by
hand.
It did not reconcile the node against the controller's declaration. It applied **the declaration
the host carries** — the genesis bundle: foundation only, *converged*. In order: it recreated the
store, tried to recreate the broker and failed on a held port, wrote the converged base filter to
disk, enabled and started its service, and stopped at the first failing action with "nothing after
it was attempted".
The base filter is the ruleset [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
exists to keep off an adopted node: input policy drop, three ports allowed. It closed the machine
to the internet for about forty-five minutes — every site, the forge's path to its database — and
was removed by hand.
Nothing about the command said any of this would happen. It printed what it did after it did it.
The node's mode was known to the host — it reports "adopted" in every report — and the
declaration it applied said "converged", and no comparison was made.
## Why it matters beyond this instance
The host has two declarations and one command that does not say which it means. The bundle is
right at genesis and stale a minute later; on an adopted node it is actively dangerous, because it
is a converged declaration for a machine that is not converged. A command an operator would
reasonably reach for under pressure — "make the machine match" — is the one that must not be run.
The operator error here was real and is recorded as such. But a design that turns "I ran the
obvious command" into a closed machine has a hole of its own, and the runbook's line "never bypass
the controller" was standing in for a refusal the host should make itself.
## Open questions
- Should `reconcile` refuse outright when the declaration it carries is older than the one the
controller last sent, or says a different mode than the node reports — naming both?
- Should a converged declaration be refused on an adopted node at the point of application,
whatever command delivered it, since the mode is a fact the host already knows?
- Should the host preview before applying from a file — the way `converge` previews — and stop at
the first action *before* running it rather than after?
- Does the bundle need to remain applicable after genesis at all, or should genesis consume it?
@@ -0,0 +1,34 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 105 — The hub of the private network is a placement, not a seat
## What was observed
Reading the registry of a one-node mesh, 2026-09-23. The private network claims a seat,
`the-private-network`, scoped to the **node** — because every node has an interface and a
mesh-scoped seat would refuse the second machine. The fact that matters, *which node is the hub the
others rendezvous at*, is a placement record (`overlay place <node> hub`) and no seat at all.
## Why it matters beyond this instance
The four foundation seats all say the same thing: *there is exactly one of me in this mesh*
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)). "There is
exactly one hub" is that shape. As a placement it is refused by nothing: two nodes can be placed as
hub, and the mesh would compute a graph with two rendezvous points and say nothing.
It also confuses the reading. Asked "who provides the private network", the registry answers with
a node-scoped seat held by every node, which is true and not what was asked.
## Open questions
- Should the hub be a mesh-scoped seat — `the-hub`, or the network's own name — claimed by the
node that is placed there, refused elsewhere?
- Does the per-node seat still say anything once the hub is a seat, or is it the interface's
presence restated?
- What else in the mesh is "exactly one" and recorded as a placement rather than a seat?
@@ -0,0 +1,33 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 106 — The vault claims no seat, so nothing refuses a second one
## What was observed
Reading the registry, 2026-09-23. The vault provides `secret` to the whole mesh and **claims no
seat**. The store, the broker, the controller and the catalogue each claim one, named after
themselves, so that a second claimant is refused at resolution
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)). The vault
was added after that record and did not inherit the rule.
## Why it matters beyond this instance
The vault holds every credential the mesh mints. A second vault, assigned by mistake or by a module
that provides `secret` itself, would answer requirements the first was answering, and nothing in
resolution would object. Of all the components to allow two of silently, this is the one to allow
least.
The rule is already written; this is an instance it was not applied to. Worth asking whether
others were missed the same way — every provider added after 0079.
## Open questions
- A mesh-scoped seat `mesh-vault`, by the 0079 convention — is there any reason not to?
- Should a provider of a mesh-scoped provision be required to claim a seat, or say explicitly that
more than one is allowed, so the omission cannot recur?