Issues 102–106 and ADR 0105 from the core migration
Two birth-address outages and a registry that would have been the third; a container that keeps a stale environment after its file changes; a host command that applied a converged declaration to an adopted node; the hub and the vault without seats. And the decision the operator made under it all: the hub adopts the predecessor's tunnel in place, key and peers and range and port.
This commit is contained in:
+66
@@ -0,0 +1,66 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-23
|
||||
located-in: [mesh-controller cmd/mesh-bootstrap, mesh-controller internal/builder, mesh-controller internal/catalogue]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 102 — An address recorded at genesis or at build does not follow the node's ports
|
||||
|
||||
## What was observed
|
||||
|
||||
The control-node, 2026-09-23, migrating the foundation's store and broker onto the ports the
|
||||
predecessor served them on — the move [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
|
||||
describes: *the ports given become that node's settings, and every place that uses them reads them
|
||||
from there.*
|
||||
|
||||
Most places did. When the store was given 6852, the bindings handed to the forge, the analytics
|
||||
service and the catalogue all said 6852. When the broker was given 5679, its consumers reconnected.
|
||||
|
||||
Three places did not, and each took the mesh down in a different way:
|
||||
|
||||
1. **The controller's own store address.** A secret written at genesis says `127.0.0.1:5432`. The
|
||||
store moved; the controller could not reach its inventory; the mesh was headless.
|
||||
2. **The controller's own broker address.** The same, for `127.0.0.1:5672`. The controller
|
||||
crash-looped and its control queue filled.
|
||||
3. **Every image reference the mesh has built.** Each recorded build reads
|
||||
`<registry-address>:5100/<module>/<artifact>@sha256:…`, and declarations carry that literal. Move the
|
||||
registry and every fresh pull of a mesh image fails — a new node, a recreate after eviction.
|
||||
Found by reading before the move; the first two were found by making it.
|
||||
|
||||
All three are the same fact: an address was **written down** when the port was decided, rather than
|
||||
**read** from the node's settings when it is used. The first two live in genesis secrets; the third
|
||||
lives in stored data, which is worse — it is baked into every build ever recorded.
|
||||
|
||||
Both outages were closed by hand with a forwarder on the old address to the new one. Two such
|
||||
forwarders are holding the control-node's mesh together while this is open. They are not the fix;
|
||||
they are the shape of the bug, made visible.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
A port that is a setting in nine places and a constant in three is a constant. The migration's
|
||||
whole method — take the predecessor's ports one service at a time — depends on every reader
|
||||
following the setting, and the readers that do not are exactly the control plane's own, which is
|
||||
the worst place for them: the failure is headlessness, and headlessness cannot be repaired through
|
||||
the mesh.
|
||||
|
||||
The image reference case will bite any mesh that changes its registry's port, or moves its registry
|
||||
to another node, or ever has two registries. It also means an image is recorded by *where it was
|
||||
pushed* rather than *what it is*: the digest is the identity, the address is a route to it, and the
|
||||
mesh stores them as one string.
|
||||
|
||||
This is the same hole [issue 085](../085-the-packages-port-given-at-genesis-is-not-a-setting/00-report.md)
|
||||
found for the packages port, fixed there for that one reader. It is not one reader; it is a
|
||||
category.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the controller read its own store and broker addresses from the node's settings at start,
|
||||
the way it composes them for every other module — and re-read them when they change, since it is
|
||||
the thing that changes them?
|
||||
- Should a recorded build store the artifact's **digest and path** only, with the registry address
|
||||
composed into the declaration from the node's current settings — so a reference is assembled
|
||||
where it is used, never stored?
|
||||
- Is there a way to find the remaining constants mechanically — every place a foundation port
|
||||
number appears as a literal — rather than one outage at a time?
|
||||
@@ -0,0 +1,45 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-23
|
||||
located-in: [mesh-host internal/apply]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 103 — A container is not recreated when a file it reads changes
|
||||
|
||||
## What was observed
|
||||
|
||||
The control-node, 2026-09-23. The store was given the predecessor's port. The host rewrote the
|
||||
forge's and the analytics service's environment files with the new port — correctly — and left
|
||||
both containers running with the old one in their environment. Both lost their database. Both
|
||||
stayed "up" and healthy-looking for the twenty minutes it took to notice, then answered 502.
|
||||
|
||||
A restart did not help: a container reads its `env-file` when it is **created**, not when it
|
||||
starts, so `docker restart` handed both containers the same stale environment. Only removing them
|
||||
and letting the host recreate them fixed it.
|
||||
|
||||
The host decides whether a container needs recreating by comparing a hash of its declared spec.
|
||||
The spec names the env file's *path*; the file's *content* is not part of it. So a change that
|
||||
alters everything the process will see alters nothing the host compares.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
Every module with an `env-file` — which is most of them — has a configuration the host writes and a
|
||||
container that reads it once. Any change to that configuration that the host applies without
|
||||
recreating the container is applied to the disk and not to the service. The mesh then reports the
|
||||
node as running what it was told, because the file is right; only the process is wrong.
|
||||
|
||||
The ports move is the obvious trigger, and it is the migration's whole method. But a rotated
|
||||
credential, a re-provisioned database, a changed binding — anything the host substitutes into a
|
||||
file a container reads — has the same shape.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the spec hash cover the **content** of every file the container mounts or reads, so a
|
||||
changed file recreates it — accepting that every such change is a restart of the service?
|
||||
- Or should the host recreate on content change only for `env-file` and mounted secrets, and leave
|
||||
bind-mounted data alone — a file the service reads at start versus a directory it reads while
|
||||
running?
|
||||
- How does the host report the difference between "the file is right" and "the process has read
|
||||
it"? Today it cannot, and that is what made this silent.
|
||||
@@ -0,0 +1,51 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-23
|
||||
located-in: [mesh-host cmd/mesh-host]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 104 — `reconcile` applies the declaration the host was installed with, and refuses nothing
|
||||
|
||||
## What was observed
|
||||
|
||||
The control-node, 2026-09-23, adopted, twelve modules assigned, two services already cut over.
|
||||
Chasing why an assignment had not finished, an operator ran the host's own `reconcile` command by
|
||||
hand.
|
||||
|
||||
It did not reconcile the node against the controller's declaration. It applied **the declaration
|
||||
the host carries** — the genesis bundle: foundation only, *converged*. In order: it recreated the
|
||||
store, tried to recreate the broker and failed on a held port, wrote the converged base filter to
|
||||
disk, enabled and started its service, and stopped at the first failing action with "nothing after
|
||||
it was attempted".
|
||||
|
||||
The base filter is the ruleset [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
|
||||
exists to keep off an adopted node: input policy drop, three ports allowed. It closed the machine
|
||||
to the internet for about forty-five minutes — every site, the forge's path to its database — and
|
||||
was removed by hand.
|
||||
|
||||
Nothing about the command said any of this would happen. It printed what it did after it did it.
|
||||
The node's mode was known to the host — it reports "adopted" in every report — and the
|
||||
declaration it applied said "converged", and no comparison was made.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
The host has two declarations and one command that does not say which it means. The bundle is
|
||||
right at genesis and stale a minute later; on an adopted node it is actively dangerous, because it
|
||||
is a converged declaration for a machine that is not converged. A command an operator would
|
||||
reasonably reach for under pressure — "make the machine match" — is the one that must not be run.
|
||||
|
||||
The operator error here was real and is recorded as such. But a design that turns "I ran the
|
||||
obvious command" into a closed machine has a hole of its own, and the runbook's line "never bypass
|
||||
the controller" was standing in for a refusal the host should make itself.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should `reconcile` refuse outright when the declaration it carries is older than the one the
|
||||
controller last sent, or says a different mode than the node reports — naming both?
|
||||
- Should a converged declaration be refused on an adopted node at the point of application,
|
||||
whatever command delivered it, since the mode is a fact the host already knows?
|
||||
- Should the host preview before applying from a file — the way `converge` previews — and stop at
|
||||
the first action *before* running it rather than after?
|
||||
- Does the bundle need to remain applicable after genesis at all, or should genesis consume it?
|
||||
@@ -0,0 +1,34 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-23
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 105 — The hub of the private network is a placement, not a seat
|
||||
|
||||
## What was observed
|
||||
|
||||
Reading the registry of a one-node mesh, 2026-09-23. The private network claims a seat,
|
||||
`the-private-network`, scoped to the **node** — because every node has an interface and a
|
||||
mesh-scoped seat would refuse the second machine. The fact that matters, *which node is the hub the
|
||||
others rendezvous at*, is a placement record (`overlay place <node> hub`) and no seat at all.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
The four foundation seats all say the same thing: *there is exactly one of me in this mesh*
|
||||
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)). "There is
|
||||
exactly one hub" is that shape. As a placement it is refused by nothing: two nodes can be placed as
|
||||
hub, and the mesh would compute a graph with two rendezvous points and say nothing.
|
||||
|
||||
It also confuses the reading. Asked "who provides the private network", the registry answers with
|
||||
a node-scoped seat held by every node, which is true and not what was asked.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the hub be a mesh-scoped seat — `the-hub`, or the network's own name — claimed by the
|
||||
node that is placed there, refused elsewhere?
|
||||
- Does the per-node seat still say anything once the hub is a seat, or is it the interface's
|
||||
presence restated?
|
||||
- What else in the mesh is "exactly one" and recorded as a placement rather than a seat?
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-23
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 106 — The vault claims no seat, so nothing refuses a second one
|
||||
|
||||
## What was observed
|
||||
|
||||
Reading the registry, 2026-09-23. The vault provides `secret` to the whole mesh and **claims no
|
||||
seat**. The store, the broker, the controller and the catalogue each claim one, named after
|
||||
themselves, so that a second claimant is refused at resolution
|
||||
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)). The vault
|
||||
was added after that record and did not inherit the rule.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
The vault holds every credential the mesh mints. A second vault, assigned by mistake or by a module
|
||||
that provides `secret` itself, would answer requirements the first was answering, and nothing in
|
||||
resolution would object. Of all the components to allow two of silently, this is the one to allow
|
||||
least.
|
||||
|
||||
The rule is already written; this is an instance it was not applied to. Worth asking whether
|
||||
others were missed the same way — every provider added after 0079.
|
||||
|
||||
## Open questions
|
||||
|
||||
- A mesh-scoped seat `mesh-vault`, by the 0079 convention — is there any reason not to?
|
||||
- Should a provider of a mesh-scoped provision be required to claim a seat, or say explicitly that
|
||||
more than one is allowed, so the omission cannot recur?
|
||||
Reference in New Issue
Block a user