Merge pull request 'Issue 102 resolved and verified on the machine; issue 097's orphan was on the host network' (#94) from issues/102-resolved-and-097-worse into main

This commit was merged in pull request #94.
This commit is contained in:
2026-09-23 22:15:49 +00:00
3 changed files with 68 additions and 6 deletions
@@ -35,10 +35,17 @@ Every other container the mesh made on this machine is declared; this is the onl
It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing
stops it, and nothing reports it.
Harmless in this instance by luck — the stranded container published no ports and held an
anonymous volume rather than the service's data — and that luck is the point. Had the rename gone
the other way, two containers of the same service would have run against one data directory, or
the old one would have kept the port the new one needed and the new one would have failed to bind.
Harmless in this instance by luck — and less harmless than it first looked. **Addendum,
2026-09-24**, after reading it properly rather than glancing at it: the stranded container was
running on the **host's own network**, so it was listening on a port on every interface of the
machine, and it held open connections to the mesh's store. It reached the store through a forwarder
that had been put in front of an old address for an unrelated reason, which is the only thing that
kept the two facts from meeting sooner.
What saved it was that its configuration named its own former database rather than the one the
service now uses, so no data was at risk. Nothing in the mesh arranged that. Had the rename gone
the other way — or had the database name not changed with it — two versions of one service would
have been writing to one database, one of them a version older than the schema.
## Why it matters beyond this instance
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-bootstrap, mesh-controller internal/builder, mesh-controller internal/catalogue]
fixed-by:
fixed-by: mesh-controller — an address is read from the node's settings where it is used, never recorded with a port
amended-design:
---
@@ -0,0 +1,55 @@
# Diagnosis and resolution — 2026-09-24
## One fault in three readers
Each was the same mistake: an address was **written down** when a port was decided, rather than
**read** where it is used.
- The control plane's own store and broker addresses were whole connection strings, sealed at
genesis. Nothing could rewrite them, so moving the store and the broker onto the predecessor's
ports left the control plane dialling numbers nothing listened on — twice, and each time the mesh
was headless until a forwarder was put in front of the old address by hand.
- Every recorded build carried the registry's address in the image reference. Found by reading
rather than by breaking, because the registry had not moved yet.
## What it is now
A port that is a node's setting is answered where it is asked. The control plane's own manifest asks
for it — one value per context, per address — and the controller fills it from the same place every
consumer's binding comes from: what the node was given. An address the mesh does not need to add
anything to comes back empty, and the value sealed at genesis stands; there is no second source of
truth to disagree.
A recorded build no longer holds an address at all. It names the artifact — its module, its name and
its digest — and the address is composed when a declaration is built, from the node that holds the
artifact store and the port that node was given. A reference is assembled where it is used.
## Verified on the machine, not only in tests
Rolled out in two steps, deliberately: the code first, with the manifest unchanged, so the running
control plane composed a declaration it fully understood; then the manifest, composed by the new
binary, which filled it in. The control plane came back holding **6852** for all three store
contexts, **5679** for the bus it consumes, the broker's management port as the node gave it, and an
empty value for the one address the node never moved — exactly the rule. Its own image reference,
recorded address-free, was routed back to the node's registry port at compose.
Then both forwarders were stopped, and with neither in place: the mesh reports healthy, a push
round-trips, a broker account is minted through the management port, and a build runs. **The
addresses follow.**
## Two things the review caught before they could happen
- A control plane older than the manifest it is handed would have passed the unfilled request
through as a literal, and the reader would have refused it as "not a port" — headless, by exactly
the mechanism this issue is about. The reader now treats an unfilled request as nothing said and
says so on the way past; the split rollout above is the belt to that brace.
- On a machine with nothing on the private network yet — a first node, mid-genesis — there is no
address for the artifact store at all, and the composed reference would have been refused. The
store is now reached by loopback in that one case, which is the only case where the machine
composing and the machine holding are the same.
## What this does not close
Two references pinned at genesis — the control plane's own image and the builder's — name a
single-segment repository with no build behind them, so nothing re-routes them until each is rebuilt
through the mesh. The registry has not moved and is not going to; when it does, rebuild both first.