Merge pull request 'Issue 094 diagnosed and resolved; 096, 097 and 098 opened from what it uncovered' (#84) from issues/094-diagnosis-and-096 into main

This commit was merged in pull request #84.
This commit is contained in:
2026-09-23 00:37:27 +00:00
5 changed files with 277 additions and 2 deletions
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-23
located-in: [mesh-controller internal/catalogue/filtering.go, mesh-controller cmd/mesh-controller/plan.go]
fixed-by:
fixed-by: mesh-controller — a node may move a port a module publishes as a mapping's machine side
amended-design:
---
@@ -0,0 +1,80 @@
# Diagnosis — 2026-09-23
## One cause, two symptoms
Both halves of the report come from the same blind spot, in two different places.
**The check that refused the setting** collected, for each entry a container publishes, only its
**last** segment. A short form publishes one number, which is the container's port and the
machine's at once, so reading the last segment is right. A long form — a machine port mapped to a
different port inside the container — has two, and the last segment is the container's. So the
machine side, the number the module names in `listens`, in `serves` and in every port
substitution, was not in the set of ports the module was considered to publish, and giving it one
was refused as naming nothing.
**The lookup that made naming the other half useless** reads the same map under two different
keys. Everything derived from what the module *declares* — the filter, the openings, the guard,
what a consumer is told, a port substituted into a file — looks the machine side up. The one
reader that rewrites the mapping handed to the container runtime looks up the container's port.
For a short form those are the same number and no one notices. For a long form they differ, so
exactly one reader ever found the entry: keyed the only way the check allowed, the container's
mapping moved and nothing else did, leaving a firewall, a set of openings and a consumer all
pointing at a port the software had left.
That is also the third symptom in the report, seen from the other end. No opening was derived for
the moved port because the opening is derived from the declared port, which still read as the old
number; and a per-node `expose` could not rescue it, because `expose` keys on the same declared
port — it widens the opening on the port nothing is on any more, and naming the real machine port
is refused as a port the module does not listen on.
## What was ruled out
The openings derivation itself. Composed with no setting at all, a module publishing a long-form
mapping on an adopted node does get its opening, on the machine side, forwarded to the container's
port. A regression test now records that, deliberately passing before the fix as well as after, so
the next reader does not go looking there.
## Answering the first open question
**Either end names the mapping, and both answer.** A module publishing `"2222:22"` may reasonably
say it listens on the port its software uses or on the port the machine serves; the mesh accepts
whichever the setting names and returns the machine port under both, so every reader finds the
same number under the key it happens to hold. This keeps working what already worked — the
container's end was the only key the old check accepted — and makes it mean the same thing.
Ambiguity is refused where it is real: one number naming two **different** mappings, and the two
ends of **one** mapping given two different machine ports. Naming both ends of one mapping with
the same number was already refused, by the rule that a machine port has one holder.
## What this does not close
The second and third open questions stand, and they are the push-blocking half: the setting is
still **stored** without a manifest in view, so a key that names nothing a module publishes is
accepted where it is set and refused at composition — where it stops the node being told anything
at all, rather than stopping that module. The fix removes one reason a key could be wrong; it does
not remove the shape of the failure. Carried to [issue 096](../096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md).
## Resolved — 2026-09-23
A given port names either end of a mapping and is answered **once**, under the end the module
names in its own `listens` — the number the plan, the filter, the openings, the guard and the
consumer all ask for. A number that names two of a module's mappings is refused, whether the
setting used it or the answer would be filed under it.
The first pass filed the answer under both ends. Review found that this is wrong wherever two
mappings share a number: the second key lands where another mapping's reader looks, the later
write wins, and a module publishing `8080:80` beside `9090:8080` had an explicit setting silently
replaced — the container moving to one number while the filter, the opening and the consumer kept
another. Reintroducing the fault, inside the fix for it. Two mappings onto one machine port came
with it, where before the setting had been safely refused. Neither was reachable against the
catalogue as it stands; both were silent, which is worse than reachable.
**Verified on the machine**, not only in tests: the forge was given `{"2222": 222}`, its container
now publishes `222:22`, the opening derived for it is `tcp-222-forwarded` to 22 from anywhere —
the node's `expose` for that port carried across the move — and nothing is opened at the number it
left. Git over ssh answers from outside again, on the port the forge's own configuration has
advertised in every clone URL all along.
Checked from the other side as well, and it matters: `SSH_PORT` in the forge's configuration says
222. While the port could not be moved, the mesh published 2222 and the forge told everyone 222.
A service can be wrong about itself without anything failing.
@@ -0,0 +1,60 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 096 — A setting that cannot work is accepted where it is set, and stops the node where it is read
## What was observed
On the control-node during the first module's migration, 2026-09-23, and reproduced since against
the controller's own tests.
A per-node port setting was accepted and stored. Composing that node's declaration then failed on
it, and because a node is told everything or nothing, **every push to that machine was refused**
until somebody found the setting and removed it. The message named a port number. It did not name
the setting, the layer it was stored in, or the node it had stopped; nothing said that a stored
statement was the reason the machine had gone quiet.
The specific reason that setting could not work is [issue
094](../094-a-port-published-as-the-machine-side-cannot-be-moved/00-report.md), and it is fixed.
This issue is the shape that surrounded it, which is not:
- **The place that stores a setting has no manifest in view.** It checks what it can without one —
that a value is a port, that ssh keeps its own, that two modules on the machine do not claim the
same machine port — and leaves anything that needs the module's own declaration to composition.
So a key naming a port the module does not publish is stored today, and a typo is stored today,
and both are found later, from the far end.
- **Composition fails the node, not the module.** One unusable statement about one module refuses
the whole declaration, so the other modules on that machine stop being told anything either —
including modules that were fine before the setting existed.
## Why it matters beyond this instance
The two together turn a typo into an outage of the control link for a machine, at a distance from
the thing that caused it. The delay is the damage: a refusal at the moment of setting is a
correction, and the same refusal a push later is a machine nobody can talk to, found by whoever
next notices it is not being updated.
It generalises past ports. Any setting whose validity depends on the module's declaration has this
shape — a value that is well-formed on its own and impossible against the manifest. Ports are
merely where the mesh found it first, because migrating a service is when settings get written.
It also touches a rule the mesh states elsewhere: what refuses, refuses early and by name. A
refusal that names a number rather than the statement that produced it cannot be acted on without
knowing the code.
## Open questions
- Should storing a setting compose it against the module's manifest first — and if so, against
which version, given the catalogue moves and a manifest that was right when the setting was
written may not be later?
- Or should the guard be at the far end: composition refuses that **module** and sends the rest of
the node, so an impossible statement costs one service and not the machine?
- Either way, what does a refusal have to name — the node, the module, the layer and the key — for
an operator to undo it without reading the source?
- Is there anything a node must never be pushed without, such that sending a partial declaration is
worse than sending none?
@@ -0,0 +1,70 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 097 — A resource whose target changes leaves the old one behind, running
## What was observed
On the control-node, 2026-09-23, four hours after the forge's module was first assigned.
A module resource named its container implicitly: with no name of its own, the host derived one
from the module and the resource's id. The manifest then gained an explicit name, because the
module had to take over a container the predecessor already ran under that name
([issue 090](../090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md)).
The host applied the change by raising a container under the new name — and left the one under the
old name **running**.
The host's own record shows why. It keeps one entry per resource id, and that entry holds the
resource's *current* target:
```
{"id": "gitea.server", "type": "container", "origin": "declared",
"holds": [2999, 2222], "target": "gitea", "applied_at": "..."}
```
There is no entry naming the old container. Rewriting the record on a target change is what
erases the only trace of what the host must now remove, so the old one cannot be found by the
thing that would have removed it.
Every other container the mesh made on this machine is declared; this is the only one that is not.
It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing
stops it, and nothing reports it.
Harmless in this instance by luck — the stranded container published no ports and held an
anonymous volume rather than the service's data — and that luck is the point. Had the rename gone
the other way, two containers of the same service would have run against one data directory, or
the old one would have kept the port the new one needed and the new one would have failed to bind.
## Why it matters beyond this instance
The mesh's promise is that a machine runs what it was told and nothing else, and that a module
removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record
that remembers only the current target breaks that for **any** resource whose target moves — and a
file is worse than a container, because a stale configuration file at the old path is still read
by whatever reads that path, silently, with no process to notice running twice.
Renaming is not exotic. It happens exactly when a module is taught to take over something that
already exists, which is every module in a migration.
It also crosses [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md):
on an adopted node the host must distinguish what it wrote from what it found, because what it
found is held and never removed. A resource that changes target turns something the host wrote
into something no record claims — which, on the next machine, is indistinguishable from something
found, and so would be kept for ever on purpose.
## Open questions
- Should the record keep every target a resource has had, and the host remove the ones it no
longer declares — and if so, for how long, given a record is also how the host knows what it may
destroy?
- Should the host refuse a target change outright, requiring the old resource to be removed by a
declaration that still names it before a new one may take the name?
- Does the same hole exist for a resource whose **id** changes while the target stays, and for a
module unassigned between the two declarations?
- What reports this? Nothing on the machine currently answers "what is running here that the mesh
did not ask for", which is the question that would have found this in seconds.
@@ -0,0 +1,65 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 098 — Taking a module replaces a configuration nobody compared, and the part that was the installation's is gone
## What was observed
Preparing the second module's cutover on the control-node, 2026-09-23. Found by reading, before
anything was taken — which is the only reason it is a report and not an incident.
The module declares the service's whole configuration file, at the same path the machine already
has one. On an adopted node that file is *found*, so it is held until the module is taken, and the
host records the original first — all exactly as
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md) says.
The two files are not the same. The machine's says, in its own words, that packages under one
private scope are **stored locally and never proxied upstream**; every other package is proxied.
The catalogue's file has the second rule and not the first. Taking the module would therefore
replace a policy statement with its absence: the private packages would start resolving through a
public upstream that has never heard of them.
Nothing in the cutover says so. `take` replaces the held file and reports that it did. There is no
step between "held" and "replaced" at which the two contents are put side by side, and the
operator is not asked. The whole-node preview names *which* held files a flip will replace; it does
not say **how they differ**, and the per-module cutover — the step that actually does it — previews
nothing at all.
The original is kept, so this is recoverable. It is not detectable: the service starts, answers,
and serves the wrong thing.
## Why it matters beyond this instance
A configuration file is where an installation keeps what is true about *it* — which packages are
private, which realm is trusted, which paths are exceptions. A catalogue module is, correctly, the
same for every mesh. Declaring the file whole makes the second overwrite the first, and the
migration is precisely when every such file changes hands.
There is no mechanism to carry the difference. A module's file content is a fixed string: no
setting is substituted into it, and the mesh's way of writing *into* a file it does not own whole
([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)) reads and
writes JSON, so a service configured in any other language cannot use it. The two remaining
options are both wrong: put one installation's private scope into a catalogue every mesh shares,
or accept the loss.
The silence is the worse half. The mesh's rule is that each step says what it changes before it
changes it; here a step changes the meaning of a running service and says only that a file was
written.
## Open questions
- Should `take` refuse, or ask, when the held file it would replace differs from what it declares —
and show the difference? What makes a difference acceptable enough to pass silently?
- Should a module be able to declare a configuration **partially**, in the service's own language,
and if so which languages must the host speak — or should such a file never be declared whole,
and always merged?
- Where does an installation's own policy live, if not in the catalogue and not in the file the
catalogue overwrites? A settings layer substituted into content would answer it; nothing
substitutes settings into content today.
- Is the kept original enough of an answer, given nothing restores it and nothing points at it
when the service starts behaving differently?