Issues 096 and 097, and 094 diagnosed: a setting stored where it cannot work, and a resource that changed target

094's cause is one blind spot read from two ends, written up in its diagnosis; the fix
answers the first open question and not the other two, which become 096. 097 was found
looking at what the forge's cutover left running.
This commit is contained in:
2026-09-23 02:20:48 +02:00
parent 1b8e5043ee
commit 235b9ea0e5
3 changed files with 185 additions and 0 deletions
@@ -0,0 +1,55 @@
# Diagnosis — 2026-09-23
## One cause, two symptoms
Both halves of the report come from the same blind spot, in two different places.
**The check that refused the setting** collected, for each entry a container publishes, only its
**last** segment. A short form publishes one number, which is the container's port and the
machine's at once, so reading the last segment is right. A long form — a machine port mapped to a
different port inside the container — has two, and the last segment is the container's. So the
machine side, the number the module names in `listens`, in `serves` and in every port
substitution, was not in the set of ports the module was considered to publish, and giving it one
was refused as naming nothing.
**The lookup that made naming the other half useless** reads the same map under two different
keys. Everything derived from what the module *declares* — the filter, the openings, the guard,
what a consumer is told, a port substituted into a file — looks the machine side up. The one
reader that rewrites the mapping handed to the container runtime looks up the container's port.
For a short form those are the same number and no one notices. For a long form they differ, so
exactly one reader ever found the entry: keyed the only way the check allowed, the container's
mapping moved and nothing else did, leaving a firewall, a set of openings and a consumer all
pointing at a port the software had left.
That is also the third symptom in the report, seen from the other end. No opening was derived for
the moved port because the opening is derived from the declared port, which still read as the old
number; and a per-node `expose` could not rescue it, because `expose` keys on the same declared
port — it widens the opening on the port nothing is on any more, and naming the real machine port
is refused as a port the module does not listen on.
## What was ruled out
The openings derivation itself. Composed with no setting at all, a module publishing a long-form
mapping on an adopted node does get its opening, on the machine side, forwarded to the container's
port. A regression test now records that, deliberately passing before the fix as well as after, so
the next reader does not go looking there.
## Answering the first open question
**Either end names the mapping, and both answer.** A module publishing `"2222:22"` may reasonably
say it listens on the port its software uses or on the port the machine serves; the mesh accepts
whichever the setting names and returns the machine port under both, so every reader finds the
same number under the key it happens to hold. This keeps working what already worked — the
container's end was the only key the old check accepted — and makes it mean the same thing.
Ambiguity is refused where it is real: one number naming two **different** mappings, and the two
ends of **one** mapping given two different machine ports. Naming both ends of one mapping with
the same number was already refused, by the rule that a machine port has one holder.
## What this does not close
The second and third open questions stand, and they are the push-blocking half: the setting is
still **stored** without a manifest in view, so a key that names nothing a module publishes is
accepted where it is set and refused at composition — where it stops the node being told anything
at all, rather than stopping that module. The fix removes one reason a key could be wrong; it does
not remove the shape of the failure. Carried to [issue 096](../096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md).
@@ -0,0 +1,60 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 096 — A setting that cannot work is accepted where it is set, and stops the node where it is read
## What was observed
On the control-node during the first module's migration, 2026-09-23, and reproduced since against
the controller's own tests.
A per-node port setting was accepted and stored. Composing that node's declaration then failed on
it, and because a node is told everything or nothing, **every push to that machine was refused**
until somebody found the setting and removed it. The message named a port number. It did not name
the setting, the layer it was stored in, or the node it had stopped; nothing said that a stored
statement was the reason the machine had gone quiet.
The specific reason that setting could not work is [issue
094](../094-a-port-published-as-the-machine-side-cannot-be-moved/00-report.md), and it is fixed.
This issue is the shape that surrounded it, which is not:
- **The place that stores a setting has no manifest in view.** It checks what it can without one —
that a value is a port, that ssh keeps its own, that two modules on the machine do not claim the
same machine port — and leaves anything that needs the module's own declaration to composition.
So a key naming a port the module does not publish is stored today, and a typo is stored today,
and both are found later, from the far end.
- **Composition fails the node, not the module.** One unusable statement about one module refuses
the whole declaration, so the other modules on that machine stop being told anything either —
including modules that were fine before the setting existed.
## Why it matters beyond this instance
The two together turn a typo into an outage of the control link for a machine, at a distance from
the thing that caused it. The delay is the damage: a refusal at the moment of setting is a
correction, and the same refusal a push later is a machine nobody can talk to, found by whoever
next notices it is not being updated.
It generalises past ports. Any setting whose validity depends on the module's declaration has this
shape — a value that is well-formed on its own and impossible against the manifest. Ports are
merely where the mesh found it first, because migrating a service is when settings get written.
It also touches a rule the mesh states elsewhere: what refuses, refuses early and by name. A
refusal that names a number rather than the statement that produced it cannot be acted on without
knowing the code.
## Open questions
- Should storing a setting compose it against the module's manifest first — and if so, against
which version, given the catalogue moves and a manifest that was right when the setting was
written may not be later?
- Or should the guard be at the far end: composition refuses that **module** and sends the rest of
the node, so an impossible statement costs one service and not the machine?
- Either way, what does a refusal have to name — the node, the module, the layer and the key — for
an operator to undo it without reading the source?
- Is there anything a node must never be pushed without, such that sending a partial declaration is
worse than sending none?
@@ -0,0 +1,70 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 097 — A resource whose target changes leaves the old one behind, running
## What was observed
On the control-node, 2026-09-23, four hours after the forge's module was first assigned.
A module resource named its container implicitly: with no name of its own, the host derived one
from the module and the resource's id. The manifest then gained an explicit name, because the
module had to take over a container the predecessor already ran under that name
([issue 090](../090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md)).
The host applied the change by raising a container under the new name — and left the one under the
old name **running**.
The host's own record shows why. It keeps one entry per resource id, and that entry holds the
resource's *current* target:
```
{"id": "gitea.server", "type": "container", "origin": "declared",
"holds": [2999, 2222], "target": "gitea", "applied_at": "..."}
```
There is no entry naming the old container. Rewriting the record on a target change is what
erases the only trace of what the host must now remove, so the old one cannot be found by the
thing that would have removed it.
Every other container the mesh made on this machine is declared; this is the only one that is not.
It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing
stops it, and nothing reports it.
Harmless in this instance by luck — the stranded container published no ports and held an
anonymous volume rather than the service's data — and that luck is the point. Had the rename gone
the other way, two containers of the same service would have run against one data directory, or
the old one would have kept the port the new one needed and the new one would have failed to bind.
## Why it matters beyond this instance
The mesh's promise is that a machine runs what it was told and nothing else, and that a module
removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record
that remembers only the current target breaks that for **any** resource whose target moves — and a
file is worse than a container, because a stale configuration file at the old path is still read
by whatever reads that path, silently, with no process to notice running twice.
Renaming is not exotic. It happens exactly when a module is taught to take over something that
already exists, which is every module in a migration.
It also crosses [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md):
on an adopted node the host must distinguish what it wrote from what it found, because what it
found is held and never removed. A resource that changes target turns something the host wrote
into something no record claims — which, on the next machine, is indistinguishable from something
found, and so would be kept for ever on purpose.
## Open questions
- Should the record keep every target a resource has had, and the host remove the ones it no
longer declares — and if so, for how long, given a record is also how the host knows what it may
destroy?
- Should the host refuse a target change outright, requiring the old resource to be removed by a
declaration that still names it before a new one may take the name?
- Does the same hole exist for a resource whose **id** changes while the target stays, and for a
module unassigned between the two declarations?
- What reports this? Nothing on the machine currently answers "what is running here that the mesh
did not ask for", which is the question that would have found this in seconds.