Issues 096 and 097, and 094 diagnosed: a setting stored where it cannot work, and a resource that changed target
094's cause is one blind spot read from two ends, written up in its diagnosis; the fix answers the first open question and not the other two, which become 096. 097 was found looking at what the forge's cutover left running.
This commit is contained in:
@@ -0,0 +1,55 @@
|
|||||||
|
# Diagnosis — 2026-09-23
|
||||||
|
|
||||||
|
## One cause, two symptoms
|
||||||
|
|
||||||
|
Both halves of the report come from the same blind spot, in two different places.
|
||||||
|
|
||||||
|
**The check that refused the setting** collected, for each entry a container publishes, only its
|
||||||
|
**last** segment. A short form publishes one number, which is the container's port and the
|
||||||
|
machine's at once, so reading the last segment is right. A long form — a machine port mapped to a
|
||||||
|
different port inside the container — has two, and the last segment is the container's. So the
|
||||||
|
machine side, the number the module names in `listens`, in `serves` and in every port
|
||||||
|
substitution, was not in the set of ports the module was considered to publish, and giving it one
|
||||||
|
was refused as naming nothing.
|
||||||
|
|
||||||
|
**The lookup that made naming the other half useless** reads the same map under two different
|
||||||
|
keys. Everything derived from what the module *declares* — the filter, the openings, the guard,
|
||||||
|
what a consumer is told, a port substituted into a file — looks the machine side up. The one
|
||||||
|
reader that rewrites the mapping handed to the container runtime looks up the container's port.
|
||||||
|
For a short form those are the same number and no one notices. For a long form they differ, so
|
||||||
|
exactly one reader ever found the entry: keyed the only way the check allowed, the container's
|
||||||
|
mapping moved and nothing else did, leaving a firewall, a set of openings and a consumer all
|
||||||
|
pointing at a port the software had left.
|
||||||
|
|
||||||
|
That is also the third symptom in the report, seen from the other end. No opening was derived for
|
||||||
|
the moved port because the opening is derived from the declared port, which still read as the old
|
||||||
|
number; and a per-node `expose` could not rescue it, because `expose` keys on the same declared
|
||||||
|
port — it widens the opening on the port nothing is on any more, and naming the real machine port
|
||||||
|
is refused as a port the module does not listen on.
|
||||||
|
|
||||||
|
## What was ruled out
|
||||||
|
|
||||||
|
The openings derivation itself. Composed with no setting at all, a module publishing a long-form
|
||||||
|
mapping on an adopted node does get its opening, on the machine side, forwarded to the container's
|
||||||
|
port. A regression test now records that, deliberately passing before the fix as well as after, so
|
||||||
|
the next reader does not go looking there.
|
||||||
|
|
||||||
|
## Answering the first open question
|
||||||
|
|
||||||
|
**Either end names the mapping, and both answer.** A module publishing `"2222:22"` may reasonably
|
||||||
|
say it listens on the port its software uses or on the port the machine serves; the mesh accepts
|
||||||
|
whichever the setting names and returns the machine port under both, so every reader finds the
|
||||||
|
same number under the key it happens to hold. This keeps working what already worked — the
|
||||||
|
container's end was the only key the old check accepted — and makes it mean the same thing.
|
||||||
|
|
||||||
|
Ambiguity is refused where it is real: one number naming two **different** mappings, and the two
|
||||||
|
ends of **one** mapping given two different machine ports. Naming both ends of one mapping with
|
||||||
|
the same number was already refused, by the rule that a machine port has one holder.
|
||||||
|
|
||||||
|
## What this does not close
|
||||||
|
|
||||||
|
The second and third open questions stand, and they are the push-blocking half: the setting is
|
||||||
|
still **stored** without a manifest in view, so a key that names nothing a module publishes is
|
||||||
|
accepted where it is set and refused at composition — where it stops the node being told anything
|
||||||
|
at all, rather than stopping that module. The fix removes one reason a key could be wrong; it does
|
||||||
|
not remove the shape of the failure. Carried to [issue 096](../096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md).
|
||||||
@@ -0,0 +1,60 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-23
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 096 — A setting that cannot work is accepted where it is set, and stops the node where it is read
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
On the control-node during the first module's migration, 2026-09-23, and reproduced since against
|
||||||
|
the controller's own tests.
|
||||||
|
|
||||||
|
A per-node port setting was accepted and stored. Composing that node's declaration then failed on
|
||||||
|
it, and because a node is told everything or nothing, **every push to that machine was refused**
|
||||||
|
until somebody found the setting and removed it. The message named a port number. It did not name
|
||||||
|
the setting, the layer it was stored in, or the node it had stopped; nothing said that a stored
|
||||||
|
statement was the reason the machine had gone quiet.
|
||||||
|
|
||||||
|
The specific reason that setting could not work is [issue
|
||||||
|
094](../094-a-port-published-as-the-machine-side-cannot-be-moved/00-report.md), and it is fixed.
|
||||||
|
This issue is the shape that surrounded it, which is not:
|
||||||
|
|
||||||
|
- **The place that stores a setting has no manifest in view.** It checks what it can without one —
|
||||||
|
that a value is a port, that ssh keeps its own, that two modules on the machine do not claim the
|
||||||
|
same machine port — and leaves anything that needs the module's own declaration to composition.
|
||||||
|
So a key naming a port the module does not publish is stored today, and a typo is stored today,
|
||||||
|
and both are found later, from the far end.
|
||||||
|
- **Composition fails the node, not the module.** One unusable statement about one module refuses
|
||||||
|
the whole declaration, so the other modules on that machine stop being told anything either —
|
||||||
|
including modules that were fine before the setting existed.
|
||||||
|
|
||||||
|
## Why it matters beyond this instance
|
||||||
|
|
||||||
|
The two together turn a typo into an outage of the control link for a machine, at a distance from
|
||||||
|
the thing that caused it. The delay is the damage: a refusal at the moment of setting is a
|
||||||
|
correction, and the same refusal a push later is a machine nobody can talk to, found by whoever
|
||||||
|
next notices it is not being updated.
|
||||||
|
|
||||||
|
It generalises past ports. Any setting whose validity depends on the module's declaration has this
|
||||||
|
shape — a value that is well-formed on its own and impossible against the manifest. Ports are
|
||||||
|
merely where the mesh found it first, because migrating a service is when settings get written.
|
||||||
|
|
||||||
|
It also touches a rule the mesh states elsewhere: what refuses, refuses early and by name. A
|
||||||
|
refusal that names a number rather than the statement that produced it cannot be acted on without
|
||||||
|
knowing the code.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Should storing a setting compose it against the module's manifest first — and if so, against
|
||||||
|
which version, given the catalogue moves and a manifest that was right when the setting was
|
||||||
|
written may not be later?
|
||||||
|
- Or should the guard be at the far end: composition refuses that **module** and sends the rest of
|
||||||
|
the node, so an impossible statement costs one service and not the machine?
|
||||||
|
- Either way, what does a refusal have to name — the node, the module, the layer and the key — for
|
||||||
|
an operator to undo it without reading the source?
|
||||||
|
- Is there anything a node must never be pushed without, such that sending a partial declaration is
|
||||||
|
worse than sending none?
|
||||||
@@ -0,0 +1,70 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-23
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 097 — A resource whose target changes leaves the old one behind, running
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
On the control-node, 2026-09-23, four hours after the forge's module was first assigned.
|
||||||
|
|
||||||
|
A module resource named its container implicitly: with no name of its own, the host derived one
|
||||||
|
from the module and the resource's id. The manifest then gained an explicit name, because the
|
||||||
|
module had to take over a container the predecessor already ran under that name
|
||||||
|
([issue 090](../090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md)).
|
||||||
|
The host applied the change by raising a container under the new name — and left the one under the
|
||||||
|
old name **running**.
|
||||||
|
|
||||||
|
The host's own record shows why. It keeps one entry per resource id, and that entry holds the
|
||||||
|
resource's *current* target:
|
||||||
|
|
||||||
|
```
|
||||||
|
{"id": "gitea.server", "type": "container", "origin": "declared",
|
||||||
|
"holds": [2999, 2222], "target": "gitea", "applied_at": "..."}
|
||||||
|
```
|
||||||
|
|
||||||
|
There is no entry naming the old container. Rewriting the record on a target change is what
|
||||||
|
erases the only trace of what the host must now remove, so the old one cannot be found by the
|
||||||
|
thing that would have removed it.
|
||||||
|
|
||||||
|
Every other container the mesh made on this machine is declared; this is the only one that is not.
|
||||||
|
It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing
|
||||||
|
stops it, and nothing reports it.
|
||||||
|
|
||||||
|
Harmless in this instance by luck — the stranded container published no ports and held an
|
||||||
|
anonymous volume rather than the service's data — and that luck is the point. Had the rename gone
|
||||||
|
the other way, two containers of the same service would have run against one data directory, or
|
||||||
|
the old one would have kept the port the new one needed and the new one would have failed to bind.
|
||||||
|
|
||||||
|
## Why it matters beyond this instance
|
||||||
|
|
||||||
|
The mesh's promise is that a machine runs what it was told and nothing else, and that a module
|
||||||
|
removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record
|
||||||
|
that remembers only the current target breaks that for **any** resource whose target moves — and a
|
||||||
|
file is worse than a container, because a stale configuration file at the old path is still read
|
||||||
|
by whatever reads that path, silently, with no process to notice running twice.
|
||||||
|
|
||||||
|
Renaming is not exotic. It happens exactly when a module is taught to take over something that
|
||||||
|
already exists, which is every module in a migration.
|
||||||
|
|
||||||
|
It also crosses [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md):
|
||||||
|
on an adopted node the host must distinguish what it wrote from what it found, because what it
|
||||||
|
found is held and never removed. A resource that changes target turns something the host wrote
|
||||||
|
into something no record claims — which, on the next machine, is indistinguishable from something
|
||||||
|
found, and so would be kept for ever on purpose.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Should the record keep every target a resource has had, and the host remove the ones it no
|
||||||
|
longer declares — and if so, for how long, given a record is also how the host knows what it may
|
||||||
|
destroy?
|
||||||
|
- Should the host refuse a target change outright, requiring the old resource to be removed by a
|
||||||
|
declaration that still names it before a new one may take the name?
|
||||||
|
- Does the same hole exist for a resource whose **id** changes while the target stays, and for a
|
||||||
|
module unassigned between the two declarations?
|
||||||
|
- What reports this? Nothing on the machine currently answers "what is running here that the mesh
|
||||||
|
did not ask for", which is the question that would have found this in seconds.
|
||||||
Reference in New Issue
Block a user