diff --git a/04-ISSUES/094-a-port-published-as-the-machine-side-cannot-be-moved/01-diagnosis.md b/04-ISSUES/094-a-port-published-as-the-machine-side-cannot-be-moved/01-diagnosis.md new file mode 100644 index 0000000..d545374 --- /dev/null +++ b/04-ISSUES/094-a-port-published-as-the-machine-side-cannot-be-moved/01-diagnosis.md @@ -0,0 +1,55 @@ +# Diagnosis — 2026-09-23 + +## One cause, two symptoms + +Both halves of the report come from the same blind spot, in two different places. + +**The check that refused the setting** collected, for each entry a container publishes, only its +**last** segment. A short form publishes one number, which is the container's port and the +machine's at once, so reading the last segment is right. A long form — a machine port mapped to a +different port inside the container — has two, and the last segment is the container's. So the +machine side, the number the module names in `listens`, in `serves` and in every port +substitution, was not in the set of ports the module was considered to publish, and giving it one +was refused as naming nothing. + +**The lookup that made naming the other half useless** reads the same map under two different +keys. Everything derived from what the module *declares* — the filter, the openings, the guard, +what a consumer is told, a port substituted into a file — looks the machine side up. The one +reader that rewrites the mapping handed to the container runtime looks up the container's port. +For a short form those are the same number and no one notices. For a long form they differ, so +exactly one reader ever found the entry: keyed the only way the check allowed, the container's +mapping moved and nothing else did, leaving a firewall, a set of openings and a consumer all +pointing at a port the software had left. + +That is also the third symptom in the report, seen from the other end. No opening was derived for +the moved port because the opening is derived from the declared port, which still read as the old +number; and a per-node `expose` could not rescue it, because `expose` keys on the same declared +port — it widens the opening on the port nothing is on any more, and naming the real machine port +is refused as a port the module does not listen on. + +## What was ruled out + +The openings derivation itself. Composed with no setting at all, a module publishing a long-form +mapping on an adopted node does get its opening, on the machine side, forwarded to the container's +port. A regression test now records that, deliberately passing before the fix as well as after, so +the next reader does not go looking there. + +## Answering the first open question + +**Either end names the mapping, and both answer.** A module publishing `"2222:22"` may reasonably +say it listens on the port its software uses or on the port the machine serves; the mesh accepts +whichever the setting names and returns the machine port under both, so every reader finds the +same number under the key it happens to hold. This keeps working what already worked — the +container's end was the only key the old check accepted — and makes it mean the same thing. + +Ambiguity is refused where it is real: one number naming two **different** mappings, and the two +ends of **one** mapping given two different machine ports. Naming both ends of one mapping with +the same number was already refused, by the rule that a machine port has one holder. + +## What this does not close + +The second and third open questions stand, and they are the push-blocking half: the setting is +still **stored** without a manifest in view, so a key that names nothing a module publishes is +accepted where it is set and refused at composition — where it stops the node being told anything +at all, rather than stopping that module. The fix removes one reason a key could be wrong; it does +not remove the shape of the failure. Carried to [issue 096](../096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md). diff --git a/04-ISSUES/096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md b/04-ISSUES/096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md new file mode 100644 index 0000000..7d3ac29 --- /dev/null +++ b/04-ISSUES/096-a-setting-that-cannot-work-is-stored-and-stops-the-node/00-report.md @@ -0,0 +1,60 @@ +--- +status: open +opened: 2026-09-23 +located-in: [] +fixed-by: +amended-design: +--- + +# 096 — A setting that cannot work is accepted where it is set, and stops the node where it is read + +## What was observed + +On the control-node during the first module's migration, 2026-09-23, and reproduced since against +the controller's own tests. + +A per-node port setting was accepted and stored. Composing that node's declaration then failed on +it, and because a node is told everything or nothing, **every push to that machine was refused** +until somebody found the setting and removed it. The message named a port number. It did not name +the setting, the layer it was stored in, or the node it had stopped; nothing said that a stored +statement was the reason the machine had gone quiet. + +The specific reason that setting could not work is [issue +094](../094-a-port-published-as-the-machine-side-cannot-be-moved/00-report.md), and it is fixed. +This issue is the shape that surrounded it, which is not: + +- **The place that stores a setting has no manifest in view.** It checks what it can without one — + that a value is a port, that ssh keeps its own, that two modules on the machine do not claim the + same machine port — and leaves anything that needs the module's own declaration to composition. + So a key naming a port the module does not publish is stored today, and a typo is stored today, + and both are found later, from the far end. +- **Composition fails the node, not the module.** One unusable statement about one module refuses + the whole declaration, so the other modules on that machine stop being told anything either — + including modules that were fine before the setting existed. + +## Why it matters beyond this instance + +The two together turn a typo into an outage of the control link for a machine, at a distance from +the thing that caused it. The delay is the damage: a refusal at the moment of setting is a +correction, and the same refusal a push later is a machine nobody can talk to, found by whoever +next notices it is not being updated. + +It generalises past ports. Any setting whose validity depends on the module's declaration has this +shape — a value that is well-formed on its own and impossible against the manifest. Ports are +merely where the mesh found it first, because migrating a service is when settings get written. + +It also touches a rule the mesh states elsewhere: what refuses, refuses early and by name. A +refusal that names a number rather than the statement that produced it cannot be acted on without +knowing the code. + +## Open questions + +- Should storing a setting compose it against the module's manifest first — and if so, against + which version, given the catalogue moves and a manifest that was right when the setting was + written may not be later? +- Or should the guard be at the far end: composition refuses that **module** and sends the rest of + the node, so an impossible statement costs one service and not the machine? +- Either way, what does a refusal have to name — the node, the module, the layer and the key — for + an operator to undo it without reading the source? +- Is there anything a node must never be pushed without, such that sending a partial declaration is + worse than sending none? diff --git a/04-ISSUES/097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md b/04-ISSUES/097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md new file mode 100644 index 0000000..8bf2f34 --- /dev/null +++ b/04-ISSUES/097-a-resource-that-changes-target-leaves-the-old-one-behind/00-report.md @@ -0,0 +1,70 @@ +--- +status: open +opened: 2026-09-23 +located-in: [] +fixed-by: +amended-design: +--- + +# 097 — A resource whose target changes leaves the old one behind, running + +## What was observed + +On the control-node, 2026-09-23, four hours after the forge's module was first assigned. + +A module resource named its container implicitly: with no name of its own, the host derived one +from the module and the resource's id. The manifest then gained an explicit name, because the +module had to take over a container the predecessor already ran under that name +([issue 090](../090-the-forge-module-does-not-take-over-the-forge-genesis-raised/00-report.md)). +The host applied the change by raising a container under the new name — and left the one under the +old name **running**. + +The host's own record shows why. It keeps one entry per resource id, and that entry holds the +resource's *current* target: + +``` +{"id": "gitea.server", "type": "container", "origin": "declared", + "holds": [2999, 2222], "target": "gitea", "applied_at": "..."} +``` + +There is no entry naming the old container. Rewriting the record on a target change is what +erases the only trace of what the host must now remove, so the old one cannot be found by the +thing that would have removed it. + +Every other container the mesh made on this machine is declared; this is the only one that is not. +It survived every reconcile since, and would survive a reboot: nothing declares it, so nothing +stops it, and nothing reports it. + +Harmless in this instance by luck — the stranded container published no ports and held an +anonymous volume rather than the service's data — and that luck is the point. Had the rename gone +the other way, two containers of the same service would have run against one data directory, or +the old one would have kept the port the new one needed and the new one would have failed to bind. + +## Why it matters beyond this instance + +The mesh's promise is that a machine runs what it was told and nothing else, and that a module +removed leaves nothing behind. Both depend on the host being able to name what it wrote. A record +that remembers only the current target breaks that for **any** resource whose target moves — and a +file is worse than a container, because a stale configuration file at the old path is still read +by whatever reads that path, silently, with no process to notice running twice. + +Renaming is not exotic. It happens exactly when a module is taught to take over something that +already exists, which is every module in a migration. + +It also crosses [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md): +on an adopted node the host must distinguish what it wrote from what it found, because what it +found is held and never removed. A resource that changes target turns something the host wrote +into something no record claims — which, on the next machine, is indistinguishable from something +found, and so would be kept for ever on purpose. + +## Open questions + +- Should the record keep every target a resource has had, and the host remove the ones it no + longer declares — and if so, for how long, given a record is also how the host knows what it may + destroy? +- Should the host refuse a target change outright, requiring the old resource to be removed by a + declaration that still names it before a new one may take the name? +- Does the same hole exist for a resource whose **id** changes while the target stays, and for a + module unassigned between the two declarations? +- What reports this? Nothing on the machine currently answers "what is running here that the mesh + did not ask for", which is the question that would have found this in seconds.