Compare commits

..
Author SHA1 Message Date
jschoubben bff32e3370 Issue 136: a module may name a program the machine does not have, and everything reports success 2026-09-28 20:50:35 +02:00
mesh-admin d6b62387f2 Merge pull request 'Issue 135: a container's mesh names are not compared' (#165) from issue/135-a-containers-mesh-names into main 2026-09-28 15:19:37 +00:00
jschoubben 4789624857 Issue 135: a container's mesh names are not compared, so a moved address is never noticed
One container restarted 2286 times over five days while the mesh reported the machine as doing what it
was told. Its overlay address was five days out of date: the host compares a container by a digest of
its spec, and the mesh's names were not in it, so a container whose image and files never changed was
left alone holding a name that no longer resolved. Forty-eight others were current only because
something else had recreated them.

The same fault as issue 045, in the field that was left out. Resolved by putting the names in the
digest.
2026-09-28 17:19:35 +02:00
mesh-admin e00862e317 Merge pull request 'ADR 0112 is accepted, and issue 134 records what it is not yet' (#164) from decision/0112-accepted-and-issue-134 into main 2026-09-28 15:12:15 +00:00
2 changed files with 162 additions and 0 deletions
@@ -0,0 +1,70 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-host internal/apply]
fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour.
amended-design:
---
# 135 — A container's mesh names are not compared, so a moved address is never noticed
## What was observed
One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and
the mesh reported the machine as doing what it was told.
Its logs said its database connected and then a query timed out. The database was reachable: the same
query from the same network, with the same credential, answered in milliseconds. What differed was the
name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the
machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held
the old one:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
A container resolves other machines and public names through the entries the mesh gives it when it is
created, and nothing re-reads them afterwards. The host compares a container against what was declared
by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and
what it reads — and **the mesh's names were not in it**. So this container matched what was declared,
was left alone, and kept an address that had not existed for five days.
Forty-eight other containers had current names. Not because anything corrected them: each had been
recreated for some other reason — a new image, a changed file — and picked up the current roster on the
way. This one's image is an upstream release that had not moved, and nothing else about it changed, so
nothing ever recreated it.
## Why it matters beyond this instance
**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md)
named, in the one field that was left out.** That issue is why the digest carries what a container
reads: "a container whose configuration had since been rewritten compared equal and was left alone —
running values the machine no longer holds, while every check reported success." The same sentence
describes this, with *names* in place of *files*.
**The failure is invisible in exactly the way that matters.** The container runs, so the machine
reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The
only account of the fault is inside the container's own log, in the words of the application rather
than of the mesh — and what it says is that a query timed out, which points at the database.
**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs
itself by accident. The victim is the module whose image is stable — which is to say, the module that
was working fine.
## What was done
The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody
made. A container whose names moved is now recreated exactly as one whose image moved.
The first apply after this recreates every container that carries mesh names — one restart each,
already the price the mesh pays for any image update — because their recorded digests predate the
field.
## What is still true
The mesh gives a container its names at creation and has no way to change them in place. That is the
container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module
that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it.
@@ -0,0 +1,92 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-catalog modules/fail2ban]
fixed-by: mesh-catalog — the intrusion-prevention module bans through an action it ships itself, already in use on every machine, instead of naming a firewall front-end two of them do not have. The instance is closed; the class in "What is still true" is not.
amended-design:
---
# 136 — A module may name a program the machine does not have, and everything reports success
## What was observed
Two machines were given the intrusion-prevention module on 2026-09-28. Both refused to start it:
```
ERROR Failed during configuration: Have not found any log file for 'recidive' jail.
ERROR Async configuration of server failed
fail2ban.service: Main process exited, code=exited, status=255/EXCEPTION
```
The jail that bans whoever keeps coming back reads the service's *own* log, and the service checks
every jail's log file while it configures itself — before it has created that log. The module
declared the jail and shipped the rotation for that log, and never declared the log. On the two
machines where it had run for years the file was simply there, so nothing had ever noticed.
That failure was loud. Fixing it uncovered a second one in the same module that is not.
The module's defaults named `ufw` as the way to ban an address. Two of these four machines have no
`ufw` — they filter with nftables — and nothing checks that until an address is banned. Asked to ban
a documentation address on such a machine, the service accepted the instruction, counted it, ran the
command, and wrote this to a log nobody reads:
```
ERROR ... -- stderr: '/bin/sh: line 5: ufw: command not found'
ERROR ... -- returned 127
ERROR Failed to execute ban jail 'sshd' action 'ufw' ... Error banning 192.0.2.99
```
No rule existed afterwards. Throughout, the unit was `active`, the module was applied, and the
machine's report said so. **A machine had been added to the mesh's intrusion prevention, reported as
protected, and was banning nobody.**
## Why it matters beyond this instance
**The two faults are the same mistake with opposite symptoms.** Both are the module assuming
something about the machine — a file that happens to exist, a program that happens to be installed.
One stopped the service, which anybody notices. The other left it running and empty, which nobody
does. A mesh that only catches the loud one is a mesh whose coverage is unknown.
**"The unit is running" was taken for "the module is doing its job".** That is the only health a
service resource has. It is the right answer for most modules and it is silent for any module whose
work happens later, on an event — a ban, a renewal, a backup, a notification. The report cannot
distinguish "protecting this machine" from "installed and inert".
**And it is exactly the naming rule, one level down.**
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) says a
definition names no node, no mesh and no host path, because the same definition has to raise a
different mesh. `ufw` is not a node name, but it is the same class of assumption: a value the module
cannot know, true on some machines and false on others, written as though it were a constant. The
module already knew how to do better a few lines away — the mesh's own address range is named there
as something the machine fills in.
## What was done
The module declares the log its own jail reads, created once and never touched again, since what
grows in it is the service's and the rotation the module already ships is what keeps it small. And
it bans through the action it ships itself, which every machine here can run, which was already in
use by the other jail on all four, and which covers a container's published port as well as the
host's own.
All four machines now run it, with both jails, and a ban lands on each — verified by banning and
unbanning a documentation address on every one.
## What is still true
**Nothing would have caught either fault before it shipped.** The control plane reads a manifest, not
a machine; `ufw` and `/var/log/…` are strings in a file it has no way to evaluate. The host could in
principle be asked whether a declared program exists, but no resource says "this file names a command
that must be there", so there is nothing to check.
**Two machines' bans from before this are stale rules in the old front-end**, which the service no
longer knows about and will never lift. They reject two addresses for ever. Harmless, and a reminder
that changing how a module enforces something leaves what it already enforced behind.
## Open questions
- What does a service resource's health mean for a module whose work is event-driven? A unit being
active is the weakest claim available, and four of this mesh's modules are of that kind.
- Should a declaration be able to say that a resource depends on a program, so the machine can refuse
what it cannot carry out rather than reporting success?
- Where should the packet filter a module bans through come from — the module's own choice, as now,
or the seat that owns the machine's filtering?