From 47896248574184e95013b3e21833d9fdffa3bf3c Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 28 Sep 2026 17:19:35 +0200 Subject: [PATCH] Issue 135: a container's mesh names are not compared, so a moved address is never noticed One container restarted 2286 times over five days while the mesh reported the machine as doing what it was told. Its overlay address was five days out of date: the host compares a container by a digest of its spec, and the mesh's names were not in it, so a container whose image and files never changed was left alone holding a name that no longer resolved. Forty-eight others were current only because something else had recreated them. The same fault as issue 045, in the field that was left out. Resolved by putting the names in the digest. --- .../00-report.md | 70 +++++++++++++++++++ 1 file changed, 70 insertions(+) create mode 100644 04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md diff --git a/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md b/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md new file mode 100644 index 0000000..8cb8964 --- /dev/null +++ b/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md @@ -0,0 +1,70 @@ +--- +status: resolved +opened: 2026-09-28 +located-in: [mesh-host internal/apply] +fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour. +amended-design: +--- + +# 135 — A container's mesh names are not compared, so a moved address is never noticed + +## What was observed + +One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and +the mesh reported the machine as doing what it was told. + +Its logs said its database connected and then a query timed out. The database was reachable: the same +query from the same network, with the same credential, answered in milliseconds. What differed was the +name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the +machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held +the old one: + +``` +umami created 2026-09-23 novox.internal:10.42.0.1 +mesh-catalog created today novox.internal:10.10.0.1 +``` + +A container resolves other machines and public names through the entries the mesh gives it when it is +created, and nothing re-reads them afterwards. The host compares a container against what was declared +by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and +what it reads — and **the mesh's names were not in it**. So this container matched what was declared, +was left alone, and kept an address that had not existed for five days. + +Forty-eight other containers had current names. Not because anything corrected them: each had been +recreated for some other reason — a new image, a changed file — and picked up the current roster on the +way. This one's image is an upstream release that had not moved, and nothing else about it changed, so +nothing ever recreated it. + +## Why it matters beyond this instance + +**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md) +named, in the one field that was left out.** That issue is why the digest carries what a container +reads: "a container whose configuration had since been rewritten compared equal and was left alone — +running values the machine no longer holds, while every check reported success." The same sentence +describes this, with *names* in place of *files*. + +**The failure is invisible in exactly the way that matters.** The container runs, so the machine +reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The +only account of the fault is inside the container's own log, in the words of the application rather +than of the mesh — and what it says is that a query timed out, which points at the database. + +**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs +itself by accident. The victim is the module whose image is stable — which is to say, the module that +was working fine. + +## What was done + +The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody +made. A container whose names moved is now recreated exactly as one whose image moved. + +The first apply after this recreates every container that carries mesh names — one restart each, +already the price the mesh pays for any image update — because their recorded digests predate the +field. + +## What is still true + +The mesh gives a container its names at creation and has no way to change them in place. That is the +container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module +that would rather re-read a roster from a file can already ask for one as a fact +([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on +it.