Files
hq/04-ISSUES/099-a-modules-image-pin-ages-into-a-downgrade/00-report.md
T

78 lines
4.3 KiB
Markdown

---
status: resolved
opened: 2026-09-23
located-in: [mesh-controller cmd/mesh-controller/adoption.go (take previews nothing), mesh-host internal/apply (the comparison and the record)]
fixed-by: mesh-host 63 (both images' creation dates), mesh-controller 201 (DOWNGRADE said; refused unless `--downgrade`)
amended-design:
---
# 099 — A module's image pin ages into a downgrade, and taking it over is where that is discovered
## What was observed
Three modules in a row, on the same machine, 2026-09-22 and 2026-09-23.
A catalogue module pins its image by digest, which is right: a module must run the same software
everywhere, and a floating tag is not a version. The digest is chosen when the module is written.
The machine whose service it takes over runs whatever *its* floating tag resolved to, and that has
moved on since.
| | the module pinned | the machine ran |
|---|---|---|
| the forge | 1.22.6 | 1.27.3 |
| the package registry | 6.10.1 | 6.10.3 |
| the analytics service | built 2026-08-20 | built 2026-09-17 |
The first was found by taking it. The service came up on the older build, refused to start against a
database its newer self had migrated, and **was down about three minutes** before the cutover was
rolled back. The other two were found by hand afterwards, only because the first had made it a
thing to check.
Nothing in the mesh compares the two. The host is holding the very container it is about to
replace — it knows the image that container runs, it reports it as held, and it replaces it without
a word about what changes. `take` says a container was replaced. It does not say *with what*.
## Why it matters beyond this instance
The pin is not wrong when it is written and it is not wrong for a fresh install; it goes wrong by
**sitting still while the world moves**, and the moment it is discovered is the moment a running
service is replaced by it. Every module in the catalogue has this property, and the ones most
likely to have drifted are the ones written earliest — which, in a migration, is most of them.
A downgrade is not an ordinary failure. A service that has already migrated its own data forward
does not start on an older build; it may also start and quietly write the older format. Rolling
back is a restore, not an undo.
The runbook now says to compare versions before every cutover, and that is the shape of rule this
repository says it should not settle for: a rule stated where a person will read it, enforced by
nothing, checked by remembering. It was checked by remembering twice and missed once, which is the
expected rate.
## Open questions
- Should the mesh refuse to replace a held container with an **older** image and make the operator
say so explicitly — and what orders two digests, given neither a digest nor a registry tells you
which came first? An image's own creation date is in its config, available before pulling.
- Should `take` preview at all? It is the step that replaces a running service and it previews
nothing; the whole-node flip previews more than the per-service cutover it is made of.
- Should the catalogue know when a module's pin is behind its upstream, the way the mesh already
knows when a module is behind its **source**? `build --behind` answers the second question and
nothing answers the first.
- Is a digest pin the right thing for a module that takes over an existing service at all, or
should a cutover be able to say *keep what is running* and record what that was?
## Decided, 2026-10-01
[ADR 0163](../../02-DECISIONS/0163-taking-a-module-over-is-a-comparison.md), rules 1 and 2: the images are compared by age and a downgrade refuses. Building follows,
host first, then the controller's `take`.
## Built, 2026-10-02
mesh-host 63 reports the found image and both images' creation dates; mesh-controller 201 says
DOWNGRADE and refuses unless `--downgrade` is said. Stays located until a take is read on an adopted
machine.
## Resolved, 2026-10-02
Closed on the operator's decision of 2026-10-02 with the built and tested code live on every machine (mesh-controller 206, mesh-host 64), not on a take read on an adopted machine: every machine of this mesh is converged, so none holds a found thing to compare, and the record's live row — ADR 0163's last — will be read at the next real adoption rather than staged. Said here so nobody later believes that row was run.