Issue 099: a module's image pin ages into a downgrade, and taking it over is where that is discovered #85

Merged
jschoubben merged 1 commits from issues/099-a-pin-ages-into-a-downgrade into main 2026-09-23 00:48:02 +00:00
@@ -0,0 +1,62 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 099 — A module's image pin ages into a downgrade, and taking it over is where that is discovered
## What was observed
Three modules in a row, on the same machine, 2026-09-22 and 2026-09-23.
A catalogue module pins its image by digest, which is right: a module must run the same software
everywhere, and a floating tag is not a version. The digest is chosen when the module is written.
The machine whose service it takes over runs whatever *its* floating tag resolved to, and that has
moved on since.
| | the module pinned | the machine ran |
|---|---|---|
| the forge | 1.22.6 | 1.27.3 |
| the package registry | 6.10.1 | 6.10.3 |
| the analytics service | built 2026-08-20 | built 2026-09-17 |
The first was found by taking it. The service came up on the older build, refused to start against a
database its newer self had migrated, and **was down about three minutes** before the cutover was
rolled back. The other two were found by hand afterwards, only because the first had made it a
thing to check.
Nothing in the mesh compares the two. The host is holding the very container it is about to
replace — it knows the image that container runs, it reports it as held, and it replaces it without
a word about what changes. `take` says a container was replaced. It does not say *with what*.
## Why it matters beyond this instance
The pin is not wrong when it is written and it is not wrong for a fresh install; it goes wrong by
**sitting still while the world moves**, and the moment it is discovered is the moment a running
service is replaced by it. Every module in the catalogue has this property, and the ones most
likely to have drifted are the ones written earliest — which, in a migration, is most of them.
A downgrade is not an ordinary failure. A service that has already migrated its own data forward
does not start on an older build; it may also start and quietly write the older format. Rolling
back is a restore, not an undo.
The runbook now says to compare versions before every cutover, and that is the shape of rule this
repository says it should not settle for: a rule stated where a person will read it, enforced by
nothing, checked by remembering. It was checked by remembering twice and missed once, which is the
expected rate.
## Open questions
- Should the mesh refuse to replace a held container with an **older** image and make the operator
say so explicitly — and what orders two digests, given neither a digest nor a registry tells you
which came first? An image's own creation date is in its config, available before pulling.
- Should `take` preview at all? It is the step that replaces a running service and it previews
nothing; the whole-node flip previews more than the per-service cutover it is made of.
- Should the catalogue know when a module's pin is behind its upstream, the way the mesh already
knows when a module is behind its **source**? `build --behind` answers the second question and
nothing answers the first.
- Is a digest pin the right thing for a module that takes over an existing service at all, or
should a cutover be able to say *keep what is running* and record what that was?