The diagnosis carried a table headed "claims that could not be substantiated", denying a module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is withdrawn in full and replaced with what is actually true, plus the two claims that remain genuinely unverified rather than disproven. The cause: one repository was searched and absence in it was written up as absence. The catalogue of the mesh being built is a separate repository, not checked out where the search ran, and all four claims were about that repository. Compounding it, the predecessor's object-store module and the one being cut over to were treated as one thing — they are different files in different repositories, one pinning a tag with no sidecar, the other a digest with two container resources. Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage described the predecessor; the open question about pinning by digest is struck, because this module already does and it made no difference — a deleted digest resolves to nothing either way. The section on why nothing broke is now scoped explicitly to the predecessor's machinery. The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct report is worse than no diagnosis — it sends the next person to the wrong place with a written record behind them.
130 lines
7.7 KiB
Markdown
130 lines
7.7 KiB
Markdown
---
|
|
status: located
|
|
opened: 2026-09-24
|
|
located-in: [mesh-catalog modules/minio]
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 113 — The object store's images were withdrawn upstream, and only a node that already holds them can still run it
|
|
|
|
## What was observed
|
|
|
|
On 2026-09-24, during a service-by-service cutover, the object-store module could not be built on
|
|
a node that did not already hold its images. Both images the module needs answer an anonymous
|
|
pull with `401 UNAUTHORIZED`:
|
|
|
|
```
|
|
<registry>/minio/minio 401 <registry>/minio/operator 200
|
|
<registry>/minio/mc 401 <registry>/minio/console 200
|
|
```
|
|
|
|
The module pins the server image **by digest**, in its `module.json`:
|
|
|
|
"image": "<registry>/minio/minio@sha256:14cea493…"
|
|
|
|
*Corrected 2026-09-24. This first said the module pinned a **tag** whose default was four and a
|
|
half years old. That described the **predecessor** mesh's object-store module — a different file
|
|
in a different repository — not the module being cut over to. The conflation, and what it cost,
|
|
is retracted in full in [the diagnosis](01-diagnosis.md).*
|
|
|
|
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
|
|
repositories. It is not an access policy that a credential could answer, and nothing about the
|
|
mesh's own registry configuration, resolver or trust settings is involved.
|
|
|
|
- The vendor removed both repositories from the main public registry on **2026-09-11**. Its API
|
|
answers `404` for the server repository while a sibling in the same namespace answers `200`.
|
|
- The secondary registry that the wider ecosystem repointed to as a stopgap **no longer lists
|
|
them either**. Sixty-eight repositories in that namespace are still public and pull normally;
|
|
the server and the client are simply absent, and the namespace is now dominated by the vendor's
|
|
commercially licensed line.
|
|
- The open-source repository was archived in **February 2026**, and the community edition has been
|
|
source-only since **October 2025**. No new images are published anywhere.
|
|
|
|
## What did not happen, and why it is recorded
|
|
|
|
The first reading of this was that the registry had *disabled anonymous pulls for the whole
|
|
vendor namespace*. That was wrong in a way worth keeping, because the evidence looked conclusive:
|
|
|
|
- The anonymous token carries `"actions": []` for the affected repositories and `['pull']` for
|
|
working ones — a real signal, but it is **also exactly what a repository that does not exist
|
|
returns**. A deliberately invented repository name in the same namespace produced a
|
|
byte-identical response. The signal cannot distinguish *revoked* from *absent*.
|
|
- The token also carries `"$disabled"`, which was read as confirmation. It appears on **every**
|
|
repository on that registry, including the ones pulling successfully. It describes image
|
|
**signing**, not access.
|
|
|
|
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
|
|
policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
|
|
|
|
## The predecessor mesh was not blocked, which is the other half
|
|
|
|
*Scope, corrected 2026-09-24: everything in this section describes the **predecessor's** delivery
|
|
machinery and its object-store module. It is what made the instance harmless, and it is why
|
|
dropping the module from the queue was unnecessary. It says nothing about how the mesh being built
|
|
resolves images, which is a different mechanism.*
|
|
|
|
A node that already holds the images runs the module normally. The node carrying the cutover holds
|
|
the pinned server image, the client, and the load-balancer image the module composes with, all
|
|
pulled years ago. Its resolved version variable matches the cached tag exactly.
|
|
|
|
This is by design and not by luck. The deploy stage pulls **best-effort** and then asserts only
|
|
that every image the composition declares **resolves locally**, precisely so that an image which
|
|
exists on the node but can no longer be fetched does not fail a deploy. The code comment naming
|
|
the precedent describes this case exactly — *"an old tag pulled years ago and since removed
|
|
upstream"* — and records that failing on the pull instead had previously made a module
|
|
undeployable while all of its images sat on the node.
|
|
|
|
So the deploy logs a warning and succeeds. Dropping the module from the cutover queue was not
|
|
necessary.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
The instance is harmless; the standing condition is not.
|
|
|
|
1. **No new node can ever provision this module.** Every node that does not already hold the
|
|
images is permanently unable to obtain them, and the same will be true of any module whose
|
|
upstream withdraws an image.
|
|
|
|
This is the failure mode of a deliberate design choice, which is why it is worth recording
|
|
rather than patching. The foundation design chooses **references over payload** — *"the bundle
|
|
names images by digest and the host fetches them"* — on the stated grounds that
|
|
*"reproducibility comes from pinning the identity of a thing rather than carrying its bytes"*
|
|
([to-be 07](../../03-DESIGN/01-to-be/07-the-foundation.md)). That reasoning is sound. It holds
|
|
only while a pinned identity stays **resolvable**, and nothing in the mesh's control guarantees
|
|
that for an image in somebody else's registry. The passage is about
|
|
the foundation bundle, and this module is not in it; but the pattern — pin the identity, fetch
|
|
the bytes on demand — is how every module gets its third-party images, so the exposure is
|
|
general even though the sentence is local.
|
|
2. **The pinned release is permanently unpatched.** It is four and a half years old, upstream is
|
|
archived, and no security fix will ever reach it.
|
|
3. **Nothing detects this class of failure.** The condition is invisible until a node without the
|
|
image tries to deploy. The very guard that correctly stops this from breaking existing nodes —
|
|
assert local resolvability, not the pull — also means a warning is the only trace, and a
|
|
warning is not a rule. A mesh cannot state that its modules are installable while the only
|
|
evidence is that they are already installed.
|
|
|
|
The third point is the general one, and it is not specific to this vendor: an image pinned against
|
|
a registry the mesh does not control — **by tag or by digest, it makes no difference** — is a
|
|
dependency with no guarantee behind it, and the mesh currently learns it has lost one only by
|
|
trying to use it.
|
|
|
|
## Open questions
|
|
|
|
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
|
|
registry at adoption, so a module's installability does not depend on an upstream's continued
|
|
goodwill? That is the fix that generalises. It costs storage and a policy about what to mirror,
|
|
and it is a deliberate move **away** from references-over-payload for third-party images
|
|
specifically — so it should be decided as such, not smuggled in as a fix.
|
|
- ~~Should a module's images be pinned **by digest** rather than by tag?~~ **Answered, and the
|
|
premise was wrong.** This module already pins by digest, and it made no difference: the
|
|
repository was deleted, so the digest resolves to nothing. A digest buys an exact, auditable
|
|
artifact; it buys no protection whatever against withdrawal. Struck rather than deleted, because
|
|
the question was asked from a mistaken reading of the manifest and that is worth seeing.
|
|
- What **checks** that every module in the catalogue is still obtainable from a node that holds
|
|
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
|
|
2026-09-11 rather than thirteen days later, mid-cutover.
|
|
- For this module specifically: replace the product. The design already says the dependency is on
|
|
**S3 the protocol, not the product** — see
|
|
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md).
|