The symptom arrived diagnosed as "the registry disabled anonymous pulls for the whole vendor namespace". It did not hold: sibling repositories in that namespace pull normally, the "$disabled" token field appears on every repository including working ones and describes signing rather than access, and "actions": [] with a 401 is byte-identical to what an invented repository name returns. Both registries' own APIs establish deletion instead. Recorded because the correction is the expensive part to rediscover, and because the instance was harmless while the standing condition is not: no node that does not already hold the images can ever provision the module again, and nothing detects that until one tries. Research 015 scopes the replacement. It is not a redesign — the foundation design already commits to S3 the protocol rather than the product, and the object store is an ordinary module, so this instantiates an existing principle. The live OIDC wiring is the requirement that gates the choice, and it is checked first.
107 lines
5.8 KiB
Markdown
107 lines
5.8 KiB
Markdown
---
|
|
status: located
|
|
opened: 2026-09-24
|
|
located-in: [hal modules/minio]
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 113 — The object store's images were withdrawn upstream, and only a node that already holds them can still run it
|
|
|
|
## What was observed
|
|
|
|
On 2026-09-24, during a service-by-service cutover, the object-store module could not be built on
|
|
a node that did not already hold its images. Both images the module needs answer an anonymous
|
|
pull with `401 UNAUTHORIZED`:
|
|
|
|
```
|
|
<registry>/minio/minio 401 <registry>/minio/operator 200
|
|
<registry>/minio/mc 401 <registry>/minio/console 200
|
|
```
|
|
|
|
The module pins a **tag**, not a digest, and the default is four and a half years old:
|
|
|
|
```
|
|
image: <registry>/minio/minio:${MINIO_VERSION:-RELEASE.2022-01-07T01-53-23Z}
|
|
```
|
|
|
|
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
|
|
repositories. It is not an access policy that a credential could answer, and nothing about the
|
|
mesh's own registry configuration, resolver or trust settings is involved.
|
|
|
|
- The vendor removed both repositories from the main public registry on **2026-09-11**. Its API
|
|
answers `404` for the server repository while a sibling in the same namespace answers `200`.
|
|
- The secondary registry that the wider ecosystem repointed to as a stopgap **no longer lists
|
|
them either**. Sixty-eight repositories in that namespace are still public and pull normally;
|
|
the server and the client are simply absent, and the namespace is now dominated by the vendor's
|
|
commercially licensed line.
|
|
- The open-source repository was archived in **February 2026**, and the community edition has been
|
|
source-only since **October 2025**. No new images are published anywhere.
|
|
|
|
## What did not happen, and why it is recorded
|
|
|
|
The first reading of this was that the registry had *disabled anonymous pulls for the whole
|
|
vendor namespace*. That was wrong in a way worth keeping, because the evidence looked conclusive:
|
|
|
|
- The anonymous token carries `"actions": []` for the affected repositories and `['pull']` for
|
|
working ones — a real signal, but it is **also exactly what a repository that does not exist
|
|
returns**. A deliberately invented repository name in the same namespace produced a
|
|
byte-identical response. The signal cannot distinguish *revoked* from *absent*.
|
|
- The token also carries `"$disabled"`, which was read as confirmation. It appears on **every**
|
|
repository on that registry, including the ones pulling successfully. It describes image
|
|
**signing**, not access.
|
|
|
|
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
|
|
policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
|
|
|
|
## The mesh was not blocked, which is the other half
|
|
|
|
A node that already holds the images runs the module normally. The node carrying the cutover holds
|
|
the pinned server image, the client, and the load-balancer image the module composes with, all
|
|
pulled years ago. Its resolved version variable matches the cached tag exactly.
|
|
|
|
This is by design and not by luck. The deploy stage pulls **best-effort** and then asserts only
|
|
that every image the composition declares **resolves locally**, precisely so that an image which
|
|
exists on the node but can no longer be fetched does not fail a deploy. The code comment naming
|
|
the precedent describes this case exactly — *"an old tag pulled years ago and since removed
|
|
upstream"* — and records that failing on the pull instead had previously made a module
|
|
undeployable while all of its images sat on the node.
|
|
|
|
So the deploy logs a warning and succeeds. Dropping the module from the cutover queue was not
|
|
necessary.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
The instance is harmless; the standing condition is not.
|
|
|
|
1. **No new node can ever provision this module.** Every node that does not already hold the
|
|
images is permanently unable to obtain them. The mesh's claim that a node can be rebuilt from
|
|
its declarations is false for this module, and will be false the same way for any module whose
|
|
upstream withdraws an image.
|
|
2. **The pinned release is permanently unpatched.** It is four and a half years old, upstream is
|
|
archived, and no security fix will ever reach it.
|
|
3. **Nothing detects this class of failure.** The condition is invisible until a node without the
|
|
image tries to deploy. The very guard that correctly stops this from breaking existing nodes —
|
|
assert local resolvability, not the pull — also means a warning is the only trace, and a
|
|
warning is not a rule. A mesh cannot state that its modules are installable while the only
|
|
evidence is that they are already installed.
|
|
|
|
The third point is the general one, and it is not specific to this vendor: an image pinned by tag
|
|
against a registry the mesh does not control is a dependency with no guarantee behind it, and the
|
|
mesh currently learns it has lost one only by trying to use it.
|
|
|
|
## Open questions
|
|
|
|
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
|
|
registry at adoption, so a module's installability does not depend on an upstream's continued
|
|
goodwill? That is the fix that generalises, and it costs storage and a policy about what to
|
|
mirror.
|
|
- Should a module's images be pinned **by digest** rather than by tag? It makes the artifact
|
|
exact and auditable, but does nothing about withdrawal — a deleted digest is just as gone.
|
|
- What **checks** that every module in the catalogue is still obtainable from a node that holds
|
|
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
|
|
2026-09-11 rather than thirteen days later, mid-cutover.
|
|
- For this module specifically: replace the product. The design already says the dependency is on
|
|
**S3 the protocol, not the product** — see
|
|
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md).
|