Files
hq/04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md
T
jochen ab7d216fc6 Issue 113 resolved by the repin, and what issue 064 did not cover
The object-store module was repinned to a maintained fork of the withdrawn server image, its
runtime sidecar built rather than pulled, and its data moved off the predecessor's live
directory. The instance is closed; the three general points the report makes are not, and What
was done says so rather than letting a resolved status imply otherwise.

Folds in the one thing a duplicate report of this symptom had that this one did not: issue 064
asked whether the build environment can reach a declared vendor image and assumed that, once
declared, it stays fetchable. Withdrawal is the case that assumption does not cover. The
duplicate is not merged — it carried the reading this report's diagnosis retracts.
2026-09-25 00:55:13 +02:00

152 lines
9.5 KiB
Markdown

---
status: resolved
opened: 2026-09-24
located-in: [mesh-catalog modules/minio]
fixed-by: mesh-catalog — the object-store module repinned to a maintained fork of the withdrawn server image, its runtime sidecar built from source rather than pulled, and its data moved off the predecessor's live directory. The standing condition this report names is not closed by it — see What was done.
amended-design:
---
# 113 — The object store's images were withdrawn upstream, and only a node that already holds them can still run it
## What was observed
On 2026-09-24, during a service-by-service cutover, the object-store module could not be built on
a node that did not already hold its images. Both images the module needs answer an anonymous
pull with `401 UNAUTHORIZED`:
```
<registry>/minio/minio 401 <registry>/minio/operator 200
<registry>/minio/mc 401 <registry>/minio/console 200
```
The module pins the server image **by digest**, in its `module.json`:
"image": "<registry>/minio/minio@sha256:14cea493…"
*Corrected 2026-09-24. This first said the module pinned a **tag** whose default was four and a
half years old. That described the **predecessor** mesh's object-store module — a different file
in a different repository — not the module being cut over to. The conflation, and what it cost,
is retracted in full in [the diagnosis](01-diagnosis.md).*
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
repositories. It is not an access policy that a credential could answer, and nothing about the
mesh's own registry configuration, resolver or trust settings is involved.
- The vendor removed both repositories from the main public registry on **2026-09-11**. Its API
answers `404` for the server repository while a sibling in the same namespace answers `200`.
- The secondary registry that the wider ecosystem repointed to as a stopgap **no longer lists
them either**. Sixty-eight repositories in that namespace are still public and pull normally;
the server and the client are simply absent, and the namespace is now dominated by the vendor's
commercially licensed line.
- The open-source repository was archived in **February 2026**, and the community edition has been
source-only since **October 2025**. No new images are published anywhere.
## What did not happen, and why it is recorded
The first reading of this was that the registry had *disabled anonymous pulls for the whole
vendor namespace*. That was wrong in a way worth keeping, because the evidence looked conclusive:
- The anonymous token carries `"actions": []` for the affected repositories and `['pull']` for
working ones — a real signal, but it is **also exactly what a repository that does not exist
returns**. A deliberately invented repository name in the same namespace produced a
byte-identical response. The signal cannot distinguish *revoked* from *absent*.
- The token also carries `"$disabled"`, which was read as confirmation. It appears on **every**
repository on that registry, including the ones pulling successfully. It describes image
**signing**, not access.
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
## The predecessor mesh was not blocked, which is the other half
*Scope, corrected 2026-09-24: everything in this section describes the **predecessor's** delivery
machinery and its object-store module. It is what made the instance harmless, and it is why
dropping the module from the queue was unnecessary. It says nothing about how the mesh being built
resolves images, which is a different mechanism.*
A node that already holds the images runs the module normally. The node carrying the cutover holds
the pinned server image, the client, and the load-balancer image the module composes with, all
pulled years ago. Its resolved version variable matches the cached tag exactly.
This is by design and not by luck. The deploy stage pulls **best-effort** and then asserts only
that every image the composition declares **resolves locally**, precisely so that an image which
exists on the node but can no longer be fetched does not fail a deploy. The code comment naming
the precedent describes this case exactly — *"an old tag pulled years ago and since removed
upstream"* — and records that failing on the pull instead had previously made a module
undeployable while all of its images sat on the node.
So the deploy logs a warning and succeeds. Dropping the module from the cutover queue was not
necessary.
## Why it matters beyond this instance
The instance is harmless; the standing condition is not.
1. **No new node can ever provision this module.** Every node that does not already hold the
images is permanently unable to obtain them, and the same will be true of any module whose
upstream withdraws an image.
This is the failure mode of a deliberate design choice, which is why it is worth recording
rather than patching. The foundation design chooses **references over payload** — *"the bundle
names images by digest and the host fetches them"* — on the stated grounds that
*"reproducibility comes from pinning the identity of a thing rather than carrying its bytes"*
([to-be 07](../../03-DESIGN/01-to-be/07-the-foundation.md)). That reasoning is sound. It holds
only while a pinned identity stays **resolvable**, and nothing in the mesh's control guarantees
that for an image in somebody else's registry. The passage is about
the foundation bundle, and this module is not in it; but the pattern — pin the identity, fetch
the bytes on demand — is how every module gets its third-party images, so the exposure is
general even though the sentence is local.
2. **The pinned release is permanently unpatched.** It is four and a half years old, upstream is
archived, and no security fix will ever reach it.
3. **Nothing detects this class of failure.** The condition is invisible until a node without the
image tries to deploy. The very guard that correctly stops this from breaking existing nodes —
assert local resolvability, not the pull — also means a warning is the only trace, and a
warning is not a rule. A mesh cannot state that its modules are installable while the only
evidence is that they are already installed.
The third point is the general one, and it is not specific to this vendor: an image pinned against
a registry the mesh does not control — **by tag or by digest, it makes no difference** — is a
dependency with no guarantee behind it, and the mesh currently learns it has lost one only by
trying to use it.
[Issue 064](../064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md) is the
nearest precedent, and it does not cover this. That issue asked whether the mesh's build
environment can **reach** a declared vendor image — a network-policy question, answered by
requiring the image be declared as a build input — and it assumed that an image, once declared,
stays fetchable. Withdrawal is the case the assumption does not cover: no network policy and no
declaration makes a deleted repository resolvable, so a module can satisfy 064 in full and still
be unbuildable on a node that holds nothing.
## What was done
The module was repinned to a maintained fork of the server image, published to a registry that
still serves it; its runtime sidecar is now built from source rather than pulled; and its data was
moved off the predecessor's live directory. The object store runs on the control-node from that
pin, and a node holding nothing can obtain it again.
That answers the instance and none of the three points above. The mesh still cannot say which of
its other pinned third-party images are still obtainable, and it would still learn of a withdrawal
only when a node without the image tried to deploy. The replacement question — S3 the protocol
rather than this product — is carried by
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md); the detection
question is carried by nothing, and is the first of the open questions below.
## Open questions
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
registry at adoption, so a module's installability does not depend on an upstream's continued
goodwill? That is the fix that generalises. It costs storage and a policy about what to mirror,
and it is a deliberate move **away** from references-over-payload for third-party images
specifically — so it should be decided as such, not smuggled in as a fix.
- ~~Should a module's images be pinned **by digest** rather than by tag?~~ **Answered, and the
premise was wrong.** This module already pins by digest, and it made no difference: the
repository was deleted, so the digest resolves to nothing. A digest buys an exact, auditable
artifact; it buys no protection whatever against withdrawal. Struck rather than deleted, because
the question was asked from a mistaken reading of the manifest and that is worth seeing.
- What **checks** that every module in the catalogue is still obtainable from a node that holds
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
2026-09-11 rather than thirteen days later, mid-cutover.
- For this module specifically: replace the product. The design already says the dependency is on
**S3 the protocol, not the product** — see
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md).