Files
hq/04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/01-diagnosis.md
T
jochen 999636e2a8 Issue 113: retract the diagnosis table — the original report was right
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.

The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.

Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.

The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
2026-09-24 18:46:47 +02:00

137 lines
7.6 KiB
Markdown

# 113 — Diagnosis
## 2026-09-24 — the trail
The symptom arrived already carrying a diagnosis: *the registry has disabled anonymous pulls for
the entire vendor namespace.* Everything below was an attempt to confirm that, and it did not
survive.
### Step 1 — the token is not the test
The reported evidence was the anonymous pull token's contents: `"actions": []` and `"$disabled"`.
A token is an intermediate artifact. The test is whether a manifest can actually be fetched with
it, so the first step was to request one:
```
GET /v2/minio/minio/manifests/<pinned tag> -> 401 UNAUTHORIZED
```
That confirmed the failure but said nothing about its scope or cause.
### Step 2 — a control ruled out the stated cause
The same request flow, in the same minute, against other repositories:
| repository | token actions | manifest |
|---|---|---|
| the vendor's server | `[]` | **401** |
| the vendor's client | `[]` | **401** |
| the vendor's operator | `['pull']` | 200 |
| the vendor's console | `['pull']` | 200 |
| the vendor's sidecar proxy | `['pull']` | 200 |
| an unrelated public project | `['pull']` | 200 |
**A namespace-wide policy is ruled out.** Two repositories fail; their siblings in the same
namespace pull normally.
### Step 3 — both pieces of the original evidence were red herrings
- `"$disabled"` is present on **every** repository on that registry, including all of the
successful ones above. It belongs to the signing context, not to authorisation. It carries no
information about this failure at all.
- `"actions": []` with a `401` is **indistinguishable from a repository that does not exist.** A
deliberately invented repository name in the vendor's namespace returned a byte-identical
response — empty actions, `401`. The signal cannot separate *access revoked* from *not there*,
so it cannot support the conclusion it was used for.
That second point turned the question from *who revoked access* to *is it still there*.
### Step 4 — the registries' own APIs establish deletion
Asked directly, rather than through the pull path:
- **Primary registry:** its repository API answers **404** for the server repository, and **200**
for a sibling in the same namespace. The repository is gone, not private.
- **Secondary registry:** a listing of the vendor's namespace returns **68 public repositories**.
The server and the client are **absent from the list**. Present are the operator, console,
sidecar proxy, benchmarking and key-management images — and a large, newer set belonging to the
vendor's commercially licensed line.
### Step 5 — upstream confirms, and dates it
The vendor deleted the community server and client from the primary registry on **2026-09-11**.
This was the last step of a staged withdrawal: free image publishing stopped in **October 2025**,
the community console UI was removed mid-2025, and the open-source repository was archived in
**February 2026**. The secondary registry was where the ecosystem repointed as a stopgap; it has
since lost the two repositories as well.
**Conclusion: the images were withdrawn, not restricted.** No credential can answer this, because
there is nothing left to authenticate against. A different image source is the only remedy.
## Why the mesh kept working, checked rather than assumed
The claim that the module "genuinely isn't ready for cutover" was tested and is false for the node
in question.
- The node **holds** the pinned server image, the client, and the load-balancer image the module
composes with — all pulled years before the withdrawal.
- The module's resolved version variable on that node matches the cached tag **exactly**, so the
composition references an image that is present.
- The composition declares **no pull policy**, so a present tag is used as-is.
- The deploy stage pulls best-effort, then asserts only that every declared image resolves
locally. A pull error becomes a logged warning when all images are present.
- The start path restarts the unit and does not pull. The one `--pull always` in the tree sits
inside a generated guide describing the **superseded** approach, not in the code that writes
units.
So a deploy of this module on that node succeeds today.
## RETRACTED — the four claims this diagnosis called unsubstantiated
*Added 2026-09-24, the same day, after the error was pointed out.*
This diagnosis originally carried a table headed *"Claims in the original report that could not be
substantiated"*, asserting that a `module.json` did not exist, that no digest pin existed, that an
all-zeros runtime digest appeared nowhere, and that a readiness document was not on disk. **The
table was wrong and it is withdrawn in full.** The original report was accurate.
| Claim, as reported | Actual finding |
|---|---|
| The manifest is a `module.json` pinning the server image by digest at line 83 | **True, and exactly.** `modules/minio/module.json`, digest pin, line 83. |
| The module's own runtime artifact carries an all-zeros placeholder digest | **True.** A second container resource pins a runtime sidecar at an all-zeros digest, meaning nothing was ever published for it. |
| A dated readiness document records several modules with placeholder digests | **Unverified, not disproven.** It is not on the machine searched. The migration record is a separate private repository that is not checked out there, so its absence locally is not evidence. |
| The cached server image is an older release than the module pins | **Unverified.** What was checked was the *predecessor's* tag pin against the cache, which did match. Whether the cached image is the digest this module pins was never checked. |
### Why it went wrong, stated plainly
**One repository was searched, and absence in it was reported as absence.** The catalogue of the
mesh being built is a **separate repository**, not checked out on the machine where the search ran.
Every one of the four claims was about that repository. The searches were real and their output was
reported honestly; the inference drawn from them was not warranted.
Compounding it, the predecessor's object-store module and the one being cut over to were treated as
the same thing. They are different files, in different repositories, with different shapes: the
predecessor's is a compose file pinning a **tag** with a version variable, and it declares no
sidecar; the one being cut over to is a JSON manifest pinning a **digest**, and it declares two
container resources. Findings about the first were written up as findings about the second.
**The lesson worth keeping, because it is not specific to this issue:** *"zero occurrences anywhere
in the tree"* is only ever as strong as the tree that was searched, and a diagnosis must name which
tree that was. This one did not, which is what let a one-repository search read as a mesh-wide fact.
A confident rebuttal of a correct report is worse than no diagnosis, because it sends the next
person looking in the wrong place with the authority of a written record behind them.
What none of this changes: the images are gone upstream, and that finding stands on the registries'
own APIs.
## What is located, and what is not
**Located:** the object-store module in the catalogue of the mesh being built — it pins, by digest,
a server image that no longer exists anywhere public, and a runtime sidecar that was never
published.
**Not located, and deliberately left open:** the general condition. The mesh has no mirror of the
third-party images its modules depend on and no check that a module is obtainable by a node
holding nothing. That is a design gap rather than a defect in this module, and it is stated as an
open question on the report rather than answered here.