Files
hq/04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/01-diagnosis.md
T
jochen d497b37e43 Issue 113 and research 015: the object store's images are gone upstream, not access-restricted
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.

Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.

Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
2026-09-24 15:27:03 +02:00

111 lines
5.6 KiB
Markdown

# 113 — Diagnosis
## 2026-09-24 — the trail
The symptom arrived already carrying a diagnosis: *the registry has disabled anonymous pulls for
the entire vendor namespace.* Everything below was an attempt to confirm that, and it did not
survive.
### Step 1 — the token is not the test
The reported evidence was the anonymous pull token's contents: `"actions": []` and `"$disabled"`.
A token is an intermediate artifact. The test is whether a manifest can actually be fetched with
it, so the first step was to request one:
```
GET /v2/minio/minio/manifests/<pinned tag> -> 401 UNAUTHORIZED
```
That confirmed the failure but said nothing about its scope or cause.
### Step 2 — a control ruled out the stated cause
The same request flow, in the same minute, against other repositories:
| repository | token actions | manifest |
|---|---|---|
| the vendor's server | `[]` | **401** |
| the vendor's client | `[]` | **401** |
| the vendor's operator | `['pull']` | 200 |
| the vendor's console | `['pull']` | 200 |
| the vendor's sidecar proxy | `['pull']` | 200 |
| an unrelated public project | `['pull']` | 200 |
**A namespace-wide policy is ruled out.** Two repositories fail; their siblings in the same
namespace pull normally.
### Step 3 — both pieces of the original evidence were red herrings
- `"$disabled"` is present on **every** repository on that registry, including all of the
successful ones above. It belongs to the signing context, not to authorisation. It carries no
information about this failure at all.
- `"actions": []` with a `401` is **indistinguishable from a repository that does not exist.** A
deliberately invented repository name in the vendor's namespace returned a byte-identical
response — empty actions, `401`. The signal cannot separate *access revoked* from *not there*,
so it cannot support the conclusion it was used for.
That second point turned the question from *who revoked access* to *is it still there*.
### Step 4 — the registries' own APIs establish deletion
Asked directly, rather than through the pull path:
- **Primary registry:** its repository API answers **404** for the server repository, and **200**
for a sibling in the same namespace. The repository is gone, not private.
- **Secondary registry:** a listing of the vendor's namespace returns **68 public repositories**.
The server and the client are **absent from the list**. Present are the operator, console,
sidecar proxy, benchmarking and key-management images — and a large, newer set belonging to the
vendor's commercially licensed line.
### Step 5 — upstream confirms, and dates it
The vendor deleted the community server and client from the primary registry on **2026-09-11**.
This was the last step of a staged withdrawal: free image publishing stopped in **October 2025**,
the community console UI was removed mid-2025, and the open-source repository was archived in
**February 2026**. The secondary registry was where the ecosystem repointed as a stopgap; it has
since lost the two repositories as well.
**Conclusion: the images were withdrawn, not restricted.** No credential can answer this, because
there is nothing left to authenticate against. A different image source is the only remedy.
## Why the mesh kept working, checked rather than assumed
The claim that the module "genuinely isn't ready for cutover" was tested and is false for the node
in question.
- The node **holds** the pinned server image, the client, and the load-balancer image the module
composes with — all pulled years before the withdrawal.
- The module's resolved version variable on that node matches the cached tag **exactly**, so the
composition references an image that is present.
- The composition declares **no pull policy**, so a present tag is used as-is.
- The deploy stage pulls best-effort, then asserts only that every declared image resolves
locally. A pull error becomes a logged warning when all images are present.
- The start path restarts the unit and does not pull. The one `--pull always` in the tree sits
inside a generated guide describing the **superseded** approach, not in the code that writes
units.
So a deploy of this module on that node succeeds today.
## Claims in the original report that could not be substantiated
Recorded because they were specific and load-bearing, and acting on them would have wasted time.
| Claim | Finding |
|---|---|
| The module's manifest is a `module.json`, pinning the server image by digest at line 83 | There is **no `module.json` anywhere** in the monorepo. The manifest is YAML and pins a **tag**. No digest pin exists. |
| The module's own runtime artifact is a placeholder with an all-zero digest | **Zero occurrences** of that image name or of an all-zero digest anywhere in the tree. The module declares no runtime or sidecar artifact. |
| A dated readiness document records six modules with placeholder digests | **No such file exists.** |
| The server image is cached locally at an older release than the module pins | The cached tag is **exactly** the pinned one, not an older release. |
None of these change the real finding, which stands: the images are gone upstream.
## What is located, and what is not
**Located:** the module in the monorepo's catalogue — it pins, by tag, an image that no longer
exists anywhere public.
**Not located, and deliberately left open:** the general condition. The mesh has no mirror of the
third-party images its modules depend on and no check that a module is obtainable by a node
holding nothing. That is a design gap rather than a defect in this module, and it is stated as an
open question on the report rather than answered here.