Issue 113 and research 015: the object store's images are gone upstream, not access-restricted #103
@@ -0,0 +1,119 @@
|
||||
---
|
||||
status: active
|
||||
initiated: 2026-09-24
|
||||
touches:
|
||||
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
|
||||
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
|
||||
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
|
||||
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
|
||||
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
|
||||
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
|
||||
- 03-DESIGN/00-as-is/03-provisioning.md
|
||||
- 03-DESIGN/01-to-be/07-the-foundation.md
|
||||
- 04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md
|
||||
---
|
||||
|
||||
# 015 — The object store after MinIO: which S3 implementation, and how the data moves
|
||||
|
||||
**The question.** The mesh's object store is MinIO. Its community edition is archived upstream,
|
||||
its server and client images have been deleted from every public registry, and the pinned release
|
||||
is four and a half years old and will never be patched
|
||||
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)).
|
||||
Which S3-compatible implementation replaces it, and what is the migration track for the data and
|
||||
the provisioning model that sit on top of it?
|
||||
|
||||
**Why now, and why not sooner.** Nothing is on fire: nodes that already hold the images keep
|
||||
running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present
|
||||
image by design. The forcing function is not an outage but a one-way door — **no node that does
|
||||
not already hold the images can ever provision the module again**, so the mesh's ability to stand
|
||||
a node up from its declarations is already broken for this module, and silently.
|
||||
|
||||
**The direction is not a departure from the design; it is the design.** The foundation document
|
||||
already states the commitment:
|
||||
|
||||
> The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object
|
||||
> store, the OCI protocol for the registry. That is what keeps the naming safe rather than a
|
||||
> commitment that cannot be revisited.
|
||||
|
||||
The object store is also **not** a foundation service — ADR 0028 removed it, and it is an
|
||||
ordinary module required through the module graph by whatever wants one. (The "exception that is
|
||||
not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's
|
||||
own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket
|
||||
is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which
|
||||
is the cheapest kind of decision to make and the strongest kind to cite.
|
||||
|
||||
## What the replacement has to carry, measured
|
||||
|
||||
Taken from the module's manifest, its composition, its tool surface, and a search for its
|
||||
consumers across the catalogue — not from assumption.
|
||||
|
||||
| Requirement | Evidence in the module today |
|
||||
|---|---|
|
||||
| S3 API | The protocol every consumer speaks; already the design's stated dependency. |
|
||||
| **OIDC login against the mesh's identity provider** | Six configuration variables are wired and populated in practice — discovery URL, client id, client secret, scopes, display name, redirect — plus a dedicated entrypoint script that blocks startup until the provider answers. This is live, not aspirational. |
|
||||
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
|
||||
| A single-node form | Declared as a flavour, for development and small nodes. |
|
||||
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
|
||||
| A tool surface | Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning. |
|
||||
| A console | Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads. |
|
||||
|
||||
**Consumers, counted:** one application module, one capture module that takes a private bucket per
|
||||
node, one workflow module's tools, and the delivery/rescue internals of the shared library. The
|
||||
surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the
|
||||
OIDC story, not spread across the catalogue.
|
||||
|
||||
## Candidates
|
||||
|
||||
Scoped to **SeaweedFS** as the primary, with the others recorded so the rejection is not
|
||||
rediscovered.
|
||||
|
||||
- **SeaweedFS** — Apache-2.0, Go, twelve-plus years of development, erasure coding, and OIDC
|
||||
support in its S3/STS layer. Chosen to scope because it is the only candidate that plausibly
|
||||
preserves the OIDC requirement above, which is the one requirement that is live and least
|
||||
substitutable.
|
||||
- **Garage** — the lightest to operate and the simplest model, but **no native identity-provider
|
||||
integration**. Adopting it means losing OIDC console login or fronting it with a proxy. A real
|
||||
functional regression against something currently in use.
|
||||
- **RustFS** — markets itself as a binary-level drop-in retaining existing data, buckets and
|
||||
configuration, which would make the data migration close to trivial. Young, and that claim is
|
||||
exactly the kind that must be verified on a copy before it is believed.
|
||||
- **Ceph RGW** — the most capable and the most operationally expensive; disproportionate to a mesh
|
||||
where the object store is an ordinary module, not a platform.
|
||||
|
||||
**The first thing to verify, because the choice turns on it:** how much of SeaweedFS's OIDC story
|
||||
is in the freely licensed build, and whether its shape — IAM/STS token exchange — can actually
|
||||
stand in for a console that redirects a human to an identity provider. If it cannot, the honest
|
||||
finding may be that **no** candidate preserves the current feature set, and the decision becomes
|
||||
which regression to accept. That question is worth answering before any migration work starts.
|
||||
|
||||
## The migration track, in outline
|
||||
|
||||
Data movement is the easy half, and deliberately reversible.
|
||||
|
||||
1. **Stand the replacement up beside the incumbent**, on its own ports and its own provision type.
|
||||
No downtime, nothing removed.
|
||||
2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client
|
||||
— the client has been withdrawn upstream too, so building the migration on it would inherit
|
||||
the same dependency this effort exists to remove.
|
||||
3. **Verify per bucket** — object counts and checksums, not a transfer exit code.
|
||||
4. **Repoint consumers through the connection the module already publishes.** Consumers read an
|
||||
API URL from the module's declared connections rather than addressing the store directly, so
|
||||
the cutover surface is that value plus the provisioning and tool handlers.
|
||||
5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
|
||||
rollback until confidence is earned.
|
||||
6. **Retire**, and only then remove the module.
|
||||
|
||||
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
|
||||
handlers**, which are written against MinIO's admin API, and the OIDC wiring.
|
||||
|
||||
## Open questions
|
||||
|
||||
- How much of the OIDC requirement survives, and in which build? See above — this gates the
|
||||
choice.
|
||||
- Does the mesh's bucket provision translate to the candidate's identity model without weakening
|
||||
what ADR 0049 says about a consumer's identity fitting the tightest backend?
|
||||
- Should this effort also answer issue 113's general question — mirroring third-party images into
|
||||
the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with
|
||||
another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
|
||||
- Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking
|
||||
while the product is being chosen, rather than reproducing a shape by default.
|
||||
@@ -0,0 +1,106 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-24
|
||||
located-in: [hal modules/minio]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 113 — The object store's images were withdrawn upstream, and only a node that already holds them can still run it
|
||||
|
||||
## What was observed
|
||||
|
||||
On 2026-09-24, during a service-by-service cutover, the object-store module could not be built on
|
||||
a node that did not already hold its images. Both images the module needs answer an anonymous
|
||||
pull with `401 UNAUTHORIZED`:
|
||||
|
||||
```
|
||||
<registry>/minio/minio 401 <registry>/minio/operator 200
|
||||
<registry>/minio/mc 401 <registry>/minio/console 200
|
||||
```
|
||||
|
||||
The module pins a **tag**, not a digest, and the default is four and a half years old:
|
||||
|
||||
```
|
||||
image: <registry>/minio/minio:${MINIO_VERSION:-RELEASE.2022-01-07T01-53-23Z}
|
||||
```
|
||||
|
||||
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
|
||||
repositories. It is not an access policy that a credential could answer, and nothing about the
|
||||
mesh's own registry configuration, resolver or trust settings is involved.
|
||||
|
||||
- The vendor removed both repositories from the main public registry on **2026-09-11**. Its API
|
||||
answers `404` for the server repository while a sibling in the same namespace answers `200`.
|
||||
- The secondary registry that the wider ecosystem repointed to as a stopgap **no longer lists
|
||||
them either**. Sixty-eight repositories in that namespace are still public and pull normally;
|
||||
the server and the client are simply absent, and the namespace is now dominated by the vendor's
|
||||
commercially licensed line.
|
||||
- The open-source repository was archived in **February 2026**, and the community edition has been
|
||||
source-only since **October 2025**. No new images are published anywhere.
|
||||
|
||||
## What did not happen, and why it is recorded
|
||||
|
||||
The first reading of this was that the registry had *disabled anonymous pulls for the whole
|
||||
vendor namespace*. That was wrong in a way worth keeping, because the evidence looked conclusive:
|
||||
|
||||
- The anonymous token carries `"actions": []` for the affected repositories and `['pull']` for
|
||||
working ones — a real signal, but it is **also exactly what a repository that does not exist
|
||||
returns**. A deliberately invented repository name in the same namespace produced a
|
||||
byte-identical response. The signal cannot distinguish *revoked* from *absent*.
|
||||
- The token also carries `"$disabled"`, which was read as confirmation. It appears on **every**
|
||||
repository on that registry, including the ones pulling successfully. It describes image
|
||||
**signing**, not access.
|
||||
|
||||
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
|
||||
policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
|
||||
|
||||
## The mesh was not blocked, which is the other half
|
||||
|
||||
A node that already holds the images runs the module normally. The node carrying the cutover holds
|
||||
the pinned server image, the client, and the load-balancer image the module composes with, all
|
||||
pulled years ago. Its resolved version variable matches the cached tag exactly.
|
||||
|
||||
This is by design and not by luck. The deploy stage pulls **best-effort** and then asserts only
|
||||
that every image the composition declares **resolves locally**, precisely so that an image which
|
||||
exists on the node but can no longer be fetched does not fail a deploy. The code comment naming
|
||||
the precedent describes this case exactly — *"an old tag pulled years ago and since removed
|
||||
upstream"* — and records that failing on the pull instead had previously made a module
|
||||
undeployable while all of its images sat on the node.
|
||||
|
||||
So the deploy logs a warning and succeeds. Dropping the module from the cutover queue was not
|
||||
necessary.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
The instance is harmless; the standing condition is not.
|
||||
|
||||
1. **No new node can ever provision this module.** Every node that does not already hold the
|
||||
images is permanently unable to obtain them. The mesh's claim that a node can be rebuilt from
|
||||
its declarations is false for this module, and will be false the same way for any module whose
|
||||
upstream withdraws an image.
|
||||
2. **The pinned release is permanently unpatched.** It is four and a half years old, upstream is
|
||||
archived, and no security fix will ever reach it.
|
||||
3. **Nothing detects this class of failure.** The condition is invisible until a node without the
|
||||
image tries to deploy. The very guard that correctly stops this from breaking existing nodes —
|
||||
assert local resolvability, not the pull — also means a warning is the only trace, and a
|
||||
warning is not a rule. A mesh cannot state that its modules are installable while the only
|
||||
evidence is that they are already installed.
|
||||
|
||||
The third point is the general one, and it is not specific to this vendor: an image pinned by tag
|
||||
against a registry the mesh does not control is a dependency with no guarantee behind it, and the
|
||||
mesh currently learns it has lost one only by trying to use it.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
|
||||
registry at adoption, so a module's installability does not depend on an upstream's continued
|
||||
goodwill? That is the fix that generalises, and it costs storage and a policy about what to
|
||||
mirror.
|
||||
- Should a module's images be pinned **by digest** rather than by tag? It makes the artifact
|
||||
exact and auditable, but does nothing about withdrawal — a deleted digest is just as gone.
|
||||
- What **checks** that every module in the catalogue is still obtainable from a node that holds
|
||||
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
|
||||
2026-09-11 rather than thirteen days later, mid-cutover.
|
||||
- For this module specifically: replace the product. The design already says the dependency is on
|
||||
**S3 the protocol, not the product** — see
|
||||
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md).
|
||||
@@ -0,0 +1,110 @@
|
||||
# 113 — Diagnosis
|
||||
|
||||
## 2026-09-24 — the trail
|
||||
|
||||
The symptom arrived already carrying a diagnosis: *the registry has disabled anonymous pulls for
|
||||
the entire vendor namespace.* Everything below was an attempt to confirm that, and it did not
|
||||
survive.
|
||||
|
||||
### Step 1 — the token is not the test
|
||||
|
||||
The reported evidence was the anonymous pull token's contents: `"actions": []` and `"$disabled"`.
|
||||
A token is an intermediate artifact. The test is whether a manifest can actually be fetched with
|
||||
it, so the first step was to request one:
|
||||
|
||||
```
|
||||
GET /v2/minio/minio/manifests/<pinned tag> -> 401 UNAUTHORIZED
|
||||
```
|
||||
|
||||
That confirmed the failure but said nothing about its scope or cause.
|
||||
|
||||
### Step 2 — a control ruled out the stated cause
|
||||
|
||||
The same request flow, in the same minute, against other repositories:
|
||||
|
||||
| repository | token actions | manifest |
|
||||
|---|---|---|
|
||||
| the vendor's server | `[]` | **401** |
|
||||
| the vendor's client | `[]` | **401** |
|
||||
| the vendor's operator | `['pull']` | 200 |
|
||||
| the vendor's console | `['pull']` | 200 |
|
||||
| the vendor's sidecar proxy | `['pull']` | 200 |
|
||||
| an unrelated public project | `['pull']` | 200 |
|
||||
|
||||
**A namespace-wide policy is ruled out.** Two repositories fail; their siblings in the same
|
||||
namespace pull normally.
|
||||
|
||||
### Step 3 — both pieces of the original evidence were red herrings
|
||||
|
||||
- `"$disabled"` is present on **every** repository on that registry, including all of the
|
||||
successful ones above. It belongs to the signing context, not to authorisation. It carries no
|
||||
information about this failure at all.
|
||||
- `"actions": []` with a `401` is **indistinguishable from a repository that does not exist.** A
|
||||
deliberately invented repository name in the vendor's namespace returned a byte-identical
|
||||
response — empty actions, `401`. The signal cannot separate *access revoked* from *not there*,
|
||||
so it cannot support the conclusion it was used for.
|
||||
|
||||
That second point turned the question from *who revoked access* to *is it still there*.
|
||||
|
||||
### Step 4 — the registries' own APIs establish deletion
|
||||
|
||||
Asked directly, rather than through the pull path:
|
||||
|
||||
- **Primary registry:** its repository API answers **404** for the server repository, and **200**
|
||||
for a sibling in the same namespace. The repository is gone, not private.
|
||||
- **Secondary registry:** a listing of the vendor's namespace returns **68 public repositories**.
|
||||
The server and the client are **absent from the list**. Present are the operator, console,
|
||||
sidecar proxy, benchmarking and key-management images — and a large, newer set belonging to the
|
||||
vendor's commercially licensed line.
|
||||
|
||||
### Step 5 — upstream confirms, and dates it
|
||||
|
||||
The vendor deleted the community server and client from the primary registry on **2026-09-11**.
|
||||
This was the last step of a staged withdrawal: free image publishing stopped in **October 2025**,
|
||||
the community console UI was removed mid-2025, and the open-source repository was archived in
|
||||
**February 2026**. The secondary registry was where the ecosystem repointed as a stopgap; it has
|
||||
since lost the two repositories as well.
|
||||
|
||||
**Conclusion: the images were withdrawn, not restricted.** No credential can answer this, because
|
||||
there is nothing left to authenticate against. A different image source is the only remedy.
|
||||
|
||||
## Why the mesh kept working, checked rather than assumed
|
||||
|
||||
The claim that the module "genuinely isn't ready for cutover" was tested and is false for the node
|
||||
in question.
|
||||
|
||||
- The node **holds** the pinned server image, the client, and the load-balancer image the module
|
||||
composes with — all pulled years before the withdrawal.
|
||||
- The module's resolved version variable on that node matches the cached tag **exactly**, so the
|
||||
composition references an image that is present.
|
||||
- The composition declares **no pull policy**, so a present tag is used as-is.
|
||||
- The deploy stage pulls best-effort, then asserts only that every declared image resolves
|
||||
locally. A pull error becomes a logged warning when all images are present.
|
||||
- The start path restarts the unit and does not pull. The one `--pull always` in the tree sits
|
||||
inside a generated guide describing the **superseded** approach, not in the code that writes
|
||||
units.
|
||||
|
||||
So a deploy of this module on that node succeeds today.
|
||||
|
||||
## Claims in the original report that could not be substantiated
|
||||
|
||||
Recorded because they were specific and load-bearing, and acting on them would have wasted time.
|
||||
|
||||
| Claim | Finding |
|
||||
|---|---|
|
||||
| The module's manifest is a `module.json`, pinning the server image by digest at line 83 | There is **no `module.json` anywhere** in the monorepo. The manifest is YAML and pins a **tag**. No digest pin exists. |
|
||||
| The module's own runtime artifact is a placeholder with an all-zero digest | **Zero occurrences** of that image name or of an all-zero digest anywhere in the tree. The module declares no runtime or sidecar artifact. |
|
||||
| A dated readiness document records six modules with placeholder digests | **No such file exists.** |
|
||||
| The server image is cached locally at an older release than the module pins | The cached tag is **exactly** the pinned one, not an older release. |
|
||||
|
||||
None of these change the real finding, which stands: the images are gone upstream.
|
||||
|
||||
## What is located, and what is not
|
||||
|
||||
**Located:** the module in the monorepo's catalogue — it pins, by tag, an image that no longer
|
||||
exists anywhere public.
|
||||
|
||||
**Not located, and deliberately left open:** the general condition. The mesh has no mirror of the
|
||||
third-party images its modules depend on and no check that a module is obtainable by a node
|
||||
holding nothing. That is a design gap rather than a defect in this module, and it is stated as an
|
||||
open question on the report rather than answered here.
|
||||
Reference in New Issue
Block a user