Merge pull request 'Research 015: reopen the comparison — the premise for narrowing to one candidate was false' (#105) from storage/015-reopen-the-candidate-comparison into main

This commit was merged in pull request #105.
This commit is contained in:
2026-09-24 16:47:18 +00:00
4 changed files with 284 additions and 47 deletions
@@ -50,7 +50,9 @@ consumers across the catalogue — not from assumption.
| Requirement | Evidence in the module today | | Requirement | Evidence in the module today |
|---|---| |---|---|
| S3 API | The protocol every consumer speaks; already the design's stated dependency. | | S3 API | The protocol every consumer speaks; already the design's stated dependency. |
| **OIDC login against the mesh's identity provider** | Six configuration variables are wired and populated in practice — discovery URL, client id, client secret, scopes, display name, redirect — plus a dedicated entrypoint script that blocks startup until the provider answers. This is live, not aspirational. | | ~~OIDC login against the mesh's identity provider~~ | **Struck 2026-09-24. Not a requirement, and it never worked.** Six variables are wired and an entrypoint blocks on the provider, which reads as a live feature. The module's own hook comment records the end state as *"policy claim missing"* — a failing login. See [01](01-candidate-comparison.md). |
| **Per-application access keys, each scoped to a bucket** | The real requirement. A "user" of the store is normally an application; the mesh already mints a credential per provisioned bucket. |
| **One live consumer using it as opaque primary storage** | A file-sync application, since early 2023: objects named by internal id, metadata in its own database. Highest-risk consumer — a live copy drifts, and its bucket name must be preserved. |
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. | | Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
| A single-node form | Declared as a flavour, for development and small nodes. | | A single-node form | Declared as a flavour, for development and small nodes. |
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). | | Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
@@ -64,34 +66,49 @@ OIDC story, not spread across the catalogue.
## Candidates ## Candidates
Scoped to **SeaweedFS** as the primary, with the others recorded so the rejection is not **Four candidates, not three.** The comparison was briefly narrowed to SeaweedFS on the strength of
rediscovered. console single sign-on; that axis turned out not to be a requirement, and the incumbent's own
maintained fork had been omitted altogether. Both errors, and why they happened, are recorded in
[01 — the candidates measured](01-candidate-comparison.md), which carries the evidence and the
requirement-by-requirement detail.
- **SeaweedFS** — Apache-2.0, Go, twelve-plus years of development, erasure coding, and OIDC In short, and only in short:
support in its S3/STS layer. Chosen to scope because it is the only candidate that plausibly
preserves the OIDC requirement above, which is the one requirement that is live and least
substitutable.
- **Garage** — the lightest to operate and the simplest model, but **no native identity-provider
integration**. Adopting it means losing OIDC console login or fronting it with a proxy. A real
functional regression against something currently in use.
- **RustFS** — markets itself as a binary-level drop-in retaining existing data, buckets and
configuration, which would make the data migration close to trivial. Young, and that claim is
exactly the kind that must be verified on a copy before it is believed.
- **Ceph RGW** — the most capable and the most operationally expensive; disproportionate to a mesh
where the object store is an ordinary module, not a platform.
**The first thing to verify, because the choice turns on it:** how much of SeaweedFS's OIDC story - **The maintained fork of the incumbent** — the community edition was archived and its images
is in the freely licensed build, and whether its shape — IAM/STS token exchange — can actually deleted, but a fork publishes, tracks CVEs, and preserves the on-disk format, S3 API and
stand in for a console that redirects a human to an identity provider. If it cannot, the honest environment surface. Costs **an image reference** where every other option costs a data
finding may be that **no** candidate preserves the current feature set, and the decision becomes migration, two rewrites and a maintenance window. Does not end the dependence on an abandoned
which regression to accept. That question is worth answering before any migration work starts. codebase; buys time to choose deliberately.
- **Garage** — its permission model *is* the requirement (per access key, per bucket), its admin
API is the closest match to how the mesh provisions, and the highest-risk consumer is
first-party documented against it. Remaining cost: no object versioning, no server-side
encryption or object locking, partial lifecycle — **unmeasured against the ten buckets, and the
one thing that could still disqualify it**.
- **SeaweedFS** — longest field record and erasure coding. Its console sign-on is a paid feature,
which is now beside the point. What weighs against it is narrower: its S3 surface is a gateway
translating onto its own file-system API, with no first-party support for the opaque consumer.
- **RustFS** — closest in shape to the incumbent, so the least porting. But it reached general
availability eight days before this was written, and carries an open defect in the credential
path. Two earlier claims about it are corrected in 01: it is **not** a drop-in that retains
existing data.
- **Ceph RGW** — remains rejected as disproportionate where the object store is an ordinary
module rather than a platform.
**This is now two decisions, not one:** whether to repoint to the fork or migrate, and — if
migrating — to which. Repointing does not foreclose migrating, which is the argument for taking it
first. On the corrected requirement the migration ranking is Garage, then SeaweedFS, and not yet
RustFS. Two measurements gate any graduation: **which S3 endpoints the consumers actually call**
(Garage cannot be ranked fairly until counted), and **whether the fork can read the incumbent's
on-disk format in place** — tested on a copy, because the migration between them is one-way. Both
are in [01](01-candidate-comparison.md#what-is-still-unmeasured).
## The migration track, in outline ## The migration track, in outline
Data movement is the easy half, and deliberately reversible. Data movement is the easy half, and deliberately reversible.
1. **Stand the replacement up beside the incumbent**, on its own ports and its own provision type. 1. **Stand the replacement up beside the incumbent**, on its own ports, its own provision type and
No downtime, nothing removed. **its own data directory**. Nothing removed. The data directory matters: reusing one the
incumbent already holds would put a fresh single-drive store on top of a live erasure set.
2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client 2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client
— the client has been withdrawn upstream too, so building the migration on it would inherit — the client has been withdrawn upstream too, so building the migration on it would inherit
the same dependency this effort exists to remove. the same dependency this effort exists to remove.
@@ -100,16 +117,22 @@ Data movement is the easy half, and deliberately reversible.
API URL from the module's declared connections rather than addressing the store directly, so API URL from the module's declared connections rather than addressing the store directly, so
the cutover surface is that value plus the provisioning and tool handlers. the cutover surface is that value plus the provisioning and tool handlers.
5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the 5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
rollback until confidence is earned. rollback until confidence is earned. For the opaque consumer this is **not optional and not
instant**: it stores objects by internal id with metadata in its own database, so a copy taken
while it runs will drift. It needs a maintenance window for the final sync, and the window is
proportional to 82,496 objects rather than to 230 GiB.
6. **Retire**, and only then remove the module. 6. **Retire**, and only then remove the module.
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
handlers**, which are written against MinIO's admin API, and the OIDC wiring. handlers**, both written against the incumbent's admin API. *The OIDC wiring was previously listed
here and is struck: it is not a requirement and it never worked.*
## Open questions ## Open questions
- How much of the OIDC requirement survives, and in which build? See above — this gates the - ~~How much of the OIDC requirement survives, and in which build?~~ **Answered, and it was the
choice. wrong question.** The console requirement does not exist, and the login it referred to never
worked. What replaced it: which S3 endpoints consumers actually call, and whether the fork reads
the incumbent's format in place.
- Does the mesh's bucket provision translate to the candidate's identity model without weakening - Does the mesh's bucket provision translate to the candidate's identity model without weakening
what ADR 0049 says about a consumer's identity fitting the tightest backend? what ADR 0049 says about a consumer's identity fitting the tightest backend?
- Should this effort also answer issue 113's general question — mirroring third-party images into - Should this effort also answer issue 113's general question — mirroring third-party images into
@@ -0,0 +1,176 @@
# 015 / 01 — The candidates measured
*Rewritten 2026-09-24. An earlier version of this document ranked the candidates on whether they
preserved single-sign-on to the object store's **console**. That was the wrong axis — it is not a
requirement — and a fourth candidate was missing entirely. Both errors are recorded at the end,
because how a comparison came to be ranked on the wrong thing is worth more than the ranking was.*
## The requirement, corrected
Taken from the operator and from the running system, not from the module's shape.
**A "user" of the object store is normally an application.** The requirement is therefore
**per-application access keys, each scoped to its own bucket** — not per-human single sign-on. The
mesh already works this way: it mints a credential for every provisioned bucket, and the consumer
reads an endpoint from the module's declared connection rather than addressing the store directly.
**The console is not a requirement.** It was the axis the previous version ranked on, and it should
not have been.
**The identity-provider login never worked.** The predecessor's module wires six OIDC variables and
blocks startup until the provider answers, which reads like a working feature. It is not: the
module's own hook comment records the end state as *"policy claim missing"* — a **failing** login,
written up as progress because it proved the provider had registered. The identity provider emits no
such claim, nothing in the module creates the mapper, and the configured scope alone would not carry
a custom one. Two days of logs show no genuine login attempts, only internet scanners failing on an
STS API version. **Nothing should be carried forward on the assumption this works**, and no
candidate should be credited or penalised for matching it.
**One consumer is live, opaque, and holds real user files.** A file-sync application has used the
store as its **primary storage** since early 2023: objects named by an internal id, with all
metadata in its own database. Three consequences — a copy taken while it runs will drift, its bucket
name must be preserved or its database references break, and it is the highest-risk consumer of the
lot.
## What is actually stored, measured
| | |
|---|---|
| Logical | **230 GiB, 82,496 objects, 10 buckets** |
| Raw on disk | **468 GiB** — eight drive directories at 59 GiB each |
| Implied scheme | 468 ÷ 230 = **2.03×**, confirming erasure coding at half parity |
| Headroom | ~1.3 TiB free on the filesystem holding it |
**All eight "drives" are directories on one filesystem on one machine.** The erasure coding is
therefore not buying independent-drive redundancy; the real failure domain is the array underneath,
which has its own. This single fact decides more of the comparison than any product feature: a
scheme's redundancy model is close to irrelevant here, and what remains is its storage overhead.
At 230 GiB with 1.3 TiB free, **storage overhead is not a deciding cost either.** Replication at
three copies would run ~690 GiB against the present 468 GiB — about **+222 GiB**, comfortably
absorbed. Erasure coding at a wider stripe would *save* roughly 146 GiB. Both are rounding errors
against the headroom, and neither should decide this.
## The candidates
Four, not three. The previous version omitted the first.
### The maintained fork of the incumbent
The community edition was archived upstream and its images deleted
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)),
but **a fork is maintained and publishing** — `pgsty/minio`, from the Pigsty project. It restores
the console stripped from the community
build, rebuilt image and package distribution, tracks CVEs, and states that it preserves the on-disk
format, the S3 API and the environment-variable surface. Verified by pulling it: it reports a
current release, permissive-to-copyleft licensing unchanged from upstream, and identifies itself as
a community fork. Adoption is real — the server image has been pulled three quarters of a million
times.
**Why it reorders the comparison.** Every other candidate costs a data migration, a provisioning
handler rewritten against a different admin API, a tool surface ported, and a maintenance window for
the opaque consumer. The fork costs **an image reference**. It also closes the issue's one-way door:
a node holding nothing can provision the module again, and patches resume.
**What it does not do** is end the dependence on a codebase its original authors abandoned. It is
maintenance mode, largely one project's effort, with no new features intended. It buys time to
choose deliberately rather than under pressure — which is worth a great deal, and is not the same as
a decision.
### Garage
**The best fit for how the mesh provisions.** Its permission model is *per access key, per bucket,
read/write/owner* — which is the requirement above stated verbatim rather than approximated. Its
admin API is a first-class REST surface with tokens scopeable to exactly the two operations a bucket
provision performs. The opaque consumer is **first-party documented** against it, for primary
storage, including client-side encryption support.
Its previously-recorded penalties mostly dissolve under the corrected requirement: it has no console
and no identity-provider integration, neither of which is wanted; and it replaces AWS-style ACLs and
bucket policies with its own per-key-per-bucket model, which is the thing being asked for.
**What genuinely remains.** It replicates rather than erasure-codes — immaterial at this volume and
on a single array, as above. It does **not implement the full span of S3 endpoints**: object
versioning is absent, object locking and server-side encryption endpoints are absent, and lifecycle
is partial. **Whether any of the ten buckets depends on those is unmeasured, and it is the one thing
that could still disqualify it.**
### SeaweedFS
Longest field record of the group, permissive licence, erasure coding, and identity-provider
integration on the S3 API through token exchange. **Its console sign-on is a paid feature** — the
admin UI itself is open, its identity integration is not. That finding is what falsified the
previous version's narrowing, and it is now largely beside the point, since the console is not a
requirement.
What weighs against it here is narrower and more specific: its S3 surface is a **gateway
translating onto its own file-system API**, with acknowledged divergence from AWS behaviour at the
edges, and there is no first-party documentation for the opaque consumer. For a store already
holding real user files in an opaque layout, first-party support is worth more than a feature list.
It also carries more moving parts than a single-machine deployment needs.
### RustFS
Closest in shape to the incumbent — a similar admin API and client compatibility, so the existing
handlers would port with least effort — under a permissive licence, with erasure coding and a
console that does integrate an identity provider.
Two things were recorded about it earlier that were **wrong, and are corrected here**: it is *not* a
binary-level drop-in that retains existing data (API compatibility and on-disk compatibility are
separate paths, and the on-disk one is preview-scoped with documented encryption limits), and it
therefore offers no shortcut around the migration. It also carries **an open defect in the exact
area the mesh depends on** — an access key created by an identity-provider user reported denied on
all S3 operations.
Decisively for now: **it reached general availability eight days before this was written.** For a
component holding 230 GiB of real user files, field record is a feature, and it does not have one.
### Ceph RGW
Remains rejected, for the reason already recorded: disproportionate where the object store is an
ordinary module rather than a platform.
## Where this leaves it
**The decision is no longer "which product replaces the incumbent".** It is two decisions, and they
can be taken in either order but should not be confused:
1. **Repoint to the maintained fork, or migrate now?** Repointing is an image reference and it
closes the issue. Migrating now costs a data copy, two rewrites and a maintenance window, and
buys independence from an abandoned codebase sooner.
2. **If migrating, which?** On the corrected requirement the ranking is **Garage first** — its
permission model *is* the requirement, its provisioning API is the closest match, and the
highest-risk consumer is first-party supported. SeaweedFS second, on field record, with a
translation-layer caveat that matters more here than its feature list. RustFS not yet, on age.
Taking (1) does not foreclose (2), and that asymmetry is the argument for taking (1) first.
## What is still unmeasured
1. **Whether any of the ten buckets needs object versioning, server-side encryption or lifecycle.**
This gates Garage specifically and nothing else here answers it.
2. **Whether the fork's release can actually read the incumbent's on-disk format in place.** The
format is claimed compatible across a multi-year gap; the migration between them is one-way, so
this is tested on a copy or not at all.
3. **Whether the opaque consumer's maintenance window is acceptable**, and how long it actually is
at 82,496 objects.
4. **Whether the eight-drive erasure-coded shape is warranted at all.** The evidence above says it
is not buying what it appears to: eight directories, one array, one machine. It looks inherited.
**Nothing graduates to a decision before 1 and 2.**
## Two errors in the previous version of this document
Recorded because the shape of both survives anonymisation and neither is unique to this effort.
**It ranked on a requirement that did not exist.** Console single sign-on was treated as the axis
because the module's configuration showed it wired up, and a wired-up configuration was read as a
used feature. It was neither used nor working. *A configured feature is not an observed one*, and
the evidence needed was the operator's answer and the logs — both cheap, neither consulted before
the ranking was written.
**It omitted the incumbent's own fork.** The whole effort began because an upstream withdrew its
images; whether anyone had continued that upstream was the first question to ask and it was not
asked. The candidate list was assembled from a search for *alternatives*, which by construction
returns things that are not the incumbent. *When a dependency dies, "who took it over" precedes
"what replaces it".*
@@ -1,7 +1,7 @@
--- ---
status: located status: located
opened: 2026-09-24 opened: 2026-09-24
located-in: [hal modules/minio] located-in: [mesh-catalog modules/minio]
fixed-by: fixed-by:
amended-design: amended-design:
--- ---
@@ -19,11 +19,14 @@ pull with `401 UNAUTHORIZED`:
<registry>/minio/mc 401 <registry>/minio/console 200 <registry>/minio/mc 401 <registry>/minio/console 200
``` ```
The module pins a **tag**, not a digest, and the default is four and a half years old: The module pins the server image **by digest**, in its `module.json`:
``` "image": "<registry>/minio/minio@sha256:14cea493…"
image: <registry>/minio/minio:${MINIO_VERSION:-RELEASE.2022-01-07T01-53-23Z}
``` *Corrected 2026-09-24. This first said the module pinned a **tag** whose default was four and a
half years old. That described the **predecessor** mesh's object-store module — a different file
in a different repository — not the module being cut over to. The conflation, and what it cost,
is retracted in full in [the diagnosis](01-diagnosis.md).*
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
repositories. It is not an access policy that a credential could answer, and nothing about the repositories. It is not an access policy that a credential could answer, and nothing about the
@@ -54,7 +57,12 @@ vendor namespace*. That was wrong in a way worth keeping, because the evidence l
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
policy, and the registries' own APIs — `404` against `200` — are what establish deletion. policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
## The mesh was not blocked, which is the other half ## The predecessor mesh was not blocked, which is the other half
*Scope, corrected 2026-09-24: everything in this section describes the **predecessor's** delivery
machinery and its object-store module. It is what made the instance harmless, and it is why
dropping the module from the queue was unnecessary. It says nothing about how the mesh being built
resolves images, which is a different mechanism.*
A node that already holds the images runs the module normally. The node carrying the cutover holds A node that already holds the images runs the module normally. The node carrying the cutover holds
the pinned server image, the client, and the load-balancer image the module composes with, all the pinned server image, the client, and the load-balancer image the module composes with, all
@@ -96,9 +104,10 @@ The instance is harmless; the standing condition is not.
warning is not a rule. A mesh cannot state that its modules are installable while the only warning is not a rule. A mesh cannot state that its modules are installable while the only
evidence is that they are already installed. evidence is that they are already installed.
The third point is the general one, and it is not specific to this vendor: an image pinned by tag The third point is the general one, and it is not specific to this vendor: an image pinned against
against a registry the mesh does not control is a dependency with no guarantee behind it, and the a registry the mesh does not control — **by tag or by digest, it makes no difference** — is a
mesh currently learns it has lost one only by trying to use it. dependency with no guarantee behind it, and the mesh currently learns it has lost one only by
trying to use it.
## Open questions ## Open questions
@@ -107,8 +116,11 @@ mesh currently learns it has lost one only by trying to use it.
goodwill? That is the fix that generalises. It costs storage and a policy about what to mirror, goodwill? That is the fix that generalises. It costs storage and a policy about what to mirror,
and it is a deliberate move **away** from references-over-payload for third-party images and it is a deliberate move **away** from references-over-payload for third-party images
specifically — so it should be decided as such, not smuggled in as a fix. specifically — so it should be decided as such, not smuggled in as a fix.
- Should a module's images be pinned **by digest** rather than by tag? It makes the artifact - ~~Should a module's images be pinned **by digest** rather than by tag?~~ **Answered, and the
exact and auditable, but does nothing about withdrawal — a deleted digest is just as gone. premise was wrong.** This module already pins by digest, and it made no difference: the
repository was deleted, so the digest resolves to nothing. A digest buys an exact, auditable
artifact; it buys no protection whatever against withdrawal. Struck rather than deleted, because
the question was asked from a mistaken reading of the manifest and that is worth seeing.
- What **checks** that every module in the catalogue is still obtainable from a node that holds - What **checks** that every module in the catalogue is still obtainable from a node that holds
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
2026-09-11 rather than thirteen days later, mid-cutover. 2026-09-11 rather than thirteen days later, mid-cutover.
@@ -86,23 +86,49 @@ in question.
So a deploy of this module on that node succeeds today. So a deploy of this module on that node succeeds today.
## Claims in the original report that could not be substantiated ## RETRACTED — the four claims this diagnosis called unsubstantiated
Recorded because they were specific and load-bearing, and acting on them would have wasted time. *Added 2026-09-24, the same day, after the error was pointed out.*
| Claim | Finding | This diagnosis originally carried a table headed *"Claims in the original report that could not be
substantiated"*, asserting that a `module.json` did not exist, that no digest pin existed, that an
all-zeros runtime digest appeared nowhere, and that a readiness document was not on disk. **The
table was wrong and it is withdrawn in full.** The original report was accurate.
| Claim, as reported | Actual finding |
|---|---| |---|---|
| The module's manifest is a `module.json`, pinning the server image by digest at line 83 | There is **no `module.json` anywhere** in the monorepo. The manifest is YAML and pins a **tag**. No digest pin exists. | | The manifest is a `module.json` pinning the server image by digest at line 83 | **True, and exactly.** `modules/minio/module.json`, digest pin, line 83. |
| The module's own runtime artifact is a placeholder with an all-zero digest | **Zero occurrences** of that image name or of an all-zero digest anywhere in the tree. The module declares no runtime or sidecar artifact. | | The module's own runtime artifact carries an all-zeros placeholder digest | **True.** A second container resource pins a runtime sidecar at an all-zeros digest, meaning nothing was ever published for it. |
| A dated readiness document records six modules with placeholder digests | **No such file exists.** | | A dated readiness document records several modules with placeholder digests | **Unverified, not disproven.** It is not on the machine searched. The migration record is a separate private repository that is not checked out there, so its absence locally is not evidence. |
| The server image is cached locally at an older release than the module pins | The cached tag is **exactly** the pinned one, not an older release. | | The cached server image is an older release than the module pins | **Unverified.** What was checked was the *predecessor's* tag pin against the cache, which did match. Whether the cached image is the digest this module pins was never checked. |
None of these change the real finding, which stands: the images are gone upstream. ### Why it went wrong, stated plainly
**One repository was searched, and absence in it was reported as absence.** The catalogue of the
mesh being built is a **separate repository**, not checked out on the machine where the search ran.
Every one of the four claims was about that repository. The searches were real and their output was
reported honestly; the inference drawn from them was not warranted.
Compounding it, the predecessor's object-store module and the one being cut over to were treated as
the same thing. They are different files, in different repositories, with different shapes: the
predecessor's is a compose file pinning a **tag** with a version variable, and it declares no
sidecar; the one being cut over to is a JSON manifest pinning a **digest**, and it declares two
container resources. Findings about the first were written up as findings about the second.
**The lesson worth keeping, because it is not specific to this issue:** *"zero occurrences anywhere
in the tree"* is only ever as strong as the tree that was searched, and a diagnosis must name which
tree that was. This one did not, which is what let a one-repository search read as a mesh-wide fact.
A confident rebuttal of a correct report is worse than no diagnosis, because it sends the next
person looking in the wrong place with the authority of a written record behind them.
What none of this changes: the images are gone upstream, and that finding stands on the registries'
own APIs.
## What is located, and what is not ## What is located, and what is not
**Located:** the module in the monorepo's catalogue — it pins, by tag, an image that no longer **Located:** the object-store module in the catalogue of the mesh being built — it pins, by digest,
exists anywhere public. a server image that no longer exists anywhere public, and a runtime sidecar that was never
published.
**Not located, and deliberately left open:** the general condition. The mesh has no mirror of the **Not located, and deliberately left open:** the general condition. The mesh has no mirror of the
third-party images its modules depend on and no check that a module is obtainable by a node third-party images its modules depend on and no check that a module is obtainable by a node