Files
hq/01-RESEARCH/015-the-object-store-after-minio/01-candidate-comparison.md
T
jochen 1d524a1fa8 Research 015: rewrite the comparison — wrong axis, and a missing candidate
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.

It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.

On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.

Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.

Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
2026-09-24 18:46:47 +02:00

177 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 015 / 01 — The candidates measured
*Rewritten 2026-09-24. An earlier version of this document ranked the candidates on whether they
preserved single-sign-on to the object store's **console**. That was the wrong axis — it is not a
requirement — and a fourth candidate was missing entirely. Both errors are recorded at the end,
because how a comparison came to be ranked on the wrong thing is worth more than the ranking was.*
## The requirement, corrected
Taken from the operator and from the running system, not from the module's shape.
**A "user" of the object store is normally an application.** The requirement is therefore
**per-application access keys, each scoped to its own bucket** — not per-human single sign-on. The
mesh already works this way: it mints a credential for every provisioned bucket, and the consumer
reads an endpoint from the module's declared connection rather than addressing the store directly.
**The console is not a requirement.** It was the axis the previous version ranked on, and it should
not have been.
**The identity-provider login never worked.** The predecessor's module wires six OIDC variables and
blocks startup until the provider answers, which reads like a working feature. It is not: the
module's own hook comment records the end state as *"policy claim missing"* — a **failing** login,
written up as progress because it proved the provider had registered. The identity provider emits no
such claim, nothing in the module creates the mapper, and the configured scope alone would not carry
a custom one. Two days of logs show no genuine login attempts, only internet scanners failing on an
STS API version. **Nothing should be carried forward on the assumption this works**, and no
candidate should be credited or penalised for matching it.
**One consumer is live, opaque, and holds real user files.** A file-sync application has used the
store as its **primary storage** since early 2023: objects named by an internal id, with all
metadata in its own database. Three consequences — a copy taken while it runs will drift, its bucket
name must be preserved or its database references break, and it is the highest-risk consumer of the
lot.
## What is actually stored, measured
| | |
|---|---|
| Logical | **230 GiB, 82,496 objects, 10 buckets** |
| Raw on disk | **468 GiB** — eight drive directories at 59 GiB each |
| Implied scheme | 468 ÷ 230 = **2.03×**, confirming erasure coding at half parity |
| Headroom | ~1.3 TiB free on the filesystem holding it |
**All eight "drives" are directories on one filesystem on one machine.** The erasure coding is
therefore not buying independent-drive redundancy; the real failure domain is the array underneath,
which has its own. This single fact decides more of the comparison than any product feature: a
scheme's redundancy model is close to irrelevant here, and what remains is its storage overhead.
At 230 GiB with 1.3 TiB free, **storage overhead is not a deciding cost either.** Replication at
three copies would run ~690 GiB against the present 468 GiB — about **+222 GiB**, comfortably
absorbed. Erasure coding at a wider stripe would *save* roughly 146 GiB. Both are rounding errors
against the headroom, and neither should decide this.
## The candidates
Four, not three. The previous version omitted the first.
### The maintained fork of the incumbent
The community edition was archived upstream and its images deleted
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)),
but **a fork is maintained and publishing** — `pgsty/minio`, from the Pigsty project. It restores
the console stripped from the community
build, rebuilt image and package distribution, tracks CVEs, and states that it preserves the on-disk
format, the S3 API and the environment-variable surface. Verified by pulling it: it reports a
current release, permissive-to-copyleft licensing unchanged from upstream, and identifies itself as
a community fork. Adoption is real — the server image has been pulled three quarters of a million
times.
**Why it reorders the comparison.** Every other candidate costs a data migration, a provisioning
handler rewritten against a different admin API, a tool surface ported, and a maintenance window for
the opaque consumer. The fork costs **an image reference**. It also closes the issue's one-way door:
a node holding nothing can provision the module again, and patches resume.
**What it does not do** is end the dependence on a codebase its original authors abandoned. It is
maintenance mode, largely one project's effort, with no new features intended. It buys time to
choose deliberately rather than under pressure — which is worth a great deal, and is not the same as
a decision.
### Garage
**The best fit for how the mesh provisions.** Its permission model is *per access key, per bucket,
read/write/owner* — which is the requirement above stated verbatim rather than approximated. Its
admin API is a first-class REST surface with tokens scopeable to exactly the two operations a bucket
provision performs. The opaque consumer is **first-party documented** against it, for primary
storage, including client-side encryption support.
Its previously-recorded penalties mostly dissolve under the corrected requirement: it has no console
and no identity-provider integration, neither of which is wanted; and it replaces AWS-style ACLs and
bucket policies with its own per-key-per-bucket model, which is the thing being asked for.
**What genuinely remains.** It replicates rather than erasure-codes — immaterial at this volume and
on a single array, as above. It does **not implement the full span of S3 endpoints**: object
versioning is absent, object locking and server-side encryption endpoints are absent, and lifecycle
is partial. **Whether any of the ten buckets depends on those is unmeasured, and it is the one thing
that could still disqualify it.**
### SeaweedFS
Longest field record of the group, permissive licence, erasure coding, and identity-provider
integration on the S3 API through token exchange. **Its console sign-on is a paid feature** — the
admin UI itself is open, its identity integration is not. That finding is what falsified the
previous version's narrowing, and it is now largely beside the point, since the console is not a
requirement.
What weighs against it here is narrower and more specific: its S3 surface is a **gateway
translating onto its own file-system API**, with acknowledged divergence from AWS behaviour at the
edges, and there is no first-party documentation for the opaque consumer. For a store already
holding real user files in an opaque layout, first-party support is worth more than a feature list.
It also carries more moving parts than a single-machine deployment needs.
### RustFS
Closest in shape to the incumbent — a similar admin API and client compatibility, so the existing
handlers would port with least effort — under a permissive licence, with erasure coding and a
console that does integrate an identity provider.
Two things were recorded about it earlier that were **wrong, and are corrected here**: it is *not* a
binary-level drop-in that retains existing data (API compatibility and on-disk compatibility are
separate paths, and the on-disk one is preview-scoped with documented encryption limits), and it
therefore offers no shortcut around the migration. It also carries **an open defect in the exact
area the mesh depends on** — an access key created by an identity-provider user reported denied on
all S3 operations.
Decisively for now: **it reached general availability eight days before this was written.** For a
component holding 230 GiB of real user files, field record is a feature, and it does not have one.
### Ceph RGW
Remains rejected, for the reason already recorded: disproportionate where the object store is an
ordinary module rather than a platform.
## Where this leaves it
**The decision is no longer "which product replaces the incumbent".** It is two decisions, and they
can be taken in either order but should not be confused:
1. **Repoint to the maintained fork, or migrate now?** Repointing is an image reference and it
closes the issue. Migrating now costs a data copy, two rewrites and a maintenance window, and
buys independence from an abandoned codebase sooner.
2. **If migrating, which?** On the corrected requirement the ranking is **Garage first** — its
permission model *is* the requirement, its provisioning API is the closest match, and the
highest-risk consumer is first-party supported. SeaweedFS second, on field record, with a
translation-layer caveat that matters more here than its feature list. RustFS not yet, on age.
Taking (1) does not foreclose (2), and that asymmetry is the argument for taking (1) first.
## What is still unmeasured
1. **Whether any of the ten buckets needs object versioning, server-side encryption or lifecycle.**
This gates Garage specifically and nothing else here answers it.
2. **Whether the fork's release can actually read the incumbent's on-disk format in place.** The
format is claimed compatible across a multi-year gap; the migration between them is one-way, so
this is tested on a copy or not at all.
3. **Whether the opaque consumer's maintenance window is acceptable**, and how long it actually is
at 82,496 objects.
4. **Whether the eight-drive erasure-coded shape is warranted at all.** The evidence above says it
is not buying what it appears to: eight directories, one array, one machine. It looks inherited.
**Nothing graduates to a decision before 1 and 2.**
## Two errors in the previous version of this document
Recorded because the shape of both survives anonymisation and neither is unique to this effort.
**It ranked on a requirement that did not exist.** Console single sign-on was treated as the axis
because the module's configuration showed it wired up, and a wired-up configuration was read as a
used feature. It was neither used nor working. *A configured feature is not an observed one*, and
the evidence needed was the operator's answer and the logs — both cheap, neither consulted before
the ranking was written.
**It omitted the incumbent's own fork.** The whole effort began because an upstream withdrew its
images; whether anyone had continued that upstream was the first question to ask and it was not
asked. The candidate list was assembled from a search for *alternatives*, which by construction
returns things that are not the incumbent. *When a dependency dies, "who took it over" precedes
"what replaces it".*