Files
hq/01-RESEARCH/015-the-object-store-after-minio/00-overview.md
T
jochen 1d524a1fa8 Research 015: rewrite the comparison — wrong axis, and a missing candidate
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.

It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.

On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.

Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.

Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
2026-09-24 18:46:47 +02:00

9.9 KiB

status, initiated, touches
status initiated touches
active 2026-09-24
02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
02-DECISIONS/0078-the-store-and-broker-are-modules.md
02-DECISIONS/0084-which-provider-serves-a-consumer.md
03-DESIGN/00-as-is/03-provisioning.md
03-DESIGN/01-to-be/07-the-foundation.md
04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md

015 — The object store after MinIO: which S3 implementation, and how the data moves

The question. The mesh's object store is MinIO. Its community edition is archived upstream, its server and client images have been deleted from every public registry, and the pinned release is four and a half years old and will never be patched (issue 113). Which S3-compatible implementation replaces it, and what is the migration track for the data and the provisioning model that sit on top of it?

Why now, and why not sooner. Nothing is on fire: nodes that already hold the images keep running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present image by design. The forcing function is not an outage but a one-way door — no node that does not already hold the images can ever provision the module again, so the mesh's ability to stand a node up from its declarations is already broken for this module, and silently.

The direction is not a departure from the design; it is the design. The foundation document already states the commitment:

The dependency is on the protocol, not the product: AMQP for the bus, S3 for the object store, the OCI protocol for the registry. That is what keeps the naming safe rather than a commitment that cannot be revisited.

The object store is also not a foundation service — ADR 0028 removed it, and it is an ordinary module required through the module graph by whatever wants one. (The "exception that is not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which is the cheapest kind of decision to make and the strongest kind to cite.

What the replacement has to carry, measured

Taken from the module's manifest, its composition, its tool surface, and a search for its consumers across the catalogue — not from assumption.

Requirement Evidence in the module today
S3 API The protocol every consumer speaks; already the design's stated dependency.
OIDC login against the mesh's identity provider Struck 2026-09-24. Not a requirement, and it never worked. Six variables are wired and an entrypoint blocks on the provider, which reads as a live feature. The module's own hook comment records the end state as "policy claim missing" — a failing login. See 01.
Per-application access keys, each scoped to a bucket The real requirement. A "user" of the store is normally an application; the mesh already mints a credential per provisioned bucket.
One live consumer using it as opaque primary storage A file-sync application, since early 2023: objects named by internal id, metadata in its own database. Highest-risk consumer — a live copy drifts, and its bucket name must be preserved.
Erasure-coded multi-node topology Four server nodes with two data directories each, behind a load balancer.
A single-node form Declared as a flavour, for development and small nodes.
Buckets as a typed provision The module declares a provision type of bucket on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084).
A tool surface Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning.
A console Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads.

Consumers, counted: one application module, one capture module that takes a private bucket per node, one workflow module's tools, and the delivery/rescue internals of the shared library. The surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the OIDC story, not spread across the catalogue.

Candidates

Four candidates, not three. The comparison was briefly narrowed to SeaweedFS on the strength of console single sign-on; that axis turned out not to be a requirement, and the incumbent's own maintained fork had been omitted altogether. Both errors, and why they happened, are recorded in 01 — the candidates measured, which carries the evidence and the requirement-by-requirement detail.

In short, and only in short:

  • The maintained fork of the incumbent — the community edition was archived and its images deleted, but a fork publishes, tracks CVEs, and preserves the on-disk format, S3 API and environment surface. Costs an image reference where every other option costs a data migration, two rewrites and a maintenance window. Does not end the dependence on an abandoned codebase; buys time to choose deliberately.
  • Garage — its permission model is the requirement (per access key, per bucket), its admin API is the closest match to how the mesh provisions, and the highest-risk consumer is first-party documented against it. Remaining cost: no object versioning, no server-side encryption or object locking, partial lifecycle — unmeasured against the ten buckets, and the one thing that could still disqualify it.
  • SeaweedFS — longest field record and erasure coding. Its console sign-on is a paid feature, which is now beside the point. What weighs against it is narrower: its S3 surface is a gateway translating onto its own file-system API, with no first-party support for the opaque consumer.
  • RustFS — closest in shape to the incumbent, so the least porting. But it reached general availability eight days before this was written, and carries an open defect in the credential path. Two earlier claims about it are corrected in 01: it is not a drop-in that retains existing data.
  • Ceph RGW — remains rejected as disproportionate where the object store is an ordinary module rather than a platform.

This is now two decisions, not one: whether to repoint to the fork or migrate, and — if migrating — to which. Repointing does not foreclose migrating, which is the argument for taking it first. On the corrected requirement the migration ranking is Garage, then SeaweedFS, and not yet RustFS. Two measurements gate any graduation: which S3 endpoints the consumers actually call (Garage cannot be ranked fairly until counted), and whether the fork can read the incumbent's on-disk format in place — tested on a copy, because the migration between them is one-way. Both are in 01.

The migration track, in outline

Data movement is the easy half, and deliberately reversible.

  1. Stand the replacement up beside the incumbent, on its own ports, its own provision type and its own data directory. Nothing removed. The data directory matters: reusing one the incumbent already holds would put a fresh single-drive store on top of a live erasure set.
  2. Copy bucket by bucket with a neutral tool. rclone rather than the incumbent's own client — the client has been withdrawn upstream too, so building the migration on it would inherit the same dependency this effort exists to remove.
  3. Verify per bucket — object counts and checksums, not a transfer exit code.
  4. Repoint consumers through the connection the module already publishes. Consumers read an API URL from the module's declared connections rather than addressing the store directly, so the cutover surface is that value plus the provisioning and tool handlers.
  5. Freeze writes, final incremental sync, flip, and keep the incumbent read-only as the rollback until confidence is earned. For the opaque consumer this is not optional and not instant: it stores objects by internal id with metadata in its own database, so a copy taken while it runs will drift. It needs a maintenance window for the final sync, and the window is proportional to 82,496 objects rather than to 230 GiB.
  6. Retire, and only then remove the module.

The genuinely new work is not the copy. It is the provisioning handler and the tool handlers, both written against the incumbent's admin API. The OIDC wiring was previously listed here and is struck: it is not a requirement and it never worked.

Open questions

  • How much of the OIDC requirement survives, and in which build? Answered, and it was the wrong question. The console requirement does not exist, and the login it referred to never worked. What replaced it: which S3 endpoints consumers actually call, and whether the fork reads the incumbent's format in place.
  • Does the mesh's bucket provision translate to the candidate's identity model without weakening what ADR 0049 says about a consumer's identity fitting the tightest backend?
  • Should this effort also answer issue 113's general question — mirroring third-party images into the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
  • Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking while the product is being chosen, rather than reproducing a shape by default.