From 1d524a1fa825478d7eb8ace42315ec56769934bc Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 24 Sep 2026 18:46:47 +0200 Subject: [PATCH] =?UTF-8?q?Research=20015:=20rewrite=20the=20comparison=20?= =?UTF-8?q?=E2=80=94=20wrong=20axis,=20and=20a=20missing=20candidate?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The previous version ranked candidates on whether they preserved single sign-on to the object store's console. That is not a requirement: a "user" of the store is normally an application, so the requirement is per-application keys scoped to buckets — which the mesh already mints. And the console login it ranked on never worked; the module's own hook comment records "policy claim missing", a failing login written up as progress. It also omitted the incumbent's own maintained fork, which changes the question from "which product replaces it" into two decisions: repoint, or migrate — and if migrating, to which. Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a maintenance window. Repointing does not foreclose migrating, which is the argument for taking it first. On the corrected requirement Garage ranks first — its per-key-per-bucket model is the requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk consumer is first-party documented against it. Its remaining gap (no versioning, no server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one thing that could still disqualify it. Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive directories on one filesystem on one machine. That last fact decides more than any feature — the erasure coding is not buying independent-drive redundancy, so the redundancy model is close to irrelevant and only storage overhead remains, which at this volume is a rounding error against the headroom. Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is not an observed one; and when a dependency dies, "who took it over" precedes "what replaces it" — searching for alternatives by construction returns things that are not the incumbent. --- .../00-overview.md | 73 +++--- .../01-candidate-comparison.md | 235 +++++++++++------- 2 files changed, 197 insertions(+), 111 deletions(-) diff --git a/01-RESEARCH/015-the-object-store-after-minio/00-overview.md b/01-RESEARCH/015-the-object-store-after-minio/00-overview.md index 6c628fb..88707ee 100644 --- a/01-RESEARCH/015-the-object-store-after-minio/00-overview.md +++ b/01-RESEARCH/015-the-object-store-after-minio/00-overview.md @@ -50,7 +50,9 @@ consumers across the catalogue — not from assumption. | Requirement | Evidence in the module today | |---|---| | S3 API | The protocol every consumer speaks; already the design's stated dependency. | -| **OIDC login against the mesh's identity provider** | Six configuration variables are wired and populated in practice — discovery URL, client id, client secret, scopes, display name, redirect — plus a dedicated entrypoint script that blocks startup until the provider answers. This is live, not aspirational. | +| ~~OIDC login against the mesh's identity provider~~ | **Struck 2026-09-24. Not a requirement, and it never worked.** Six variables are wired and an entrypoint blocks on the provider, which reads as a live feature. The module's own hook comment records the end state as *"policy claim missing"* — a failing login. See [01](01-candidate-comparison.md). | +| **Per-application access keys, each scoped to a bucket** | The real requirement. A "user" of the store is normally an application; the mesh already mints a credential per provisioned bucket. | +| **One live consumer using it as opaque primary storage** | A file-sync application, since early 2023: objects named by internal id, metadata in its own database. Highest-risk consumer — a live copy drifts, and its bucket name must be preserved. | | Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. | | A single-node form | Declared as a flavour, for development and small nodes. | | Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). | @@ -64,40 +66,49 @@ OIDC story, not spread across the catalogue. ## Candidates -**The comparison is open across three candidates.** It was briefly narrowed to SeaweedFS; that -narrowing did not survive measurement and was reopened on 2026-09-24. The evidence, the full -requirement-by-requirement table and what each option costs are in -[01 — the candidates measured](01-candidate-comparison.md). +**Four candidates, not three.** The comparison was briefly narrowed to SeaweedFS on the strength of +console single sign-on; that axis turned out not to be a requirement, and the incumbent's own +maintained fork had been omitted altogether. Both errors, and why they happened, are recorded in +[01 — the candidates measured](01-candidate-comparison.md), which carries the evidence and the +requirement-by-requirement detail. In short, and only in short: -- **SeaweedFS** — Apache-2.0, the longest field record, erasure coding. Its **console OIDC is an - Enterprise feature**; the free build authenticates the console with a local username and - password. This is what falsified the original narrowing. -- **Garage** — the best match for how the mesh provisions, with a first-class admin REST API - scoped to exactly bucket and key creation. Costs: replication rather than erasure coding, no - full S3 endpoint span, and neither console nor identity integration. -- **RustFS** — the only candidate preserving the live behaviour, with a MinIO-shaped console and - documented OIDC against Keycloak, under Apache-2.0. Costs: it reached GA eight days before this - was written, and an open defect is reported in the credential path the mesh's bucket provision - depends on. +- **The maintained fork of the incumbent** — the community edition was archived and its images + deleted, but a fork publishes, tracks CVEs, and preserves the on-disk format, S3 API and + environment surface. Costs **an image reference** where every other option costs a data + migration, two rewrites and a maintenance window. Does not end the dependence on an abandoned + codebase; buys time to choose deliberately. +- **Garage** — its permission model *is* the requirement (per access key, per bucket), its admin + API is the closest match to how the mesh provisions, and the highest-risk consumer is + first-party documented against it. Remaining cost: no object versioning, no server-side + encryption or object locking, partial lifecycle — **unmeasured against the ten buckets, and the + one thing that could still disqualify it**. +- **SeaweedFS** — longest field record and erasure coding. Its console sign-on is a paid feature, + which is now beside the point. What weighs against it is narrower: its S3 surface is a gateway + translating onto its own file-system API, with no first-party support for the opaque consumer. +- **RustFS** — closest in shape to the incumbent, so the least porting. But it reached general + availability eight days before this was written, and carries an open defect in the credential + path. Two earlier claims about it are corrected in 01: it is **not** a drop-in that retains + existing data. - **Ceph RGW** — remains rejected as disproportionate where the object store is an ordinary module rather than a platform. -**The question that decided the original narrowing has been answered** — SeaweedFS's console OIDC -is not in the free build — and the answer was *"no candidate preserves the current feature set -for free"*, exactly the outcome this effort said was possible. The decision is therefore not -which product is best in the abstract but **which cost is acceptable**, and two measurements -gate it: whether an authenticating proxy is an acceptable answer to console single sign-on, and -which S3 endpoints consumers actually call. Both are named in -[01](01-candidate-comparison.md#what-is-still-unmeasured). Nothing graduates before them. +**This is now two decisions, not one:** whether to repoint to the fork or migrate, and — if +migrating — to which. Repointing does not foreclose migrating, which is the argument for taking it +first. On the corrected requirement the migration ranking is Garage, then SeaweedFS, and not yet +RustFS. Two measurements gate any graduation: **which S3 endpoints the consumers actually call** +(Garage cannot be ranked fairly until counted), and **whether the fork can read the incumbent's +on-disk format in place** — tested on a copy, because the migration between them is one-way. Both +are in [01](01-candidate-comparison.md#what-is-still-unmeasured). ## The migration track, in outline Data movement is the easy half, and deliberately reversible. -1. **Stand the replacement up beside the incumbent**, on its own ports and its own provision type. - No downtime, nothing removed. +1. **Stand the replacement up beside the incumbent**, on its own ports, its own provision type and + **its own data directory**. Nothing removed. The data directory matters: reusing one the + incumbent already holds would put a fresh single-drive store on top of a live erasure set. 2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client — the client has been withdrawn upstream too, so building the migration on it would inherit the same dependency this effort exists to remove. @@ -106,16 +117,22 @@ Data movement is the easy half, and deliberately reversible. API URL from the module's declared connections rather than addressing the store directly, so the cutover surface is that value plus the provisioning and tool handlers. 5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the - rollback until confidence is earned. + rollback until confidence is earned. For the opaque consumer this is **not optional and not + instant**: it stores objects by internal id with metadata in its own database, so a copy taken + while it runs will drift. It needs a maintenance window for the final sync, and the window is + proportional to 82,496 objects rather than to 230 GiB. 6. **Retire**, and only then remove the module. The genuinely new work is not the copy. It is the **provisioning handler** and the **tool -handlers**, which are written against MinIO's admin API, and the OIDC wiring. +handlers**, both written against the incumbent's admin API. *The OIDC wiring was previously listed +here and is struck: it is not a requirement and it never worked.* ## Open questions -- How much of the OIDC requirement survives, and in which build? See above — this gates the - choice. +- ~~How much of the OIDC requirement survives, and in which build?~~ **Answered, and it was the + wrong question.** The console requirement does not exist, and the login it referred to never + worked. What replaced it: which S3 endpoints consumers actually call, and whether the fork reads + the incumbent's format in place. - Does the mesh's bucket provision translate to the candidate's identity model without weakening what ADR 0049 says about a consumer's identity fitting the tightest backend? - Should this effort also answer issue 113's general question — mirroring third-party images into diff --git a/01-RESEARCH/015-the-object-store-after-minio/01-candidate-comparison.md b/01-RESEARCH/015-the-object-store-after-minio/01-candidate-comparison.md index f2e00ab..d5fba1d 100644 --- a/01-RESEARCH/015-the-object-store-after-minio/01-candidate-comparison.md +++ b/01-RESEARCH/015-the-object-store-after-minio/01-candidate-comparison.md @@ -1,107 +1,176 @@ -# 015 / 01 — The candidates measured, and why the first pick did not survive +# 015 / 01 — The candidates measured -*Dated 2026-09-24. This document reopens the comparison the overview had narrowed to one -candidate. It exists because the requirement the narrowing rested on turned out not to be met.* +*Rewritten 2026-09-24. An earlier version of this document ranked the candidates on whether they +preserved single-sign-on to the object store's **console**. That was the wrong axis — it is not a +requirement — and a fourth candidate was missing entirely. Both errors are recorded at the end, +because how a comparison came to be ranked on the wrong thing is worth more than the ranking was.* -## What was asked, and what came back +## The requirement, corrected -The overview named one question as deciding the choice: **how much of SeaweedFS's OIDC story is -in the freely licensed build, and can its shape stand in for a console that redirects a human to -an identity provider.** SeaweedFS was scoped as primary *because it was believed to be the only -candidate that preserved that*. The answer is no, and the premise fails with it. +Taken from the operator and from the running system, not from the module's shape. -| SeaweedFS capability | Apache-2.0 build | Enterprise (proprietary, per-TB) | -|---|---|---| -| OIDC on the S3 API, via STS/JWT token exchange | **yes** | yes | -| Admin UI sign-in | **username and password only** | **OIDC SSO** (Keycloak and others) | +**A "user" of the object store is normally an application.** The requirement is therefore +**per-application access keys, each scoped to its own bucket** — not per-human single sign-on. The +mesh already works this way: it mints a credential for every provisioned bucket, and the consumer +reads an endpoint from the module's declared connection rather than addressing the store directly. -The admin UI is Apache-2.0 and real; its **identity provider integration is not**. It sits behind -the commercial licence together with point-in-time recovery, automatic erasure-code repair, -native multi-tenancy and S3 QoS limiting. +**The console is not a requirement.** It was the axis the previous version ranked on, and it should +not have been. -So the free build gives OIDC for *programmatic* access and a locally-authenticated console. The -behaviour in use today — a human redirected to the mesh's identity provider to reach the object -store's console — is not preserved. The narrowing was sound reasoning on wrong information. +**The identity-provider login never worked.** The predecessor's module wires six OIDC variables and +blocks startup until the provider answers, which reads like a working feature. It is not: the +module's own hook comment records the end state as *"policy claim missing"* — a **failing** login, +written up as progress because it proved the provider had registered. The identity provider emits no +such claim, nothing in the module creates the mapper, and the configured scope alone would not carry +a custom one. Two days of logs show no genuine login attempts, only internet scanners failing on an +STS API version. **Nothing should be carried forward on the assumption this works**, and no +candidate should be credited or penalised for matching it. -## The three-way comparison, as measured +**One consumer is live, opaque, and holds real user files.** A file-sync application has used the +store as its **primary storage** since early 2023: objects named by an internal id, with all +metadata in its own database. Three consequences — a copy taken while it runs will drift, its bucket +name must be preserved or its database references break, and it is the highest-risk consumer of the +lot. -Against the requirement set in the overview, which was taken from the module rather than assumed. +## What is actually stored, measured -| | SeaweedFS | Garage | RustFS | -|---|---|---|---| -| Licence | Apache-2.0 | AGPL-3.0 | Apache-2.0 | -| Maturity | twelve-plus years | several years, stable | **GA 2026-09-16** | -| Redundancy | erasure coding | **replication only** — a replica count, commonly 3, so ~3× raw per byte | erasure coding | -| S3 API | broad | **explicitly not the full span of endpoints** | broad, MinIO-shaped | -| Console | `weed admin`, in the free build | **none official** (a third-party UI exists) | yes, modelled on MinIO's | -| Console OIDC | Enterprise only | **none** | **yes** — documented against Keycloak | -| Programmatic bucket + credential creation | `weed shell` / IAM API | **full admin REST API**, scoped tokens for `CreateBucket` and `CreateKey` | admin API under its own `v3` namespace; client-compatible with MinIO's | -| Permission model | IAM users, groups, policies | **its own**: per access key, per bucket, read/write/owner — no AWS-style ACLs or bucket policies | MinIO-shaped IAM: users, groups, policies, service accounts | +| | | +|---|---| +| Logical | **230 GiB, 82,496 objects, 10 buckets** | +| Raw on disk | **468 GiB** — eight drive directories at 59 GiB each | +| Implied scheme | 468 ÷ 230 = **2.03×**, confirming erasure coding at half parity | +| Headroom | ~1.3 TiB free on the filesystem holding it | -**No candidate is free of cost.** That is the finding, and it is why the comparison is reopened -rather than resolved here. +**All eight "drives" are directories on one filesystem on one machine.** The erasure coding is +therefore not buying independent-drive redundancy; the real failure domain is the array underneath, +which has its own. This single fact decides more of the comparison than any product feature: a +scheme's redundancy model is close to irrelevant here, and what remains is its storage overhead. -## What each one actually costs +At 230 GiB with 1.3 TiB free, **storage overhead is not a deciding cost either.** Replication at +three copies would run ~690 GiB against the present 468 GiB — about **+222 GiB**, comfortably +absorbed. Erasure coding at a wider stripe would *save* roughly 146 GiB. Both are rounding errors +against the headroom, and neither should decide this. -**SeaweedFS** — the safe engineering choice. Longest field record, erasure coding, permissive -licence. The cost is the console regression: either accept local credentials for the console, pay -per TB, or front it with an authenticating proxy (see below). +## The candidates -**Garage** — the best fit for how the mesh *provisions*. Its admin API is a first-class REST -surface with scoped tokens for exactly the two operations the bucket provision performs, which is -a better match than any other candidate. Three costs, and they are not small: redundancy is -replication, so the storage bill is a multiple rather than erasure coding's overhead; it does not -implement the full S3 endpoint span, which has to be checked against what consumers actually call; -and there is neither a console nor identity integration, so the console requirement is not -regressed but *removed*. +Four, not three. The previous version omitted the first. -**RustFS** — the only candidate that preserves the live behaviour. Apache-2.0, erasure coding, a -MinIO-shaped console with **documented OIDC against Keycloak**, and an admin API close enough to -MinIO's that the existing tool handlers are the least work to port. Against that: +### The maintained fork of the incumbent -- **It reached GA on 2026-09-16 — eight days before this was written.** For a component that - holds data, field record is a feature, and it does not have one yet. -- **The overview's characterisation of it was wrong and is corrected here.** It was recorded as a - binary-level drop-in retaining existing data, which would make migration trivial. In fact API - compatibility and on-disk compatibility are *separate* paths, and the on-disk one is - preview-scoped with documented encryption limitations. It should not be relied on. This costs - nothing in practice — the migration track already moves data over the S3 API, not on disk — but - the claim should not survive into a decision record. -- **There is an open defect in precisely the area the mesh depends on.** An access key created by - an OIDC/SSO user is reported denied on all S3 operations, the embedded policy being ignored. - The mesh mints credentials for every provisioned bucket. Whether this touches the mesh's path — - which provisions with administrative credentials rather than as an SSO user — is **unverified, - and must be tested before RustFS is chosen**, not after. +The community edition was archived upstream and its images deleted +([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)), +but **a fork is maintained and publishing** — `pgsty/minio`, from the Pigsty project. It restores +the console stripped from the community +build, rebuilt image and package distribution, tracks CVEs, and states that it preserves the on-disk +format, the S3 API and the environment-variable surface. Verified by pulling it: it reports a +current release, permissive-to-copyleft licensing unchanged from upstream, and identifies itself as +a community fork. Adoption is real — the server image has been pulled three quarters of a million +times. -**Ceph RGW** stays rejected, for the reason already recorded: disproportionate where the object -store is an ordinary module rather than a platform. +**Why it reorders the comparison.** Every other candidate costs a data migration, a provisioning +handler rewritten against a different admin API, a tool surface ported, and a maintenance window for +the opaque consumer. The fork costs **an image reference**. It also closes the issue's one-way door: +a node holding nothing can provision the module again, and patches resume. -## The option that changes the trade +**What it does not do** is end the dependence on a codebase its original authors abandoned. It is +maintenance mode, largely one project's effort, with no new features intended. It buys time to +choose deliberately rather than under pressure — which is worth a great deal, and is not the same as +a decision. -The console regression is not necessarily the product's to solve. An authenticating proxy in -front of whatever console exists — against the identity provider the mesh already runs, behind -the reverse proxy it already runs — restores single sign-on at the proxy layer for **SeaweedFS or -Garage**, without a commercial licence and without betting on the youngest candidate. +### Garage -If that holds, the decision stops being *which regression to accept* and becomes a straight -comparison of storage properties, where SeaweedFS's field record and erasure coding are the -strongest position. **It is unverified.** It is the cheapest next measurement available and it -should be made before the record is written. +**The best fit for how the mesh provisions.** Its permission model is *per access key, per bucket, +read/write/owner* — which is the requirement above stated verbatim rather than approximated. Its +admin API is a first-class REST surface with tokens scopeable to exactly the two operations a bucket +provision performs. The opaque consumer is **first-party documented** against it, for primary +storage, including client-side encryption support. + +Its previously-recorded penalties mostly dissolve under the corrected requirement: it has no console +and no identity-provider integration, neither of which is wanted; and it replaces AWS-style ACLs and +bucket policies with its own per-key-per-bucket model, which is the thing being asked for. + +**What genuinely remains.** It replicates rather than erasure-codes — immaterial at this volume and +on a single array, as above. It does **not implement the full span of S3 endpoints**: object +versioning is absent, object locking and server-side encryption endpoints are absent, and lifecycle +is partial. **Whether any of the ten buckets depends on those is unmeasured, and it is the one thing +that could still disqualify it.** + +### SeaweedFS + +Longest field record of the group, permissive licence, erasure coding, and identity-provider +integration on the S3 API through token exchange. **Its console sign-on is a paid feature** — the +admin UI itself is open, its identity integration is not. That finding is what falsified the +previous version's narrowing, and it is now largely beside the point, since the console is not a +requirement. + +What weighs against it here is narrower and more specific: its S3 surface is a **gateway +translating onto its own file-system API**, with acknowledged divergence from AWS behaviour at the +edges, and there is no first-party documentation for the opaque consumer. For a store already +holding real user files in an opaque layout, first-party support is worth more than a feature list. +It also carries more moving parts than a single-machine deployment needs. + +### RustFS + +Closest in shape to the incumbent — a similar admin API and client compatibility, so the existing +handlers would port with least effort — under a permissive licence, with erasure coding and a +console that does integrate an identity provider. + +Two things were recorded about it earlier that were **wrong, and are corrected here**: it is *not* a +binary-level drop-in that retains existing data (API compatibility and on-disk compatibility are +separate paths, and the on-disk one is preview-scoped with documented encryption limits), and it +therefore offers no shortcut around the migration. It also carries **an open defect in the exact +area the mesh depends on** — an access key created by an identity-provider user reported denied on +all S3 operations. + +Decisively for now: **it reached general availability eight days before this was written.** For a +component holding 230 GiB of real user files, field record is a feature, and it does not have one. + +### Ceph RGW + +Remains rejected, for the reason already recorded: disproportionate where the object store is an +ordinary module rather than a platform. + +## Where this leaves it + +**The decision is no longer "which product replaces the incumbent".** It is two decisions, and they +can be taken in either order but should not be confused: + +1. **Repoint to the maintained fork, or migrate now?** Repointing is an image reference and it + closes the issue. Migrating now costs a data copy, two rewrites and a maintenance window, and + buys independence from an abandoned codebase sooner. +2. **If migrating, which?** On the corrected requirement the ranking is **Garage first** — its + permission model *is* the requirement, its provisioning API is the closest match, and the + highest-risk consumer is first-party supported. SeaweedFS second, on field record, with a + translation-layer caveat that matters more here than its feature list. RustFS not yet, on age. + +Taking (1) does not foreclose (2), and that asymmetry is the argument for taking (1) first. ## What is still unmeasured -Named so the effort is not closed while pretending otherwise. +1. **Whether any of the ten buckets needs object versioning, server-side encryption or lifecycle.** + This gates Garage specifically and nothing else here answers it. +2. **Whether the fork's release can actually read the incumbent's on-disk format in place.** The + format is claimed compatible across a multi-year gap; the migration between them is one-way, so + this is tested on a copy or not at all. +3. **Whether the opaque consumer's maintenance window is acceptable**, and how long it actually is + at 82,496 objects. +4. **Whether the eight-drive erasure-coded shape is warranted at all.** The evidence above says it + is not buying what it appears to: eight directories, one array, one machine. It looks inherited. -1. Whether an authenticating proxy in front of a console is acceptable as the mesh's answer to - console SSO — a design question as much as a technical one, since it moves identity out of the - product and into the edge for this module. -2. Which S3 endpoints the consumers actually call, checked against Garage's supported span rather - than assumed. Until that is counted, Garage cannot be fairly ranked. -3. Whether RustFS's OIDC-credential defect touches the credential path the bucket provision uses. -4. What the redundancy change costs in real terms — replication against erasure coding at the - sizes actually stored, rather than as a ratio in the abstract. -5. Whether the four-node erasure-coded topology is still warranted at all, which the overview - already raised and nothing here answers. +**Nothing graduates to a decision before 1 and 2.** -**No candidate should be graduated to a decision until 1 and 2 are measured.** 3 gates RustFS -specifically. The effort stays open. +## Two errors in the previous version of this document + +Recorded because the shape of both survives anonymisation and neither is unique to this effort. + +**It ranked on a requirement that did not exist.** Console single sign-on was treated as the axis +because the module's configuration showed it wired up, and a wired-up configuration was read as a +used feature. It was neither used nor working. *A configured feature is not an observed one*, and +the evidence needed was the operator's answer and the logs — both cheap, neither consulted before +the ranking was written. + +**It omitted the incumbent's own fork.** The whole effort began because an upstream withdrew its +images; whether anyone had continued that upstream was the first question to ask and it was not +asked. The candidate list was assembled from a search for *alternatives*, which by construction +returns things that are not the incumbent. *When a dependency dies, "who took it over" precedes +"what replaces it".*