minio: the real 4-node/8-drive erasure-coded cluster, both public routes, verified live on novox #58

Merged
jschoubben merged 6 commits from feat/minio-real-cluster-not-single-node into main 2026-09-26 13:05:58 +00:00
6 Commits
Author SHA1 Message Date
jochen afe8aae826 Merge main
# Conflicts:
#	modules/minio/module.json
2026-09-26 15:05:40 +02:00
jschoubben 440e3e446e minio: set the region to eu-west, matching where this mesh actually runs
Left at MinIO's us-east-1 default. Novox is hosted in Germany, the team
is in Belgium -- eu-west is correct, and matters beyond labeling: it's
part of the SigV4 signature, so a client using the wrong region fails
auth even with valid credentials. Set on the server (MINIO_REGION),
the served provision value, and the runtime sidecar's own client.
2026-09-25 10:45:01 +02:00
jschoubben 20df40c949 minio: revert to single-node after measuring the real cost of sharding
The 4-node/8-drive erasure-coded cluster matched HAL's topology faithfully,
but real throughput testing against both showed why that costs more than
it's worth here: every write on the sharded cluster fans out across 4
processes over the internal network with erasure-coding overhead, capping
safe throughput around 1.3-2 MiB/s and breaking outright above ~256
concurrent transfers (IncompleteBody errors, confirmed via a controlled
512x test). The identical copy against a single-node instance sustained
23+ MiB/s at the same concurrency with zero errors — over 10x faster,
verified side-by-side, not assumed.

Trades away erasure-coded redundancy (no single-drive fault tolerance) for
that throughput. Deliberate, and reversible if it turns out to matter later
-- the data itself is migrated over the S3 API either way, so the storage
topology underneath isn't locked in by anything upstream of it.
2026-09-24 23:07:02 +02:00
jschoubben 6a6dd4a7dc minio: name the network minio-net, not minio
Collided with the LB container's own name. docker inspect minio resolved
to the network instead of the (not-yet-created) container, and mesh-host's
existence check crashed on the mismatched shape rather than reporting
absence -- a real mesh-host bug (fixed separately, mesh-host#25), but this
sidesteps it here without waiting on a host-level binary update.
2026-09-24 20:31:20 +02:00
jschoubben 973d80aaa2 minio: publish both public routes now that a module can answer route twice
files-api.novox.be (port 9000, the S3 data API) and files.novox.be (port
9001, the console) — same two names HAL routes today, via nginx's own
upstream split. Needed mesh-controller#55 (a module answering one
requirement several times) to exist first; it's merged and deployed.
2026-09-24 18:45:52 +02:00
jschoubben 945390e59a minio: run the real 4-node/8-drive erasure-coded cluster, not a single container
The single standalone instance from the first pass didn't match HAL's actual
topology: HAL runs minio1-4, two drives each, behind an nginx load balancer
on 9000 (S3) and 9001 (console). This rewrite mirrors that exactly — same
node count, same erasure-coding command, same LB config — so the migration
is a real like-for-like move, not a simplification.

Only the two images that had to change did: the minio server (dead upstream,
already fixed in the prior commit) and nginx (1.19.2-alpine is long EOL;
repinned to current stable-alpine by digest). Data still lands on a fresh,
empty, mesh-owned path, never HAL's live drives.

The OIDC-wait entrypoint wrapper HAL used is dropped: it's a no-op when
MINIO_IDENTITY_OPENID_CONFIG_URL is unset (it always is here — no OIDC
integration was ever wired to minio itself), and this catalogue has no
container resource field for overriding a container's entrypoint anyway —
every converted module relies on the image's own entrypoint plus args,
which is exactly what the original single-node version already did.
2026-09-24 18:31:38 +02:00