Commit Graph
100 Commits
Author SHA1 Message Date
jschoubben 30b7ce429c automx: the schema its seed writes belongs to one automx2, so that one is named
An unpinned pip install took the latest automx2, whose schema grew a
column (server.prio) the module's own seeding SQL predates — a 500 on
every autoconfig request against a table the seed had just written.
2021.6 is what the proven image runs; the seed and the software agree
again. Regenerating the seed for a newer automx2 is its own change,
made deliberately, not by whatever pip resolved this week.
2026-09-25 23:59:42 +02:00
jschoubben b8a50ee7a0 Merge pull request 'mailu: the containers are named by the vocabulary 2024.06 reads' (#79) from fix/mailu-speaks-2024-06-addresses into main 2026-09-25 21:56:40 +00:00
jschoubben 4e37b3e84d mailu: the containers are named by the vocabulary 2024.06 reads
The front resolves its upstreams from *_ADDRESS, defaulting to the bare
compose service names — admin, antispam — and reads the 1.9-era HOST_*
not at all. Phase 1 never noticed because the predecessor's service
names WERE the defaults; the mesh's containers are mailu-*, and the
front answered 502 asking docker for a name nothing carries. The dead
vocabulary goes; every upstream is named as the container actually is.
2026-09-25 23:56:28 +02:00
jschoubben 7cb7659fbe Merge pull request 'mailu: every container asks the module's own resolver, pinned where they can find it' (#78) from feat/mailu-points-at-its-own-resolver into main 2026-09-25 21:49:30 +00:00
jschoubben 047f228fe3 mailu: every container asks the module's own resolver, pinned where they can find it
2024.06's admin refuses to serve behind a resolver that does not
validate DNSSEC — found live as an unhealthy admin, 454s on submission
and a 500 webmail, with the runtime's forwarder validating nothing. The
unbound this module always shipped becomes reachable: pinned at the
predecessor's own address on the module network (the subnet the mesh
adopted), and named as dns by the nine containers that resolve anything.
Stands on mesh-host #26, which gave the vocabulary these two fields.
2026-09-25 23:49:17 +02:00
jschoubben 40aa93dc2b Merge pull request 'mailu: the admin API answers on 8080 since 2024.06' (#77) from fix/mailu-admin-answers-on-8080 into main 2026-09-25 21:45:11 +00:00
jschoubben 5462317183 mailu: the admin API answers on 8080 since 2024.06
1.9's admin served on 80; 2024.06's gunicorn listens on 8080, and the
runtime's fetch failed against the old port the moment the broker
credential let it try.
2026-09-25 23:45:00 +02:00
jschoubben cb9d43b2a1 Merge pull request 'automx: the launcher's wrapper is written by the build that ships it' (#76) from fix/automx-carries-its-own-wrapper into main 2026-09-25 21:44:46 +00:00
jschoubben aa8c4253f1 automx: the launcher's wrapper is written by the build that ships it
The predecessor's image carried .venv/scripts/flask.sh, created by a
build step that never made it into the files this module holds — the
image worked and its recipe could not reproduce it, caught the moment
the mesh built it from source (exit 127 crash loop at cutover). The
wrapper is now written explicitly, verbatim from the proven image, so
the recipe is the whole truth about the image.
2026-09-25 23:44:34 +02:00
jschoubben 4863eef586 Merge pull request 'mailu: pin exactly what phase 1 verified' (#75) from fix/mailu-pins-what-phase-one-verified into main 2026-09-25 21:34:26 +00:00
jschoubben 75f993e92b mailu: pin exactly what phase 1 verified
Phase 1 (MAILU-CUTOVER.md) upgraded the live stack 1.9→2024.06 and its
lessons land here as pins: every image is the digest running and
verified tonight — webmail under the name 2024.06 actually uses (the
draft's roundcube pin misled a whole hop), antivirus on the upstream
clamav image with the signature DB in its own directory (the mailu-built
image ended at 2.0), and the front trusting the proxy address the live
config actually names. The welcome-mail texts ride along for parity.

TLS_FLAVOR stays letsencrypt deliberately where the live .env says cert:
the cert files are copies whose HAL-era renewal hook died with HAL
(expiry Nov 27); mailu managing its own issuance through the existing
ACME passthrough is the fix, and the files remain on disk as the
fallback flavor if first issuance misbehaves during the window.
2026-09-25 23:34:13 +02:00
jschoubben 621246996d Merge pull request 'mailu: 2024.06 closes 110/143/587 by default; parity says open them' (#74) from fix/mailu-ports-parity into main 2026-09-25 20:56:39 +00:00
jschoubben e3d8bd0726 mailu: 2024.06 closes 110/143/587 by default; parity says open them
PORTS defaults to 25,80,443,465,993,995,4190 in 2024.06 — submission on
587 among the closed, which is what every client of this server uses.
The same parity decision the listens already state, now stated where the
software reads it.
2026-09-25 22:56:27 +02:00
jschoubben a708666bad Merge pull request 'mailu: the manifest matches the machine, provides smtp, and carries automx' (#73) from feat/mailu-becomes-real into main 2026-09-25 20:40:08 +00:00
jschoubben ed5d1386ce mailu: the manifest matches the machine, provides smtp, and carries automx
Five gaps between the draft and what actually runs, each verified live
before being written down:

- front published bare 80 — the machine port Traefik holds; now the
  predecessor's own mappings (7080:80, 7443:443) plus the 110/143/995
  parity ports the draft dropped. Pruning legacy protocols is its own
  deliberate change, not a cutover side effect.
- TLS_FLAVOR said cert, which nothing supplies; live is letsencrypt —
  mailu runs its own certbot, state already on disk, HTTP-01 answered
  through a path-scoped route contribution (priority above the web one).
- the web route said http:7080, the redirect-loop shape; it now says
  what the hand-authored file always knew: https 7443, insecure.
- automx was absent entirely: the autoconfig responder is now a second
  artifact (its Containerfile moved in from the predecessor's images
  dir, base declared per ADR 0097), a container on a real data dir —
  the anonymous-volume loss of 2026-08-10 stays fixed — and the three
  public names are route contributions.
- and the reason this moved ahead of de-spiegel: mailu now provides
  smtp. A consumer contributes the account it sends as; the provisioner
  creates <account>@<domain> via the admin API and applies the minted
  password every reconcile (ADR 0048). The domain is served on the
  binding so a consumer composes its own login from mesh facts.

route-adapter learns to say no: a contribution over https, scoped to a
path, or carrying a policy is skipped aloud rather than written into a
file shape that cannot say it — plain http into a TLS listener was the
concrete wrong file this prevents. The hand-authored files keep covering
those routes until the mesh's own proxy takes over, exactly as today.
2026-09-25 22:39:29 +02:00
jschoubben 3875987656 Merge pull request 'gitea: the package team may read code, and its units are reconciled' (#72) from fix/gitea-package-team-reads-code into main 2026-09-25 19:59:25 +00:00
jschoubben 311f7f1fdb gitea: the package team may read code, and its units are reconciled
The builder's first credentialed clone of a private repository answered
'not found': the packages team named only repo.packages in its
units_map, which is exhaustive — so members had no code unit at all, and
gitea hides what a user cannot read. One credential answering npm and
git alike was the whole design of the builder's grant; the team now says
so.

And found teams are patched, not just returned: a team is configuration
the reconcile loop owns, the same as a user's password, so a unit this
code gains reaches the team that already exists rather than only the
next mesh raised from scratch.
2026-09-25 21:59:09 +02:00
jschoubben bbd1723612 Merge pull request 'route-proxy binds the mesh's own authority for internal names; gitea's internal-API refusal joins its route' (#70) from feat/route-proxy-internal-acme into main 2026-09-25 19:55:32 +00:00
jschoubben c7964b7285 Merge pull request 'gitea: a consumer's user is actually created, and a failed create says why' (#71) from fix/gitea-provisioner-user-create into main 2026-09-25 19:54:08 +00:00
jschoubben 02ccf31a71 gitea: a consumer's user is actually created, and a failed create says why
The builder's package-registry grant — the first this provider ever
received — retried for a day saying only that an edit 404'd. Two faults
under it: the API refuses an email without a dotted domain, so
`@localhost` failed validation at create (the CLI that made mesh-admin
accepts it, which is why the admin exists and no consumer did); and
ensureUser read that 422 as 'already exists' and went on to edit a user
that was never made, burying the create's own message. The address is now
gitea's own hidden-address shape, and the edit path is taken only for a
user that is actually there.
2026-09-25 21:42:38 +02:00
jschoubben c2353fc0a6 gitea: the internal-API refusal is part of the route, not a file beside the proxy
The 2026-09-12 incident response blocked /api/internal by hand in the
predecessor's dynamic directory, with a note that its durable home is the
mesh's routing. A route carries the policy applied to a request (ADR
0108), so the refusal now travels with the grant: route-proxy enforces
it on both the public name and the internal alias the moment it serves
this route, and the adapter skips it aloud (no port, nothing to write)
while the predecessor's own file still stands. The hand-authored file
retires with the proxy it configures.
2026-09-25 20:51:44 +02:00
jschoubben 962cba7c04 route-proxy: internal names are certified by the mesh's own authority
Two name spaces, two authorities (08-connectivity §2): a public name is
certified by a public CA, an internal one by the mesh's own. step-ca now
offers that second seat as internal-acme-ca beside its existing acme-ca,
and route-proxy requires both — the server dispatches by which authority
may certify the name at all, so an .internal alias stops being plain-HTTP
only without ever asking a public CA for a name it cannot validate.
2026-09-25 20:36:41 +02:00
jschoubben bfe99f8c78 Merge pull request 'minio: declare the route contribution it has always needed, scoped from PR #58' (#68) from fix/minio-declares-its-real-route-contribution into main 2026-09-25 16:19:39 +00:00
jschoubben f0aa9e5fed minio: declare the route contribution it has always needed, scoped from PR #58
files-api.novox.be and files.novox.be worked earlier tonight from route-
adapter-generated files, but minio's module.json on main never actually
carried a route requirement — that capability has been sitting in PR #58
the whole time, bundled with an unrelated network rename and console
redirect URL that need their own calmer review. This is just the two
routes: requires: route, contributes.route.api/.console (ContributesMany,
proven working via mesh-controller #55/#57), and the console port (9001)
actually published and declared in listens.

Found assigning route-proxy for the first time tonight: its own routes
file, generated the identical way route-adapter's always was, had four
hostnames in it instead of six — nothing served files-api/files at all,
which would have been a real, silent outage the moment Traefik stopped.
2026-09-25 18:19:27 +02:00
jschoubben cb38bf08a6 Merge pull request 'route-proxy: the trust step skips the fetch when the CA names no roots to get' (#67) from fix/route-proxy-trust-skips-when-there-is-nothing-to-fetch into main 2026-09-25 16:15:09 +00:00
jschoubben 95a0a5672c route-proxy: the trust step skips the fetch when the CA names no roots to get
Composed ACME_ROOTS unconditionally from ${bound:acme-ca:roots} even when
that field is empty — public-acme's own case, where an empty roots means
'the system trust store', not 'fetch from the bare authority host'. The
run-once step wget'd https://acme-v02.api.letsencrypt.org:443 (host, no
path) for two minutes every apply and failed, blocking every resource
after it — found live tonight, assigning route-proxy for the first time.

Carries the raw, uncomposed roots value alongside the composed URL
(ACME_ROOTS_PATH) so the step can tell 'nothing to fetch' apart from 'the
authority didn't answer' — a distinction the composed URL alone cannot
make. Empty copies the image's own system CA bundle to /ca/root.crt
instead of fetching one, so ACME_CA_BUNDLE stays the one path it has
always been rather than needing to become conditional itself.
2026-09-25 18:14:58 +02:00
jschoubben d243b56942 Merge pull request 'route-proxy: the server container resolves its own built artifact' (#66) from fix/route-proxy-server-uses-its-real-artifact into main 2026-09-25 16:09:19 +00:00
jschoubben 4c5e69903f route-proxy: the server container resolves its own built artifact
Was still pinned to the scaffold's placeholder digest (mesh-route-
proxy@sha256:0000...0000) even after the build+context work landed —
never caught because nothing had assigned route-proxy before tonight.
artifact: server, matching trust's own reference a few lines up and
every other built module in the catalogue.
2026-09-25 18:09:09 +02:00
jschoubben 547937034f Merge pull request 'public-acme: the roots field is named roots, not root' (#65) from fix/public-acme-field-name-matches-what-route-proxy-reads into main 2026-09-25 16:08:39 +00:00
jschoubben 3147fac08b public-acme: the roots field is named roots, not root
step-ca (the other acme-ca provider) already spells it correctly; route-
proxy's own template reads ${bound:acme-ca:roots}. Found live, assigning
public-acme for the first time tonight: the mesh refused the push outright
rather than composing a broken binding — 'route-proxy asks its acme-ca
binding for roots, and what answers it says ... root'. Empty stays empty:
a public CA's root is the system trust store already, per route-proxy's
own design (an empty ACME_CA_BUNDLE means exactly that).
2026-09-25 18:08:28 +02:00
jschoubben 7a96287d99 Merge pull request 'route-proxy: declare the build context its own Dockerfile has always needed' (#64) from feat/route-proxy-declares-its-real-context into main 2026-09-25 15:58:20 +00:00
jschoubben dd5b973e66 route-proxy: declare the build context its own Dockerfile has always needed
The Dockerfile's own comment already said it — 'the build context is the
mesh-controller repository root' — but nothing in the manifest actually
said so to the mesh, so every build attempt used mesh-catalog's own
directory instead and failed with 'stat go.mod: file does not exist'.
Never caught before because route-proxy has never been assigned anywhere.
Uses the context mechanism just added (mesh-controller#62), proven
working tonight on builder's own self-build.
2026-09-25 17:58:10 +02:00
jschoubben 72b1413497 Merge pull request 'builder is a real built module now, not handed over' (#63) from feat/builder-is-a-real-built-module into main 2026-09-25 15:54:58 +00:00
jschoubben 8369fe22b8 builder is a real built module now, not handed over
Its own image ('mesh-builder@sha256:0000...0000', later manually pinned to
a real digest tonight when the placeholder blocked a push) was never
produced by anything the mesh tracks — cmd/mesh-builder lives in
mesh-controller's own repository, and nothing declared how to build an
image from it. Uses the same context mechanism route-proxy does (mesh-
controller#62): the Dockerfile compiles ./cmd/mesh-builder from a clone of
mesh-controller's repository, not a vendored copy.

Unlike mesh-controller's own FROM scratch (ADR 0006 — nothing to audit
but one binary), the build machine's whole job is shelling out to git and
docker, so its runtime is Alpine with both installed from the base's own
packages, not fetched on their own.

Bootstrapped live tonight: a manual build got the new image running long
enough to build itself properly through the pipeline it had just gained,
and mesh-controller itself needed the same upgrade first (it parses
manifests too, and rejected the new context field with the old binary) —
genesis's own kind of ordering problem, solved by hand exactly once.
2026-09-25 17:54:46 +02:00
jschoubben 19e0a8469d Merge pull request 'minio: install mc in the runtime image, declare its region' (#62) from fix/minio-provisioner-missing-mc-cli into main 2026-09-25 15:25:18 +00:00
jschoubben 271911753f Merge pull request 'builder: consume package-registry as a real mesh grant, not a hand-faked one' (#61) from fix/builder-real-package-registry-grant into main 2026-09-25 15:25:04 +00:00
jschoubben 4e2dd0adae Merge pull request 'gitea: listens its real internal ssh port (22), not 2222' (#60) from fix/gitea-ssh-port-manifest-mismatch into main 2026-09-25 15:24:49 +00:00
jschoubben 1b02dc1f66 builder: qualify its own image with the registry host
Bare mesh-builder@sha256:... is only resolvable for a module with a build
section — the mesh's own build step rewrites the reference to a real
registry path as part of resolving build.artifacts. builder is handed
over, not built, so nothing ever rewrites it: pushed as written, docker
read it literally and tried Docker Hub. Took the live node's mesh-builder
down for the length of one push-and-fix (docker: pull access denied for
mesh-builder, repository does not exist). novox.internal:5100, not the
literal external IP docker inspect showed live, for the same reason
addresses generally don't get hardcoded in this catalogue.
2026-09-25 17:10:56 +02:00
jschoubben 755b0a5599 builder: consume package-registry as a real mesh grant, not a hand-faked one
The 'package-binding' resource was a hardcoded JSON fragment standing in
for a real grant — {"provision": "package-registry", "from": "gitea",
"at": "127.0.0.1", ...} written as if it were mesh-resolved, when nothing
resolved it. Declares requires: package-registry properly instead, with
binds/secrets pointing at the same file paths the resource used to
manually author, so the mesh mints the grant and writes it there.

npm-password renamed to package-registry.secret: it's gitea's generic
user+password, not npm-specific — the same credential works for basic
auth against cargo/PyPI/Go package endpoints too, once gitea's manifest
grows them (novox/hq ADR 0109).

Known gap, not fixed here (novox/hq issue 117): this makes builder
correct for the steady state but breaks a genesis bootstrap — gitea's own
image is built by builder, so builder cannot yet hold this grant the
first time either has to exist. Filed rather than silently accepted.
2026-09-25 17:08:12 +02:00
jschoubben ae6dfea2c9 gitea: listens its real internal ssh port (22), not 2222
2222 never meant anything — no container of gitea's publishes it, the
software never listens on it, and nobody could say where it came from.
The container's real internal sshd is 22 (gitea's own default, unmodified
— every other module's listens.port already means the container's real
internal port, this one didn't). ports now declares 222:22 directly: 222
is the real, fixed, always-known public git-ssh port, so it needs no
per-node setting to reach — unlike an arbitrary auto-assigned port, this
one has external consumers who already know the number.
2026-09-25 17:07:58 +02:00
jschoubben b3de7e7944 minio: declare region eu-west, don't leave it as live-only drift
serves.s3-bucket.region still said us-east-1 (PR #58 already fixed this,
unmerged) while the live mesh-minio sidecar had MESH_MINIO_REGION=eu-west
set out-of-band, not in the manifest at all — and the actual minio server
had no region configured whatsoever (mc admin config get region: empty),
apparently lost across a container recreation since nothing declared it.
Every s3-bucket consumer binding ${bound:s3-bucket:region} was reading
the stale us-east-1 declaration regardless of what was actually live.

Declares MINIO_REGION on the server container and MESH_MINIO_REGION on
the sidecar, matching the static serves declaration, so this is mesh-
managed and durable rather than a manual mc admin config or docker env
override that the next recreation silently drops.
2026-09-25 15:10:30 +02:00
jschoubben 5d13a5078c minio: install mc in the runtime image — the provisioner needs it to run
mesh-minio's s3-bucket provisioner shells out to mc to create buckets and
service accounts on the live server, but mc was never in this module's own
runtime image, only in minio's own. It's been silently retrying 'spawn mc
ENOENT' forever, so every s3-bucket grant reached the control-plane layer
(store.json, sealed secret) without the credential ever actually existing
on minio — nextcloud's live instance just hit this as InvalidAccessKeyId
on a real user session.

Copies mc from minio's own image (docker.io/pgsty/minio, already pinned
and pulled as this module's server container) rather than introducing a
new base — mc there is a working, already-verified binary. /usr/bin/mc is
a symlink to mcli; both are copied so it resolves.
2026-09-25 14:37:23 +02:00
jschoubben 61eca75315 nextcloud: pull the S3 region from the binding, not a hardcoded literal
the s3-bucket binding already carries serves.region (same mechanism as
at/port); ${bound:s3-bucket:region} tracks whatever minio is actually
configured with instead of a copy that can drift
2026-09-25 14:15:40 +02:00
jschoubben df0e9017e0 nextcloud: set OBJECTSTORE_S3_REGION to match minio's actual region
minio runs with MINIO_REGION=eu-west; nextcloud's S3 config never set a
region, so every object write (avatars, file writes) failed signature
validation with AuthorizationHeaderMalformed, surfacing as Internal Server
Error on real page loads
2026-09-25 14:12:05 +02:00
jschoubben b60dd5e3e3 nextcloud: use a module-owned mesh-admin account instead of 'admin'
the migrated data has no literal 'admin' account — HAL's real admin login
is a personal account (jochens), not a generic one. Resetting that would
touch a real user's own credential, so the module gets its own dedicated
admin-group service account instead, same pattern as the minio per-module
service accounts. mesh-admin was created once by hand on novox to match
this manifest for the already-migrated data; a genuinely fresh install
seeds it automatically via NEXTCLOUD_ADMIN_USER/NEXTCLOUD_ADMIN_PASSWORD.
2026-09-25 13:51:40 +02:00
jschoubben b865978e19 nextcloud: use a named build stage for docker:cli, not ARG-in-COPY-from
the legacy (non-BuildKit) docker build this host runs doesn't expand ARGs
inside COPY --from — only FROM. Give it its own named stage instead
2026-09-25 13:49:02 +02:00
jschoubben cdbc850c52 nextcloud: declare docker:cli as a pinned build.on base
build refused to reach docker:cli implicitly (novox/hq ADR 0097); pin it by
digest and thread it through as DOCKER_CLI, redeclared in the final stage
since args declared before the first FROM don't carry past it
2026-09-25 13:48:00 +02:00
jschoubben e2b3723dd9 nextcloud: fix hardcoded sidecar port and missing docker CLI in runtime image
- MESH_NEXTCLOUD_URL hardcoded :80 instead of the mesh-assigned ${port:80}
- sidecar's occ() shells to docker exec but the docker CLI binary was never
  present in the runtime image, only the mounted socket
2026-09-25 13:47:03 +02:00
jschoubben 5f74c41a3e nextcloud: give the admin password a real _FILE variant, not an env-file
The mesh's own check caught it: an env-file-loaded secret still reaches
the process environment, readable via docker inspect and /proc (hq
04-ISSUES/041) -- the same class of exposure the file-based delivery
exists to avoid. Added MESH_NEXTCLOUD_ADMIN_PASSWORD_FILE support to the
client, matching the pattern the minio client already uses, and mounted
the sealed admin secret directly rather than writing it into an env-file.
2026-09-25 13:42:21 +02:00
jschoubben 4329ca9392 nextcloud: deliver the admin password to the sidecar too
The sidecar's own client needs MESH_NEXTCLOUD_ADMIN_PASSWORD to list
shares over the OCS API, but the runtime container's env/volumes never
carried it -- only the server container did. Delivered the same way every
other sealed value in this manifest already is: a generated env-file with
the ${secret:admin} substitution, not a raw value in the container's env.
2026-09-25 13:41:17 +02:00
jschoubben 066254e5c4 nextcloud: pin an image whose PHP matches this Nextcloud version
The pinned digest resolved to a PHP 8.5.10 image; Nextcloud 30 refuses to
run above PHP 8.4. Repinned to the current digest for the nextcloud:30 tag
(matches HAL's own NEXTCLOUD_VERSION), which carries PHP 8.3.28 -- the
same version the data being migrated was actually running under.
2026-09-25 13:39:16 +02:00
jschoubben 49c5903861 Merge pull request 'minio: repin to pgsty's fork, build the runtime sidecar, move data off HAL's live drive' (#57) from fix/minio-repin-and-move-off-hal-data into main 2026-09-24 16:15:09 +00:00
jschoubben 35e8a3fe19 minio: repin to pgsty's fork, build the runtime sidecar, move data off HAL's live drive
minio/minio and minio/mc were pulled from all public registries on 2026-09-11;
pgsty's fork is the working replacement (hq issue 113). The runtime sidecar
had no Dockerfile and no build section at all — added, following the same
tsc-over-client/tools/provisioner shape every other converted module uses.

The data resource pointed straight at /services/minio/data/data1-1, one of
HAL's live 8-drive erasure-coded array — starting this module would have
written into production storage the mesh doesn't own. Moved to a fresh,
empty, mesh-owned directory; the actual data migration happens over the S3
API (rclone), not by sharing a disk path.
2026-09-24 17:55:49 +02:00
jschoubben 5c3781d8c2 Merge pull request '4 modules: persistent data is a directory bind, never a named volume' (#54) from fix/persistent-data-is-a-directory-bind-not-a-volume into main 2026-09-24 14:24:15 +00:00
jschoubben 6a3e0ab9d3 4 modules: persistent data is a directory bind, never a named volume
hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a
directory bind (./db-data, ./data) for exactly this data -- the mesh's
adoption of them, three weeks ago, switched to a named Docker volume
instead, and searxng/distribution followed the same pattern since.

A named volume survives ordinary container recreation, same as a
directory bind -- that was never the problem. The problem is everything
else: docker rm -fv, docker volume rm, and docker system prune --volumes
all target it (one flag away from the docker rm -f this migration already
uses routinely); it is invisible to every tool this migration has used all
night (ls, find, grep across /var/lib, /services); and nothing outside
Docker's own volume machinery can back it up or notice it growing.

mesh-store carries the sharpest version: every database migrated tonight,
including keycloak's, live inside it.

Data already copied and verified on novox before this merges:
- mesh-registry-data -> /var/lib/mesh-registry (11G, diff -rq clean)
- mesh-broker-data -> /var/lib/mesh-broker (37M, cp -a)
- mesh-broker-tls -> /var/lib/mesh-broker-tls (12K, cp -a)
- mesh-store-data -> /var/lib/mesh-store (1.7G) -- mesh-store stopped
  cleanly first, so the final copy is crash-consistent, not a live-file
  copy of a running postgres; diff -rq clean after.
- searxng/valkey: not yet assigned anywhere, manifest-only fix, nothing
  to copy.

Old named volumes left in place, not deleted, as the rollback path.
2026-09-24 16:12:03 +02:00
jschoubben 3e8ecfa4eb Merge pull request '11 modules: a container's ports say the software's port, not the machine's' (#53) from fix/091-ports-machine-side-is-the-mesh-to-assign into main 2026-09-24 13:53:34 +00:00
jschoubben 12f885ffc5 11 modules: a container's ports say the software's port, not the machine's
hq issue 091, measured 2026-09-22: 14 of 46 modules with containers fix the
machine side of a published port in their own manifest -- a fact about one
machine (which port HAL happened to publish it on) written into a
definition meant for any node. ADR 0038 already says the mesh assigns the
machine side and a module says only what it needs; the machinery already
does it (internal/inventory/ports.go's PortFor, declaration.go's
publishedOn rewrites a bare port automatically). These 14 just never
complied.

Fixed 11 of them -- stripped to the bare software port, letting assignment
take over: de-spiegel, gitea, hello-web (the demo module), mailu, mssql,
n8n, novox.be, only-office, photos, photos-eef, photos-filip.

Left alone, the two defensible kinds the issue names: postgres/lavinmq/
distribution (foundation, genesis-rewritten per ADR 0100 -- the number in
the manifest is a default, not a claim) and unifi (protocol/
device-discovery fixes the number; nothing else can find the controller).

gitea was a live one, not just a tidiness fix: its manifest said 2222:22,
but the actual adopted, running container is on 222. Harmless while held,
but the next 'take' would have recreated it on the wrong port and broken
SSH git access. Recorded 222 as novox's own setting for it (settings set
gitea -node novox) so that doesn't happen.
2026-09-24 15:45:58 +02:00
jschoubben 7551a65579 Merge pull request 'postgres: the provisioner connects to mesh-store's actual port, not a hardcoded 5432' (#52) from fix/postgres-provisioner-connects-to-the-actual-store-port into main 2026-09-24 13:38:45 +00:00
jschoubben 135101ce6e postgres: fix the port override properly — a twin variable, not a raw substitution
The first commit on this branch embedded ${seat:mesh-store:5432} directly
inside MESH_PROVISION_POSTGRES's URL. That placeholder resolves to an empty
string whenever mesh-store is on its own default port (novox/hq
internal/catalogue/seat_into.go: 'the mesh raised on the catalogue's own
ports never gives them a setting at all') -- which produces a malformed
connection string (host:/postgres) on exactly the common case, a fresh,
non-adopted mesh. It only worked here because novox's mesh-store happens to
be adopted at a non-default port.

mesh-controller's own manifest already has the right shape for this --
internal/envfile/port.go's NAME / NAME_PORT twin, composed by
mesh-controller's Placed(): the base value keeps its own port; a separate
_PORT variable carries the override, spliced in only when it says
something, and left alone -- not a fault -- when it's an unfilled
placeholder (a manifest ahead of the running controller).

Reverted the manifest to its original base value, added
MESH_PROVISION_POSTGRES_PORT as the twin, and taught client.ts (the only
consumer -- both tools/ and provisioner/ import it) the same precedence
envfile.Placed uses. Checked: no other file in the module reads
MESH_PROVISION_POSTGRES directly.
2026-09-24 15:27:16 +02:00
jschoubben 088022d63b postgres: the provisioner connects to mesh-store's actual port, not a hardcoded 5432
MESH_PROVISION_POSTGRES was a literal connection string naming port 5432 --
correct only when mesh-store happens to run on the mesh's own default. On
novox, mesh-store was adopted in place at HAL's original port (6852), and
the provisioner has been retrying-and-failing against 127.0.0.1:5432 ever
since, for every consumer including ones that already exist (gitea, umami,
mesh-catalog), not just a new grant.

Fixed with the same ${seat:mesh-store:5432} template mesh-controller's own
manifest already uses for the identical connection. No other module needed
this fix checked -- lavinmq's MESH_PROVISION_LAVINMQ already used
127.0.0.1:15672 unconditionally, but the broker's management port is fixed
by the module itself (127.0.0.1:15672 in lavinmq/module.json's own ports),
not by adoption, so it isn't the same bug.
2026-09-24 15:24:07 +02:00
jschoubben 415ad7a168 Merge pull request 'gitea: the token needs read:user, not just write:repository and write:issue' (#51) from fix/gitea-user-repos-scope into main 2026-09-24 09:49:36 +00:00
jschoubben aaf341a2d8 gitea: the token needs read:user, not just write:repository and write:issue
Deployed #49 and the watcher immediately broke: GET /user/repos answered 403,
'required=[read:user]' — confirmed live against the running forge (1.27.3).
That route sits under gitea's user scope category despite listing
repositories, not repository as assumed.

Also gives the fake forge real scope enforcement on /user/repos, which is
why the original PR's test suite didn't catch this: it only checked the
token's value was valid, never that it carried the required scope.
2026-09-24 11:43:46 +02:00
jschoubben 8a458aa5c4 Merge pull request 'The forge mints its own API token with the admin account the vault delivers' (#49) from fix/gitea-mints-its-token into main 2026-09-24 08:41:13 +00:00
jschoubben d036b43bfc Merge main into fix/gitea-mints-its-token to pick up the resolver conversion and artifact-store seat fix 2026-09-24 10:40:57 +02:00
jschoubben 5a94b78891 Merge pull request 'Convert the resolver from the module it replaces: forward, answer at 127.0.0.1, point the runtime at it' (#50) from convert/dnsmasq-from-hal into main 2026-09-23 23:13:08 +00:00
jschoubben 3d81bf41c6 Convert the resolver from the module it replaces: forward, answer at 127.0.0.1, point the runtime at it
Read against hal/modules/dnsmasq-app (hal dnsmasq-app conversion, hq 08-connectivity).
On the machines it runs, the predecessor's dnsmasq answers every name: the mesh's own
itself, the rest forwarded to 1.1.1.1 and 8.8.8.8, its module's defaults; resolv.conf
names it alone at 127.0.0.1, and the container runtime's dns is the machine's tunnel
address, so the host and every container resolve the world through it. The nox module
forwarded nothing, listened on 127.0.0.55 — a convention of its own beside the one every
machine already followed — and read a machines file that the module's `facts` already
asks the mesh for, so taking it would have left an adopted machine with a resolv.conf
pointing at an address nothing answered on, and no upstream for anything else.

Now the resolver keeps `no-resolv` (the documented loop — finding its own address in
resolv.conf and becoming its own upstream — stays impossible) and forwards to the same
two explicit upstreams; listens on mesh0 and 127.0.0.1, which systemd-resolved does not
hold; requires `mesh-addressing`, since its data is the mesh's addresses; and writes the
runtime's `dns` into daemon.json beside whatever the machine had (ADR 0102), at this
machine's own address — `${machine:address}`, new in the controller. The runtime is not
restarted for it: it reads the key at start, not on reload, and a restart stops every
container; on the machine this replaces the value is already there.

resolv-conf names the resolver alone, as the predecessor's file did; its placeholder second
line was a fallback nothing ever reached. resolved-split-dns follows the address. The
mesh's suffix as a local domain comes with the machines file, so a mesh name the resolver
does not know is refused here rather than asked upstream. mDNS is not carried: no module
does it and the design says mesh names are not multicast names.
2026-09-24 01:09:23 +02:00
jschoubben ed094031fd The forge mints its own API token with the admin account the vault delivers
The runtime beside the forge served 0 tools: GiteaClient.fromEnv required a token
(settings or MESH_GITEA_TOKEN), nobody had one to give — the mesh raised the forge —
and putting one in settings would store a secret in plaintext in the inventory. So
the fifteen tools registered nothing and the watcher logged "not watching".

What the mesh does deliver is the admin account: a login the manifest names and a
password the vault minted and the host unsealed into a file (ADR 0086). That is
enough to mint a token, so the module does (hq issue 100, the forge's tools):
POST /users/{admin}/tokens over basic auth, scoped to write:repository and
write:issue — the least the tools and the repo watcher need — kept at 0600 in the
module's own state (/var/lib/mesh/gitea/state, a new directory resource the runtime
mounts writable), read back on the next start, and minted afresh when the forge
answers 401 to it or the kept file is gone. A forge whose data came from the
predecessor has no mesh-admin: that is reported in plain words on every poll until
it clears, once per reason, not crash-looped. A configured token still wins and is
never minted over.

The mint happens on the first call, not at registration: a contributor is
synchronous, and a forge not yet answering must not keep the runtime from serving.
One source per kept file in a process, or the watcher and the tools would each
renew on a 401 and drop the other's token by name.
2026-09-23 23:48:36 +02:00
jschoubben 4d8be01c5f Merge pull request 'The artifact store's seat is one per mesh' (#48) from fix/the-artifact-store-is-one-per-mesh into main 2026-09-23 21:42:25 +00:00
jschoubben 833d8ec3ec The artifact store's seat is one per mesh
Review of the registry work found it node-scoped: a second `distribution` on another
machine resolved cleanly there, and only afterwards did the mesh notice `artifact-store`
offered by two nodes, with every consumer elsewhere refusing to choose. Worse, a node-scoped
requirement with one candidate installs that candidate, so anything that wanted the store
beside it would have raised a fresh, empty store on the wrong machine first.

There is one store in a mesh, which was already the effective rule; the claim now says it
where a second one is assigned, by name, instead of leaving consumers to discover it.
2026-09-23 23:40:54 +02:00
jschoubben 9f2c678355 Merge pull request 'Convert mssql from the module it replaces, not from a blank page' (#46) from convert/mssql-from-hal into main 2026-09-23 18:22:51 +00:00
jschoubben cae0d00a7d Convert mssql from the module it replaces, not from a blank page
Read against hal/modules/mssql: the predecessor publishes 4848:1433 and its connections
block hands every consumer localhost:4848; its data has lived in /services/mssql/data
since it was installed — 4.6 GB of it. The nox module had 1433, db-data, and a pin three
builds behind, so taking it would have moved the port every consumer was told about,
started on an empty directory, and downgraded the engine at the same time.

db-data was not a convention: only this module and mongodb used it, and both predecessors
say data. Now nothing moves at the cutover.
2026-09-23 20:22:34 +02:00
jschoubben c001e054fd Merge pull request 'Pin the analytics module to the version the machine it replaces is running' (#45) from fix/umami-matches-production into main 2026-09-23 00:48:13 +00:00
jschoubben b69edbb16f Pin the analytics module to the version the machine it replaces is running
The pin was a month old — 2026-08-20 against the 2026-09-17 image the container runs.
Third module in a row whose pin had aged into a downgrade; the pattern is hq issue 099.
2026-09-23 02:47:23 +02:00
jschoubben b0d48d5c54 Merge pull request 'Pin the package registry to the version the machine it replaces is running' (#44) from fix/verdaccio-matches-production into main 2026-09-23 00:37:31 +00:00
jschoubben a3eca8f4e6 Pin the package registry to the version the machine it replaces is running: 6.10.3, not 6.10.1
The rule the forge's cutover taught: a module that takes over a running service must not
carry an older image than the one running, or the cutover is a downgrade nobody asked for.
6.10.4 exists; moving to it is an upgrade and its own act.
2026-09-23 02:24:07 +02:00
jschoubben f11534102d Merge pull request 'Pin the forge to the version the machine it replaces is running' (#43) from fix/gitea-matches-production into main 2026-09-23 00:28:26 +02:00
jschoubben a29475617a Pin the forge to the version the machine it replaces is running: 1.27.3, not 1.22.6 2026-09-23 00:28:11 +02:00
jschoubben 176bbd6085 Merge pull request 'route-adapter: provide route by writing into the predecessor's proxy (hq ADR 0104)' (#42) from feat/route-adapter into main 2026-09-23 00:12:48 +02:00
jschoubben bd5b349a0d route-adapter: provide route by writing into the predecessor's proxy (hq ADR 0104)
A node being adopted cannot take a web module: every module reachable by
name requires route, the mesh's only provider of it binds the two public
ports, and the predecessor's proxy holds them and serves every public
name there. Stopping the predecessor to break the circle darkens every
name at once, with every certificate to re-obtain in the same window.

So this answers the same provision without binding anything. It provides
route and receives the same contributions file, and writes each
contribution as one route file where the predecessor's file provider
reads, naming the predecessor's own certificate resolver so no
certificate is asked for. It removes a file it wrote when its
contribution goes and never touches a file it did not write — the name
and a marker inside both have to say it is the mesh's.

A step, not a daemon: run-once, re-run by restart-on over the received
file and the settings. The predecessor's dynamic directory is a node
setting, because it is a fact about one machine.

Migration scaffolding with a stated end: assigned only on an adopted
node, deleted when the predecessor's proxy retires.
2026-09-23 00:11:21 +02:00
jschoubben 4f4a0750de Merge pull request 'The store keeps the vector extension the predecessor's database had' (#41) from fix/store-keeps-pgvector into main 2026-09-22 23:30:34 +02:00
jschoubben 5eb1baf5fe Keep the vector extension the predecessor's database had: the store runs the pgvector build of the same major 2026-09-22 23:30:21 +02:00
jschoubben d0fed5b982 Merge pull request 'The forge's own address follows the port the node gave it (hq issue 088)' (#40) from feat/forge-address into main 2026-09-22 23:29:47 +02:00
jschoubben cf30d4b3d0 The forge's sidecar dials the port the machine put the forge on (hq issue 088)
MESH_GITEA_URL named 3000 outright. The forge's container publishes 3000 without
fixing the machine side, so the mesh assigns it — and on a node given that port
as a setting it is the operator's number. The mapping that lets the forge go on
binding 3000 does nothing for a caller dialling the machine's loopback, so the
sidecar dialled a port nothing was listening on wherever the two differed.

${port:3000} is the mesh's answer to exactly that question, and it now resolves
in a container's environment as it always has in a file. The forge declares it
listens on 3000, so it may ask.

Still names 3000: the route contribution (hq issue 089) and the `listens` and
`ports` entries, which are the software's own number and belong there.
2026-09-22 22:36:34 +02:00
jschoubben 5706c00566 Merge pull request 'The builder's package binding is settable per node (hq issue 085)' (#39) from feat/packages-port into main 2026-09-22 21:59:57 +02:00
jschoubben b2e29871fd Protect the address the builder sends its registry password to
`at` was settable. A node setting could point the builder's package binding at
any host, and the builder sends its registry credential there as basic auth — so
a setting meant for a port was a way to hand the password to somebody else.

novox/hq 04-ISSUES/085
2026-09-22 21:56:51 +02:00
jschoubben d63f006cb8 The builder's package binding takes its port from the node, not from the manifest
The forge is raised by hand at genesis, before any module provides
`package-registry`, so the builder carries a binding instead of resolving one —
and the port in it was rewritten, as text, by the installer. A later
registration of this manifest from the catalogue put 3000 back, silently, and
pointed the builder and its registry credential at whatever holds that port.

Made settable instead: the port is a per-node setting the controller holds, and
the catalogue's number is only its default. `provision`, `from` and `as` are
protected — a setting here is about where the forge answers, never about who the
binding is with.

novox/hq 04-ISSUES/085, ADR 0100
2026-09-22 21:39:54 +02:00
jschoubben 6e88982498 Merge pull request 'Adoption mode: guards, and a filter unit that never flushes the ruleset (hq ADR 0100, 0103)' (#38) from feat/adoption-mode into main 2026-09-22 21:01:56 +02:00
jschoubben 90104d0818 Reload the filter on a rule change rather than restart it, so the node is never unfiltered in between (hq ADR 0102) 2026-09-22 19:47:43 +02:00
jschoubben 84012fab2e Make the stock nftables unit's stop delete only the mesh's table on nodes that still have it enabled (hq ADR 0100) 2026-09-22 18:06:15 +02:00
jschoubben 0c37d7389d Guard the store and management ports on adopted nodes, and load the filter through a unit that never flushes the ruleset (hq ADR 0100) 2026-09-22 17:32:32 +02:00
jschoubben f4097b3c57 Merge pull request 'n8n and baserow no longer take the shared cache (issue 081)' (#37) from multiple-fixes into main 2026-09-22 13:53:14 +02:00
jschoubben 431627c510 baserow: its data directory is 0755, as the image ships it — the cache it runs for itself does so as another user, who must reach its own directory beneath (issue 081) 2026-09-22 13:46:01 +02:00
jschoubben 93634041db baserow: its texts no longer name the shared cache it stopped taking 2026-09-22 12:34:58 +02:00
jschoubben 24cbac816d n8n and baserow no longer take the shared cache: neither can keep its keys and channels under the login it is granted — baserow runs its own, n8n in its shipped mode needs none (novox/hq issue 081) 2026-09-22 12:26:33 +02:00
jschoubben b2a48558ea Merge pull request 'redis: a consumer's ACL loses the dangerous category (issue 080)' (#36) from multiple-fixes into main 2026-09-22 02:19:14 +02:00
jschoubben c71bbd497f redis: INFO is allowed back — client libraries ask it at connect, and it reads nothing a consumer keeps 2026-09-22 02:03:23 +02:00
jschoubben 421f4d8577 redis: a consumer's ACL loses the dangerous category — a key pattern does not confine FLUSHALL (novox/hq issue 080) 2026-09-22 01:39:29 +02:00
jschoubben 54af5e8536 Merge pull request 'route-proxy: the trust gate names the binding it reads; the server names the gate (ADR 0099)' (#35) from multiple-fixes into main 2026-09-21 23:58:48 +02:00
jschoubben 6fb2003b8d route-proxy: the trust gate names the binding it reads and runs again when the authority moved; the server follows it (ADR 0099) 2026-09-21 23:25:06 +02:00
jschoubben 64d62978f6 Merge pull request 'Multiple fixes: five modules keep their secrets from the vault (ADR 0094), the authority makes its own root (076), the builder on the host network, route-proxy declares its bases' (#34) from multiple-fixes into main 2026-09-21 22:58:24 +02:00