postgres: declare the data directory's real owner; keycloak: use the port template #55

Merged
jschoubben merged 4 commits from fix/postgres-owner-and-keycloak-port-template into main 2026-09-26 13:05:10 +00:00
Owner

Two follow-ups to PR #54, both found live on novox after it merged.

postgres — the real root cause of the two mesh-store crash-loops tonight. Its data directory has always had split ownership: everything inside pgdata/ is owned by UID 999 (the pgvector image's actual runtime user), while only the top-level mount point happened to be 70:70. Confirmed against the untouched original named volume — this isn't something the directory-bind conversion introduced, it was there from the start. It was invisible while the directory's mode was 1777 (world-accessible, docker's own volume default), and broke the moment mode: 0700 (this module's own declared value) was actually enforced, locking out the process that owns everything inside it. mesh-store crash-looped twice before this was found: once at container creation, once mid-session on a checkpoint — the second time after ownership had already been checked and looked correct at the top level, which is what made it hard to find. Fixed with owner: "999:70" on the directory resource — mesh-host's idsOf() accepts a raw uid:gid, no host user needs to exist.

Checked mesh-broker and mesh-registry for the same pattern — both are uniformly owned 0:0 throughout, top-level and nested, no split. Those images run as root internally, so they were never at risk from this; no change needed there.

keycloak — same bug class as the postgres connection-string fix earlier tonight. MESH_KEYCLOAK_URL was hardcoded to :8080, but this node's port override (settings set keycloak {"ports":{"8080":28080}}) means the real published port is 28080. Fixed using the mesh's own ${port:8080} template — exactly the mechanism internal/catalogue/port_into.go describes for a sidecar dialling its own server over the machine's loopback.

Already built and live on novox; mesh-store verified stable on the correct ownership post-deploy, all four databases and keycloak's row counts unchanged.

Two follow-ups to PR #54, both found live on `novox` after it merged. **postgres — the real root cause of the two `mesh-store` crash-loops tonight.** Its data directory has *always* had split ownership: everything inside `pgdata/` is owned by UID `999` (the `pgvector` image's actual runtime user), while only the top-level mount point happened to be `70:70`. Confirmed against the untouched original named volume — this isn't something the directory-bind conversion introduced, it was there from the start. It was invisible while the directory's mode was `1777` (world-accessible, docker's own volume default), and broke the moment `mode: 0700` (this module's own declared value) was actually enforced, locking out the process that owns everything inside it. `mesh-store` crash-looped twice before this was found: once at container creation, once mid-session on a checkpoint — the second time *after* ownership had already been checked and looked correct at the top level, which is what made it hard to find. Fixed with `owner: "999:70"` on the directory resource — `mesh-host`'s `idsOf()` accepts a raw `uid:gid`, no host user needs to exist. Checked `mesh-broker` and `mesh-registry` for the same pattern — both are uniformly owned `0:0` throughout, top-level and nested, no split. Those images run as root internally, so they were never at risk from this; no change needed there. **keycloak — same bug class as the postgres connection-string fix earlier tonight.** `MESH_KEYCLOAK_URL` was hardcoded to `:8080`, but this node's port override (`settings set keycloak {"ports":{"8080":28080}}`) means the real published port is `28080`. Fixed using the mesh's own `${port:8080}` template — exactly the mechanism `internal/catalogue/port_into.go` describes for a sidecar dialling its own server over the machine's loopback. Already built and live on `novox`; `mesh-store` verified stable on the correct ownership post-deploy, all four databases and `keycloak`'s row counts unchanged.
jschoubben added 1 commit 2026-09-24 14:41:51 +00:00
postgres: mesh-store's data directory has always had split ownership --
everything inside pgdata/ is owned by UID 999 (the pgvector image's real
runtime user), while only the top-level mount point happened to be 70:70.
Invisible while the directory's mode was 1777 (world-accessible, from the
named volume this replaced); broke the moment mode: 0700 was enforced,
locking out the actual owning process. mesh-store crash-looped on
Permission denied twice before this was found -- once at container
creation, once mid-session on a checkpoint, after ownership looked correct
by every check that didn't look inside pgdata/ specifically.

keycloak: MESH_KEYCLOAK_URL was hardcoded to :8080, but the module's own
port override (settings set keycloak {ports:{8080:28080}} on novox) means
the real published port is 28080. Same bug class as the postgres
connection-string fix earlier tonight -- now using the mesh's own
 template instead, which is exactly the mechanism
internal/catalogue/port_into.go describes for a sidecar dialling its own
server over the machine's loopback.
jschoubben added 1 commit 2026-09-24 15:25:22 +00:00
Reported: files.novox.be's login button redirects to http://keycloak.novox.be,
not https. HAL's original config (/services/keycloak/docker-compose.yml) set
three settings the mesh's manifest never carried over:

  KC_HOSTNAME: keycloak.novox.be
  KC_HOSTNAME_STRICT_HTTPS: true
  KC_PROXY: edge

Without KC_PROXY: edge, Keycloak has no way to know it sits behind a
TLS-terminating reverse proxy (traefik) -- it generates URLs from what it
directly sees, which is plain HTTP from traefik's backend connection. Same
pattern as the named-volume conversion: the shape was rebuilt from general
knowledge of what a keycloak container needs, not from what this
installation's own working config actually had.
jschoubben added 1 commit 2026-09-24 15:27:09 +00:00
The previous commit on this branch used KC_PROXY=edge and
KC_HOSTNAME_STRICT_HTTPS=true, carried over from HAL's config -- but
HAL ran an older Keycloak using the v1 hostname provider. This image
(26.0.8) defaults to Hostname v2, which warned 'options [proxy,
hostname-strict-https] are still in use, please review your
configuration' and kept generating http:// URLs regardless -- verified
against /realms/Novox/.well-known/openid-configuration directly, not
just the login button, after the first fix deployed.

v2's actual shape (keycloak.org/server/hostname): KC_HOSTNAME is a full
URL, not a bare hostname -- the scheme in the URL is what tells Keycloak
to generate https, not a separate strict-https flag. KC_PROXY_HEADERS
replaces KC_PROXY: xforwarded to trust traefik's X-Forwarded-* headers,
which it sends by default.
Author
Owner

Reviewed; two of the three changes are ready, one needs your call.

Ready. ${port:8080} in the runtime's MESH_KEYCLOAK_URL replaces a hardcoded port — exactly the pattern issue 091 established. And the store's data directory gaining owner: "999:70" declares what the directory must actually be, which is what ADR 0030's ownership rule wants stated rather than inherited from whatever the runtime creates.

Needs a call. KC_HOSTNAME: "https://<label>.<public-domain>" puts one mesh's hostname into a module definition. It is not avoidable today — Keycloak generates absolute URLs and cannot derive them behind a proxy, and nothing in the manifest vocabulary yields a public name: ${machine:…} resolves at, name and address, all private-network, and the composed <label>.<public-domain> never leaves the controller. Recorded as hq issue 122, with the two other instances of the same workaround in #58 and #59.

So the options are: merge as a stopgap, with 122 carrying the mechanism; or hold this line and leave the identity provider generating wrong URLs behind the proxy until 122 has an answer. The other two changes could also be split out and merged now, leaving only the hostname waiting.

84 commits have landed on main since this branch was cut; it still merges clean.

Reviewed; two of the three changes are ready, one needs your call. **Ready.** `${port:8080}` in the runtime's `MESH_KEYCLOAK_URL` replaces a hardcoded port — exactly the pattern issue 091 established. And the store's data directory gaining `owner: "999:70"` declares what the directory must actually be, which is what ADR 0030's ownership rule wants stated rather than inherited from whatever the runtime creates. **Needs a call.** `KC_HOSTNAME: "https://<label>.<public-domain>"` puts one mesh's hostname into a module definition. It is not avoidable today — Keycloak generates absolute URLs and cannot derive them behind a proxy, and nothing in the manifest vocabulary yields a public name: `${machine:…}` resolves `at`, `name` and `address`, all private-network, and the composed `<label>.<public-domain>` never leaves the controller. Recorded as **hq issue 122**, with the two other instances of the same workaround in #58 and #59. So the options are: merge as a stopgap, with 122 carrying the mechanism; or hold this line and leave the identity provider generating wrong URLs behind the proxy until 122 has an answer. The other two changes could also be split out and merged now, leaving only the hostname waiting. 84 commits have landed on `main` since this branch was cut; it still merges clean.
jschoubben added 1 commit 2026-09-26 13:05:04 +00:00
jschoubben merged commit fa91be4941 into main 2026-09-26 13:05:10 +00:00
jschoubben deleted branch fix/postgres-owner-and-keycloak-port-template 2026-09-26 13:05:11 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: novox/mesh-catalog#55