Files
mesh-catalog/modules/keycloak
jochen 1fd2914ae1 Say which modules wait for a person's push, and announce the files a merge deleted (hq ADR 0236)
With a gate on the first machine and a rollback after it, a module's build rolls out by
default. The ones kept back say why: the network path a rollback could not cross, the
providers every consumer on a machine drops with, and the stores holding the photos.
A merge's deleted files are announced, so a module whose manifest went is forgotten
rather than asked to build (the public-acme plan failure).
2026-10-06 18:39:13 +02:00
..

keycloak

The mesh's identity provider: one Keycloak server that provides the oidc-client provision to every module that logs a person in. Each consumer is given one confidential OpenID Connect client in the realm named by the assignment's issuer setting, under the client id the mesh derived for it and the secret the mesh minted (ADR 0048); its redirect is its contributed callback under the names the mesh composed for its endpoint. A client the mesh did not make — no mesh.provisioned attribute — is never adopted, changed or deleted.

The admin keeps the mesh's password

The manifest mints the admin own-secret, and the server takes it from its environment only when it creates its master realm. A database that was adopted, restored or moved already has one, and its admin keeps the password it had. Every admin call then fails with 401 invalid_grant, and with it every consumer's client. On 2026-10-05 that went on for a day, about 31,000 failures, seen only in the journal (hq issue 179). Both times the fix was the same, done by hand.

The module now does that fix itself. The guard in the provider:

  • checks that admin logs in with the mesh's secret: at start (every 15 s until the server answers), then every 5 minutes, and at once when the admin API refuses the credentials;
  • on a refusal — invalid_grant, including a missing or disabled admin — and only then, repairs it inside the keycloak container with Keycloak's own recovery: kc.sh bootstrap-admin user makes a temporary admin (on a free management port; the server holds 9000), kcadm.sh creates or re-enables admin if it must and sets its password to the mesh's, and the temporary admin is deleted. Both passwords go in on the exec's standard input. Neither is on a command line or printed;
  • checks again, and says REPAIRED the admin … in the log and emits admin.repaired;
  • when it cannot, logs COULD NOT REPAIR … with the step that failed, emits admin.unrepaired, and brakes: the next automatic attempt comes 10 minutes later and the wait doubles each time, up to 6 hours. If the temporary admin may be left behind, the log and the event say so.

While the admin is refused, the provisioner does not call Keycloak. Each attempt would be one more failed admin login, and enough of those lock the account. It still counts each attempt as a failure of the consumer, so the provider's standing (below) reports credentials-rejected.

keycloak_admin_check reports the state, the last repair and the brake. With repair: true it repairs a refused admin straight away, ignoring the brake, because a person asked.

A consumer failing for minutes is announced

The provisioner loop (harness.go) is shared, byte for byte, with postgres. A consumer whose create, check or secret keeps failing for 5 minutes with no success in between is announced as provisioner.failing, with the consumer, its machine and the error's class. The announcement repeats every 15 minutes while the failure lasts. provisioner.recovered follows the first success (ADR 0224), and the controller shows the latest one in status.

A consumer the mesh no longer asks for is retired, not deleted (novox/hq ADR 0230): its client is disabled — Keycloak refuses its authorization and token requests — and marked with mesh.retired and mesh.retired-why; its secret, redirects and mappers are kept, and asked for again it is enabled as it was. Retiring waits five passes and ten minutes, and for a person's retire approve when it is more than three clients or more than half of those held. Only cleanup delete removes a client, and only a disabled one the mesh made. The tools provisioner_retirement, provisioner_retire_approve, provisioner_retire_reject and provisioner_delete are what the controller's retire and cleanup verbs ask. MESH_KEYCLOAK_LIVE_URL and MESH_KEYCLOAK_LIVE_PASSWORD run live_retire_test.go against a throwaway server.

Tools

The realm, user, client, group and role tools (keycloak_list_realms, keycloak_create_user, keycloak_list_clients, keycloak_assign_user_role, …), and keycloak_admin_check. A tool that writes something announces it: user.created, user.deleted, password.reset, client.created, group.created, role.created.

Where the code lives

One Go bundle, cmd/keycloak-provider, launched by the node's runtime. It speaks MCP over stdio through the Go SDK, and was ported from TypeScript in 2026-10. It reaches the server on MESH_KEYCLOAK_URL, reads the mesh's admin secret from MESH_KEYCLOAK_PASSWORD_FILE at every check, and reaches the container named by MESH_KEYCLOAK_CONTAINER through the container-runtime capability.

Tests

go test ./... runs against a fake Keycloak and a fake container. It covers:

  • repair on refusal, with no password in argv;
  • no repair while the server is unreachable;
  • the brake, and the operator overriding it;
  • the provisioner going quiet while the admin is refused;
  • the OIDC client rules;
  • the harness and its standing;
  • harness_same_test.go, which fails when this module's harness.go and postgres's differ.

live_test.go runs the repair against a real, throwaway Keycloak 26 container whose admin keeps an older password. The file's comment has the commands.