From 1e02593adf859df662f62fb31201355f6b7be8d4 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 6 Oct 2026 00:14:38 +0200 Subject: [PATCH] Issue 179 recurred; ADR 0224: a provider that keeps failing a consumer is a problem the controller reports The identity provider's admin lost the mesh's password again when its database moved, and 31,000 silent failures followed. Record the recurrence, the rule that makes a failing provider visible in status, and the module's self-repair. --- ...mer-is-a-problem-the-controller-reports.md | 134 ++++++++++++++++++ 02-DECISIONS/README.md | 1 + 03-DESIGN/01-to-be/19-the-module-protocol.md | 17 ++- .../00-report.md | 62 +++++++- 4 files changed, 209 insertions(+), 5 deletions(-) create mode 100644 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md diff --git a/02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md b/02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md new file mode 100644 index 0000000..e0cf01d --- /dev/null +++ b/02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md @@ -0,0 +1,134 @@ +--- +topic: the mesh +status: accepted +date: 2026-10-06 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0040-what-a-module-is.md +--- + +# 224. A provider that keeps failing a consumer is a problem the controller reports + +## Context + +**A provider's provisioner is the only thing in the mesh that knows whether a provision was made.** +The controller composes a grant, delivers the contributions file and the minted secret, and the node +reports that it applied every resource. Whether the provider then made the consumer's database, client +or bucket is known to the provisioner loop alone ([ADR 0040](0040-what-a-module-is.md), +[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)), and until now it +said so only in its own journal. + +**On 2026-10-05 that cost a day.** The identity provider's database was moved that morning. Its +realm's administrator kept the password it had before, the module's minted one was inert, and its +provisioner failed every consumer every five seconds — about 31,000 refused logins from shortly after +midnight until it was repaired by hand that night. Every surface the mesh has said the mesh was well: +the machines applied what they were sent, the modules were current, `status` printed its all-well +sentence. The same fault had been found and repaired by hand four days earlier +([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)). + +[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md) +already taught that "the mesh and the machines agree" is not "it works", and `status` says so under its +all-well line. This is the narrower, checkable half of that gap: not whether a consumer can reach its +provider, which only dialling answers, but whether the provider has told the mesh it cannot do its job. + +## Considered Options + +1. **Leave it to the journal, and to a person reading it.** Rejected: that is what happened, twice. +2. **The controller dials every provision.** Rejected for now: it is issue 145's open question, it + needs the controller to hold or borrow every consumer's credential, and it would still not say + *why* a provider fails. +3. **The provider reports its standing through the node's report.** Rejected: the node-engine applies + resources and knows nothing of what a module's code does after it starts; the provisioner runs in the + node's tool runtime, which speaks to the bus, not to the host. +4. **The provider announces, as an event, a consumer it keeps failing; the controller keeps the newest + word and `status` names it.** Chosen. + +## Decision + +**1. A provider announces a consumer it keeps failing.** When a provisioner has failed one consumer — +its create, its periodic check, or reading the consumer's minted secret — for five minutes without a +single success in between, it emits `provisioner.failing`, naming the consumer, the consumer's +machine, the provision, the class of error and the error's first words, since when, and how many +attempts. It says it again every fifteen minutes while it lasts. The first success after that is +`provisioner.recovered`; so is a consumer the mesh stopped asking for, and so is the first success +for each consumer after the provider starts, because a provider restarted after announcing a failure +has forgotten it. + +- **The classes** are `credentials-rejected`, `unreachable`, `secret-unreadable` and `refused` (any + other refusal), read from the error's words unless the provider's own code classes it better. They are + what a person reading `status` needs before opening the journal. +- **No secret travels.** The error is the provider's own text with the consumer's password removed, as + its log line already is. + +**2. Every provider may say it, whatever its manifest lists.** The permission to publish the two events +is derived for every module that receives contributions; no manifest declares them. A provider whose +manifest forgot them would otherwise have its announcement refused by the bus and fail as silently as +before. + +**3. The controller follows both events from every module, keeps the newest failing word per provider +module, its machine and the consumer, and removes it on recovery.** Who said it is read from the subject +the bus let the provider publish on, never from the body. A recovery that arrives while the store is +away is held and delivered again, because it is said once. The controller's subscription names the two +events with a wildcard for the emitter — the one pattern on its list — rather than a list of providers +somebody would have to extend. + +**4. `status` names every consumer a provider still assigned where it ran says it keeps failing**, and +such a consumer breaks the all-well sentence. Its JSON carries them as `failing`; `node show` lists those +whose provider or consumer is on that machine. A provider no longer assigned there has nothing running +to fail anybody, and is not asked about. A standing not said again for thirty minutes is shown with how +long it has been silent: the provider stopped saying anything, and its last word is all the mesh has. + +**5. Where a provider's failure has one known cause it can repair safely, it repairs it.** The identity +provider's administrator refusing the mesh's secret is the first: the module now checks the +administrator's login on start, every five minutes and whenever its provisioner is refused, and repairs +a refusal through the server's own bootstrap command — a temporary administrator, the real one's +password set to the mesh's, the temporary one removed, nothing printed — then checks again and says +what it did. A repair that fails is braked, from ten minutes doubling to six hours, and announced; while +the administrator is refused the provisioner stops asking the server, so a lockout policy is not +provoked, and its consumers are still announced failing. The mechanism is the module's +([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)), +the rule is this record's: **detected automatically, repaired where safe, loud where not.** + +## Consequences + +- **The loop is carried by every Go provider until the Go SDK has one.** The provisioner loop is the + SDK's ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)); the Go SDK has none yet, so the two Go + providers carry identical copies and a test in each fails when they differ. A TypeScript provider + announces nothing until the TypeScript SDK's loop does the same; until then its consumers fail as + silently as before, and the next provider to be ported to Go closes that gap for itself. +- **The controller's event consumer takes one message at a time**, so a standing can wait behind a + build being acted on for some minutes. Fifteen-minute repetition makes that harmless for a failure; + a recovery is held, never dropped. +- **The store keeps one row per provider module, its machine and consumer**, through a numbered + migration. +- **The rollout order matters.** The controller that derives the permission and follows the events + first; the catalogue's providers second — an older controller refuses nothing that matters, but + every announcement is then refused by the bus and logged by the provider. + +## How it is checked + +| Rule | Checked by | +|---|---| +| A consumer failed for five minutes without a success is announced, again every fifteen, and recovered on the first success | provider loop tests in mesh-catalog, postgres and keycloak (`standing_test.go`) | +| A failing periodic check and an unreadable secret count, not only a failing create | the same tests | +| A withdrawn consumer and the first success after a start are announced recovered | the same tests | +| The two Go providers carry the same loop | `harness_same_test.go` in each, comparing the files | +| Every module that receives contributions is granted the two events, and no other module is | mesh-controller inventory test over the derived declaration and the bus permissions | +| The controller may hear the two events from any module, and not every event | the same test over the controller's permissions | +| The emitter is read from the subject; a failing word is kept, a recovery cleared, and a recovery held while the store is away | mesh-controller link tests with a fake store | +| One row per provider, machine and consumer; recovered removes only its own | an inventory test against a live database, through the migration | +| `status` names it and is not all-well; the JSON carries it; `node show` names it on both machines; an unassigned provider's word is not a problem | mesh-controller command tests against a live database | +| The identity provider repairs a refused administrator, verifies, brakes a failed repair, never puts a secret in a command line, and stops asking the server while refused | keycloak module tests with a fake server and a fake container runtime; a live test against a throwaway server when asked for | +| Live: after rollout, `status` shows nothing failing on a healthy mesh; a provider made to fail a consumer for five minutes appears, and disappears on its next success | read through the console after rollout | + +## References + +- [Issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md) + — the failure, twice, and the repair this record makes automatic. +- [Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md) + — the scope of the all-well sentence, which this narrows and does not close. +- [ADR 0040](0040-what-a-module-is.md), [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), + [ADR 0039](0039-what-the-sdk-holds-and-refuses.md), [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md). +- [The module protocol](../03-DESIGN/01-to-be/19-the-module-protocol.md), provisioning, amended alongside. +- mesh-controller `internal/link/standing.go`, `internal/inventory/standing.go` and its migration, + `cmd/mesh-controller/standing.go`; mesh-catalog `modules/postgres` and `modules/keycloak`. diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 99a269d..3a52a78 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -198,6 +198,7 @@ python3 00-META/checks/index.py fail if stale - **0219** — [The build queue is controlled through the controller and the build seat](0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md) - **0221** — [A push sends no build a policy or a plan holds back, except to the machine it names](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md) - **0222** — [A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md) +- **0224** — [A provider that keeps failing a consumer is a problem the controller reports](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md) ### Its tiers, from the bottom up diff --git a/03-DESIGN/01-to-be/19-the-module-protocol.md b/03-DESIGN/01-to-be/19-the-module-protocol.md index 33f9a45..44b3337 100644 --- a/03-DESIGN/01-to-be/19-the-module-protocol.md +++ b/03-DESIGN/01-to-be/19-the-module-protocol.md @@ -5,8 +5,11 @@ code: - mesh-sdk src - mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout) - mesh-controller internal/link -updated: 2026-09-26 + - mesh-controller cmd/mesh-controller (status) + - mesh-catalog modules/postgres, modules/keycloak (the Go provisioner loop) +updated: 2026-10-06 decisions: + - 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md - 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md - 02-DECISIONS/0106-the-bus-is-nats.md - 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md @@ -232,6 +235,17 @@ A provider ships the provisioner that creates instances of what it offers same provision, so a grant addressed to a node alone does not name a consumer, and withdrawing one would take another's away. +### A provider says which consumer it keeps failing + +Whether a provision was made is known to the provider's loop alone. A consumer the loop has failed for +five minutes without one success — its create, its periodic check, or reading the secret the mesh +minted for it — is announced as `provisioner.failing`, naming the consumer, its machine, the class of +error and since when, and again every fifteen minutes while it lasts; the first success, a withdrawal, +and the first success after the provider restarts are `provisioner.recovered`. Every module that +receives contributions may publish both, derived and never declared. The controller keeps the newest +failing word per provider, machine and consumer, and `status` names each one until it recovers +([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)). + ### Checked, and it agrees (2026-09-16) This looked like the sharpest disagreement and was not one. The live wire is the contributions file @@ -255,4 +269,5 @@ two implementations drift from while both believe they conform. | The envelope is the envelope | An emitted event is compared header by header against a fixture; a missing required header fails, an unknown `x-` header is accepted. | | Delivery is at-least-once | A fixture delivered twice is handled once. | | A grant names a consumer | A grant fixture is read by every implementation and yields the same module and the same node. | +| A provider that keeps failing a consumer says so | The loop's tests drive five minutes of failure to one `provisioner.failing` and a success to `provisioner.recovered`; the controller's tests carry it from the bus to `status`. | | A partial SDK is legitimate | An implementation claiming the floor and events passes those suites and is listed for them; a module using tools in that language is refused at build time with the reason. | diff --git a/04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md b/04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md index bdcb63b..8894fce 100644 --- a/04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md +++ b/04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md @@ -1,9 +1,9 @@ --- -status: resolved +status: located opened: 2026-10-01 -located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm)] -fixed-by: done by hand on 2026-10-01 through the server's own bootstrap command — the admin's password set to the value the mesh minted; no code changed -amended-design: +located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh, and the same database moved on 2026-10-05), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm), every provider's provisioner loop (a consumer failed for a day said so only in a journal), mesh-controller status (nothing read what a provider could not do)] +fixed-by: twice by hand through the server's own bootstrap command (2026-10-01, 2026-10-05); the safety nets in mesh-controller PR #70, mesh-host PR #28 and mesh-catalog PR #80 (ADR 0224), resolved when they are merged and rolled out +amended-design: 03-DESIGN/01-to-be/19-the-module-protocol.md --- # 179 — An adopted identity provider's admin never took the secret the mesh minted @@ -44,6 +44,60 @@ only the download client's: a credential the mesh cannot make is the operator's *How it was checked:* `keycloak_list_realms` through the console answers; `keycloak_list_clients` on the realm lists both mesh-named clients; the server's log stops the five-second login error. +## Recurred, 2026-10-05 + +**The identity provider's database was moved that morning, and the admin's old password came back +with it.** The record the 2026-10-01 repair wrote lived in the database; the database that came up +on the new store was taken from before that repair, so the realm's `admin` again held a password that +predates the mesh, and the server — which applies the minted variable only when it creates its master +realm — did not touch it. From shortly after midnight the provisioner failed every consumer every five +seconds with *401 invalid_grant, Invalid user credentials*: about 31,000 refused logins until it was +repaired by hand, the same way as the first time, at about 23:55. In those twenty-three hours nothing +anywhere said so except the provider's own journal. The machines applied what they were sent, the +module was current, and `status` printed its all-well sentence. + +**The repair, as done both times**, inside the server's container and with nothing printed: the +server's `bootstrap-admin user` command creates a temporary administrator, from a password generated +in the container and handed over in an environment variable, on a management port other than the +default (the running server holds that one); the admin client logs in as it and sets `admin`'s password +to the value the mesh minted; the temporary administrator and every temporary file are removed. + +## What makes it not happen silently again (ADR 0224) + +Twice by hand is a pattern, and the operator's direction was that it never happen again: detected +automatically, repaired automatically where safe, loud where not, and tested. +[ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md) +records the rule; three safety nets implement it. + +1. **The identity provider repairs its own admin.** Its code — ported from TypeScript to Go with this + change — checks that `admin` logs in with the mesh's secret once the server answers, every five + minutes after, and at once whenever the provisioner is refused. A refusal is repaired exactly as + above, by the module, inside the server's container through the container runtime it already + declares, with the secrets passed on standard input and never on a command line; then the login is + checked again and one line and an `admin.repaired` event say what was done and why. A repair that + fails is announced as `admin.unrepaired`, said loudly in the journal with a pointer here, and braked — + ten minutes, doubling to six hours — and while the admin is refused the provisioner stops asking the + server, so a lockout policy is never provoked. `keycloak_admin_check` reports the admin's state and, + with `repair`, repairs it now. +2. **A provider that keeps failing a consumer says so.** The provisioner loop announces a consumer it + has failed for five minutes without a success — create, the periodic check or an unreadable secret — + as `provisioner.failing`, with the class of error, and again every fifteen minutes; its recovery as + `provisioner.recovered`. +3. **`status` names it.** The controller keeps each provider's newest failing word per consumer, and + `status`, its JSON and `node show` list it; it breaks the all-well sentence. On 2026-10-05 the first + line of `status` would have been the identity provider failing both consumers with + `credentials-rejected`, five minutes after midnight. + +The first is specific to this module and is the repair; the second and third are for every provider, +and are what would have said it on the day even if the repair had not existed. + +*How it is checked:* the module's tests drive the repair against a fake server and a fake container +runtime — repaired and verified on a refusal, never while the server is unreachable, braked after a +failed repair, no secret on any command line, the provisioner short-circuited while refused; the loop's +tests drive the announcements; the controller's tests take a failing word from the bus to `status` and +`node show` and back out on recovery. Live, after rollout: a deliberately wrong admin password on a +throwaway server is repaired within five minutes, and `status` stays all-well while it is. + ## Open The manifest still says the admin password is the mesh's to mint, which is true of a fresh install