Issue 179 recurred; ADR 0224: a provider that keeps failing a consumer is a problem the controller reports
The identity provider's admin lost the mesh's password again when its database moved, and 31,000 silent failures followed. Record the recurrence, the rule that makes a failing provider visible in status, and the module's self-repair.
This commit is contained in:
+134
@@ -0,0 +1,134 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-06
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0040-what-a-module-is.md
|
||||
---
|
||||
|
||||
# 224. A provider that keeps failing a consumer is a problem the controller reports
|
||||
|
||||
## Context
|
||||
|
||||
**A provider's provisioner is the only thing in the mesh that knows whether a provision was made.**
|
||||
The controller composes a grant, delivers the contributions file and the minted secret, and the node
|
||||
reports that it applied every resource. Whether the provider then made the consumer's database, client
|
||||
or bucket is known to the provisioner loop alone ([ADR 0040](0040-what-a-module-is.md),
|
||||
[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)), and until now it
|
||||
said so only in its own journal.
|
||||
|
||||
**On 2026-10-05 that cost a day.** The identity provider's database was moved that morning. Its
|
||||
realm's administrator kept the password it had before, the module's minted one was inert, and its
|
||||
provisioner failed every consumer every five seconds — about 31,000 refused logins from shortly after
|
||||
midnight until it was repaired by hand that night. Every surface the mesh has said the mesh was well:
|
||||
the machines applied what they were sent, the modules were current, `status` printed its all-well
|
||||
sentence. The same fault had been found and repaired by hand four days earlier
|
||||
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
|
||||
|
||||
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
already taught that "the mesh and the machines agree" is not "it works", and `status` says so under its
|
||||
all-well line. This is the narrower, checkable half of that gap: not whether a consumer can reach its
|
||||
provider, which only dialling answers, but whether the provider has told the mesh it cannot do its job.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Leave it to the journal, and to a person reading it.** Rejected: that is what happened, twice.
|
||||
2. **The controller dials every provision.** Rejected for now: it is issue 145's open question, it
|
||||
needs the controller to hold or borrow every consumer's credential, and it would still not say
|
||||
*why* a provider fails.
|
||||
3. **The provider reports its standing through the node's report.** Rejected: the node-engine applies
|
||||
resources and knows nothing of what a module's code does after it starts; the provisioner runs in the
|
||||
node's tool runtime, which speaks to the bus, not to the host.
|
||||
4. **The provider announces, as an event, a consumer it keeps failing; the controller keeps the newest
|
||||
word and `status` names it.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. A provider announces a consumer it keeps failing.** When a provisioner has failed one consumer —
|
||||
its create, its periodic check, or reading the consumer's minted secret — for five minutes without a
|
||||
single success in between, it emits `provisioner.failing`, naming the consumer, the consumer's
|
||||
machine, the provision, the class of error and the error's first words, since when, and how many
|
||||
attempts. It says it again every fifteen minutes while it lasts. The first success after that is
|
||||
`provisioner.recovered`; so is a consumer the mesh stopped asking for, and so is the first success
|
||||
for each consumer after the provider starts, because a provider restarted after announcing a failure
|
||||
has forgotten it.
|
||||
|
||||
- **The classes** are `credentials-rejected`, `unreachable`, `secret-unreadable` and `refused` (any
|
||||
other refusal), read from the error's words unless the provider's own code classes it better. They are
|
||||
what a person reading `status` needs before opening the journal.
|
||||
- **No secret travels.** The error is the provider's own text with the consumer's password removed, as
|
||||
its log line already is.
|
||||
|
||||
**2. Every provider may say it, whatever its manifest lists.** The permission to publish the two events
|
||||
is derived for every module that receives contributions; no manifest declares them. A provider whose
|
||||
manifest forgot them would otherwise have its announcement refused by the bus and fail as silently as
|
||||
before.
|
||||
|
||||
**3. The controller follows both events from every module, keeps the newest failing word per provider
|
||||
module, its machine and the consumer, and removes it on recovery.** Who said it is read from the subject
|
||||
the bus let the provider publish on, never from the body. A recovery that arrives while the store is
|
||||
away is held and delivered again, because it is said once. The controller's subscription names the two
|
||||
events with a wildcard for the emitter — the one pattern on its list — rather than a list of providers
|
||||
somebody would have to extend.
|
||||
|
||||
**4. `status` names every consumer a provider still assigned where it ran says it keeps failing**, and
|
||||
such a consumer breaks the all-well sentence. Its JSON carries them as `failing`; `node show` lists those
|
||||
whose provider or consumer is on that machine. A provider no longer assigned there has nothing running
|
||||
to fail anybody, and is not asked about. A standing not said again for thirty minutes is shown with how
|
||||
long it has been silent: the provider stopped saying anything, and its last word is all the mesh has.
|
||||
|
||||
**5. Where a provider's failure has one known cause it can repair safely, it repairs it.** The identity
|
||||
provider's administrator refusing the mesh's secret is the first: the module now checks the
|
||||
administrator's login on start, every five minutes and whenever its provisioner is refused, and repairs
|
||||
a refusal through the server's own bootstrap command — a temporary administrator, the real one's
|
||||
password set to the mesh's, the temporary one removed, nothing printed — then checks again and says
|
||||
what it did. A repair that fails is braked, from ten minutes doubling to six hours, and announced; while
|
||||
the administrator is refused the provisioner stops asking the server, so a lockout policy is not
|
||||
provoked, and its consumers are still announced failing. The mechanism is the module's
|
||||
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)),
|
||||
the rule is this record's: **detected automatically, repaired where safe, loud where not.**
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The loop is carried by every Go provider until the Go SDK has one.** The provisioner loop is the
|
||||
SDK's ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)); the Go SDK has none yet, so the two Go
|
||||
providers carry identical copies and a test in each fails when they differ. A TypeScript provider
|
||||
announces nothing until the TypeScript SDK's loop does the same; until then its consumers fail as
|
||||
silently as before, and the next provider to be ported to Go closes that gap for itself.
|
||||
- **The controller's event consumer takes one message at a time**, so a standing can wait behind a
|
||||
build being acted on for some minutes. Fifteen-minute repetition makes that harmless for a failure;
|
||||
a recovery is held, never dropped.
|
||||
- **The store keeps one row per provider module, its machine and consumer**, through a numbered
|
||||
migration.
|
||||
- **The rollout order matters.** The controller that derives the permission and follows the events
|
||||
first; the catalogue's providers second — an older controller refuses nothing that matters, but
|
||||
every announcement is then refused by the bus and logged by the provider.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| A consumer failed for five minutes without a success is announced, again every fifteen, and recovered on the first success | provider loop tests in mesh-catalog, postgres and keycloak (`standing_test.go`) |
|
||||
| A failing periodic check and an unreadable secret count, not only a failing create | the same tests |
|
||||
| A withdrawn consumer and the first success after a start are announced recovered | the same tests |
|
||||
| The two Go providers carry the same loop | `harness_same_test.go` in each, comparing the files |
|
||||
| Every module that receives contributions is granted the two events, and no other module is | mesh-controller inventory test over the derived declaration and the bus permissions |
|
||||
| The controller may hear the two events from any module, and not every event | the same test over the controller's permissions |
|
||||
| The emitter is read from the subject; a failing word is kept, a recovery cleared, and a recovery held while the store is away | mesh-controller link tests with a fake store |
|
||||
| One row per provider, machine and consumer; recovered removes only its own | an inventory test against a live database, through the migration |
|
||||
| `status` names it and is not all-well; the JSON carries it; `node show` names it on both machines; an unassigned provider's word is not a problem | mesh-controller command tests against a live database |
|
||||
| The identity provider repairs a refused administrator, verifies, brakes a failed repair, never puts a secret in a command line, and stops asking the server while refused | keycloak module tests with a fake server and a fake container runtime; a live test against a throwaway server when asked for |
|
||||
| Live: after rollout, `status` shows nothing failing on a healthy mesh; a provider made to fail a consumer for five minutes appears, and disappears on its next success | read through the console after rollout |
|
||||
|
||||
## References
|
||||
|
||||
- [Issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)
|
||||
— the failure, twice, and the repair this record makes automatic.
|
||||
- [Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
— the scope of the all-well sentence, which this narrows and does not close.
|
||||
- [ADR 0040](0040-what-a-module-is.md), [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md),
|
||||
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md), [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md).
|
||||
- [The module protocol](../03-DESIGN/01-to-be/19-the-module-protocol.md), provisioning, amended alongside.
|
||||
- mesh-controller `internal/link/standing.go`, `internal/inventory/standing.go` and its migration,
|
||||
`cmd/mesh-controller/standing.go`; mesh-catalog `modules/postgres` and `modules/keycloak`.
|
||||
@@ -198,6 +198,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0219** — [The build queue is controlled through the controller and the build seat](0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md)
|
||||
- **0221** — [A push sends no build a policy or a plan holds back, except to the machine it names](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)
|
||||
- **0222** — [A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md)
|
||||
- **0224** — [A provider that keeps failing a consumer is a problem the controller reports](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
|
||||
|
||||
### Its tiers, from the bottom up
|
||||
|
||||
|
||||
@@ -5,8 +5,11 @@ code:
|
||||
- mesh-sdk src
|
||||
- mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout)
|
||||
- mesh-controller internal/link
|
||||
updated: 2026-09-26
|
||||
- mesh-controller cmd/mesh-controller (status)
|
||||
- mesh-catalog modules/postgres, modules/keycloak (the Go provisioner loop)
|
||||
updated: 2026-10-06
|
||||
decisions:
|
||||
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
|
||||
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
|
||||
- 02-DECISIONS/0106-the-bus-is-nats.md
|
||||
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
|
||||
@@ -232,6 +235,17 @@ A provider ships the provisioner that creates instances of what it offers
|
||||
same provision, so a grant addressed to a node alone does not name a consumer, and withdrawing one
|
||||
would take another's away.
|
||||
|
||||
### A provider says which consumer it keeps failing
|
||||
|
||||
Whether a provision was made is known to the provider's loop alone. A consumer the loop has failed for
|
||||
five minutes without one success — its create, its periodic check, or reading the secret the mesh
|
||||
minted for it — is announced as `provisioner.failing`, naming the consumer, its machine, the class of
|
||||
error and since when, and again every fifteen minutes while it lasts; the first success, a withdrawal,
|
||||
and the first success after the provider restarts are `provisioner.recovered`. Every module that
|
||||
receives contributions may publish both, derived and never declared. The controller keeps the newest
|
||||
failing word per provider, machine and consumer, and `status` names each one until it recovers
|
||||
([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)).
|
||||
|
||||
### Checked, and it agrees (2026-09-16)
|
||||
|
||||
This looked like the sharpest disagreement and was not one. The live wire is the contributions file
|
||||
@@ -255,4 +269,5 @@ two implementations drift from while both believe they conform.
|
||||
| The envelope is the envelope | An emitted event is compared header by header against a fixture; a missing required header fails, an unknown `x-` header is accepted. |
|
||||
| Delivery is at-least-once | A fixture delivered twice is handled once. |
|
||||
| A grant names a consumer | A grant fixture is read by every implementation and yields the same module and the same node. |
|
||||
| A provider that keeps failing a consumer says so | The loop's tests drive five minutes of failure to one `provisioner.failing` and a success to `provisioner.recovered`; the controller's tests carry it from the bus to `status`. |
|
||||
| A partial SDK is legitimate | An implementation claiming the floor and events passes those suites and is listed for them; a module using tools in that language is refused at build time with the reason. |
|
||||
|
||||
+58
-4
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-10-01
|
||||
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm)]
|
||||
fixed-by: done by hand on 2026-10-01 through the server's own bootstrap command — the admin's password set to the value the mesh minted; no code changed
|
||||
amended-design:
|
||||
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh, and the same database moved on 2026-10-05), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm), every provider's provisioner loop (a consumer failed for a day said so only in a journal), mesh-controller status (nothing read what a provider could not do)]
|
||||
fixed-by: twice by hand through the server's own bootstrap command (2026-10-01, 2026-10-05); the safety nets in mesh-controller PR #70, mesh-host PR #28 and mesh-catalog PR #80 (ADR 0224), resolved when they are merged and rolled out
|
||||
amended-design: 03-DESIGN/01-to-be/19-the-module-protocol.md
|
||||
---
|
||||
|
||||
# 179 — An adopted identity provider's admin never took the secret the mesh minted
|
||||
@@ -44,6 +44,60 @@ only the download client's: a credential the mesh cannot make is the operator's
|
||||
*How it was checked:* `keycloak_list_realms` through the console answers; `keycloak_list_clients`
|
||||
on the realm lists both mesh-named clients; the server's log stops the five-second login error.
|
||||
|
||||
## Recurred, 2026-10-05
|
||||
|
||||
**The identity provider's database was moved that morning, and the admin's old password came back
|
||||
with it.** The record the 2026-10-01 repair wrote lived in the database; the database that came up
|
||||
on the new store was taken from before that repair, so the realm's `admin` again held a password that
|
||||
predates the mesh, and the server — which applies the minted variable only when it creates its master
|
||||
realm — did not touch it. From shortly after midnight the provisioner failed every consumer every five
|
||||
seconds with *401 invalid_grant, Invalid user credentials*: about 31,000 refused logins until it was
|
||||
repaired by hand, the same way as the first time, at about 23:55. In those twenty-three hours nothing
|
||||
anywhere said so except the provider's own journal. The machines applied what they were sent, the
|
||||
module was current, and `status` printed its all-well sentence.
|
||||
|
||||
**The repair, as done both times**, inside the server's container and with nothing printed: the
|
||||
server's `bootstrap-admin user` command creates a temporary administrator, from a password generated
|
||||
in the container and handed over in an environment variable, on a management port other than the
|
||||
default (the running server holds that one); the admin client logs in as it and sets `admin`'s password
|
||||
to the value the mesh minted; the temporary administrator and every temporary file are removed.
|
||||
|
||||
## What makes it not happen silently again (ADR 0224)
|
||||
|
||||
Twice by hand is a pattern, and the operator's direction was that it never happen again: detected
|
||||
automatically, repaired automatically where safe, loud where not, and tested.
|
||||
[ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
|
||||
records the rule; three safety nets implement it.
|
||||
|
||||
1. **The identity provider repairs its own admin.** Its code — ported from TypeScript to Go with this
|
||||
change — checks that `admin` logs in with the mesh's secret once the server answers, every five
|
||||
minutes after, and at once whenever the provisioner is refused. A refusal is repaired exactly as
|
||||
above, by the module, inside the server's container through the container runtime it already
|
||||
declares, with the secrets passed on standard input and never on a command line; then the login is
|
||||
checked again and one line and an `admin.repaired` event say what was done and why. A repair that
|
||||
fails is announced as `admin.unrepaired`, said loudly in the journal with a pointer here, and braked —
|
||||
ten minutes, doubling to six hours — and while the admin is refused the provisioner stops asking the
|
||||
server, so a lockout policy is never provoked. `keycloak_admin_check` reports the admin's state and,
|
||||
with `repair`, repairs it now.
|
||||
2. **A provider that keeps failing a consumer says so.** The provisioner loop announces a consumer it
|
||||
has failed for five minutes without a success — create, the periodic check or an unreadable secret —
|
||||
as `provisioner.failing`, with the class of error, and again every fifteen minutes; its recovery as
|
||||
`provisioner.recovered`.
|
||||
3. **`status` names it.** The controller keeps each provider's newest failing word per consumer, and
|
||||
`status`, its JSON and `node show` list it; it breaks the all-well sentence. On 2026-10-05 the first
|
||||
line of `status` would have been the identity provider failing both consumers with
|
||||
`credentials-rejected`, five minutes after midnight.
|
||||
|
||||
The first is specific to this module and is the repair; the second and third are for every provider,
|
||||
and are what would have said it on the day even if the repair had not existed.
|
||||
|
||||
*How it is checked:* the module's tests drive the repair against a fake server and a fake container
|
||||
runtime — repaired and verified on a refusal, never while the server is unreachable, braked after a
|
||||
failed repair, no secret on any command line, the provisioner short-circuited while refused; the loop's
|
||||
tests drive the announcements; the controller's tests take a failing word from the bus to `status` and
|
||||
`node show` and back out on recovery. Live, after rollout: a deliberately wrong admin password on a
|
||||
throwaway server is repaired within five minutes, and `status` stays all-well while it is.
|
||||
|
||||
## Open
|
||||
|
||||
The manifest still says the admin password is the mesh's to mint, which is true of a fresh install
|
||||
|
||||
Reference in New Issue
Block a user