Merge pull request 'Issue 179 recurred; ADR 0224: a provider that keeps failing a consumer is a problem the controller reports' (#119) from issues/179-recurred-and-safety-nets into main

This commit was merged in pull request #119.
This commit is contained in:
2026-10-05 22:43:01 +00:00
4 changed files with 209 additions and 5 deletions
@@ -0,0 +1,134 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0040-what-a-module-is.md
---
# 224. A provider that keeps failing a consumer is a problem the controller reports
## Context
**A provider's provisioner is the only thing in the mesh that knows whether a provision was made.**
The controller composes a grant, delivers the contributions file and the minted secret, and the node
reports that it applied every resource. Whether the provider then made the consumer's database, client
or bucket is known to the provisioner loop alone ([ADR 0040](0040-what-a-module-is.md),
[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)), and until now it
said so only in its own journal.
**On 2026-10-05 that cost a day.** The identity provider's database was moved that morning. Its
realm's administrator kept the password it had before, the module's minted one was inert, and its
provisioner failed every consumer every five seconds — about 31,000 refused logins from shortly after
midnight until it was repaired by hand that night. Every surface the mesh has said the mesh was well:
the machines applied what they were sent, the modules were current, `status` printed its all-well
sentence. The same fault had been found and repaired by hand four days earlier
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
already taught that "the mesh and the machines agree" is not "it works", and `status` says so under its
all-well line. This is the narrower, checkable half of that gap: not whether a consumer can reach its
provider, which only dialling answers, but whether the provider has told the mesh it cannot do its job.
## Considered Options
1. **Leave it to the journal, and to a person reading it.** Rejected: that is what happened, twice.
2. **The controller dials every provision.** Rejected for now: it is issue 145's open question, it
needs the controller to hold or borrow every consumer's credential, and it would still not say
*why* a provider fails.
3. **The provider reports its standing through the node's report.** Rejected: the node-engine applies
resources and knows nothing of what a module's code does after it starts; the provisioner runs in the
node's tool runtime, which speaks to the bus, not to the host.
4. **The provider announces, as an event, a consumer it keeps failing; the controller keeps the newest
word and `status` names it.** Chosen.
## Decision
**1. A provider announces a consumer it keeps failing.** When a provisioner has failed one consumer —
its create, its periodic check, or reading the consumer's minted secret — for five minutes without a
single success in between, it emits `provisioner.failing`, naming the consumer, the consumer's
machine, the provision, the class of error and the error's first words, since when, and how many
attempts. It says it again every fifteen minutes while it lasts. The first success after that is
`provisioner.recovered`; so is a consumer the mesh stopped asking for, and so is the first success
for each consumer after the provider starts, because a provider restarted after announcing a failure
has forgotten it.
- **The classes** are `credentials-rejected`, `unreachable`, `secret-unreadable` and `refused` (any
other refusal), read from the error's words unless the provider's own code classes it better. They are
what a person reading `status` needs before opening the journal.
- **No secret travels.** The error is the provider's own text with the consumer's password removed, as
its log line already is.
**2. Every provider may say it, whatever its manifest lists.** The permission to publish the two events
is derived for every module that receives contributions; no manifest declares them. A provider whose
manifest forgot them would otherwise have its announcement refused by the bus and fail as silently as
before.
**3. The controller follows both events from every module, keeps the newest failing word per provider
module, its machine and the consumer, and removes it on recovery.** Who said it is read from the subject
the bus let the provider publish on, never from the body. A recovery that arrives while the store is
away is held and delivered again, because it is said once. The controller's subscription names the two
events with a wildcard for the emitter — the one pattern on its list — rather than a list of providers
somebody would have to extend.
**4. `status` names every consumer a provider still assigned where it ran says it keeps failing**, and
such a consumer breaks the all-well sentence. Its JSON carries them as `failing`; `node show` lists those
whose provider or consumer is on that machine. A provider no longer assigned there has nothing running
to fail anybody, and is not asked about. A standing not said again for thirty minutes is shown with how
long it has been silent: the provider stopped saying anything, and its last word is all the mesh has.
**5. Where a provider's failure has one known cause it can repair safely, it repairs it.** The identity
provider's administrator refusing the mesh's secret is the first: the module now checks the
administrator's login on start, every five minutes and whenever its provisioner is refused, and repairs
a refusal through the server's own bootstrap command — a temporary administrator, the real one's
password set to the mesh's, the temporary one removed, nothing printed — then checks again and says
what it did. A repair that fails is braked, from ten minutes doubling to six hours, and announced; while
the administrator is refused the provisioner stops asking the server, so a lockout policy is not
provoked, and its consumers are still announced failing. The mechanism is the module's
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)),
the rule is this record's: **detected automatically, repaired where safe, loud where not.**
## Consequences
- **The loop is carried by every Go provider until the Go SDK has one.** The provisioner loop is the
SDK's ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)); the Go SDK has none yet, so the two Go
providers carry identical copies and a test in each fails when they differ. A TypeScript provider
announces nothing until the TypeScript SDK's loop does the same; until then its consumers fail as
silently as before, and the next provider to be ported to Go closes that gap for itself.
- **The controller's event consumer takes one message at a time**, so a standing can wait behind a
build being acted on for some minutes. Fifteen-minute repetition makes that harmless for a failure;
a recovery is held, never dropped.
- **The store keeps one row per provider module, its machine and consumer**, through a numbered
migration.
- **The rollout order matters.** The controller that derives the permission and follows the events
first; the catalogue's providers second — an older controller refuses nothing that matters, but
every announcement is then refused by the bus and logged by the provider.
## How it is checked
| Rule | Checked by |
|---|---|
| A consumer failed for five minutes without a success is announced, again every fifteen, and recovered on the first success | provider loop tests in mesh-catalog, postgres and keycloak (`standing_test.go`) |
| A failing periodic check and an unreadable secret count, not only a failing create | the same tests |
| A withdrawn consumer and the first success after a start are announced recovered | the same tests |
| The two Go providers carry the same loop | `harness_same_test.go` in each, comparing the files |
| Every module that receives contributions is granted the two events, and no other module is | mesh-controller inventory test over the derived declaration and the bus permissions |
| The controller may hear the two events from any module, and not every event | the same test over the controller's permissions |
| The emitter is read from the subject; a failing word is kept, a recovery cleared, and a recovery held while the store is away | mesh-controller link tests with a fake store |
| One row per provider, machine and consumer; recovered removes only its own | an inventory test against a live database, through the migration |
| `status` names it and is not all-well; the JSON carries it; `node show` names it on both machines; an unassigned provider's word is not a problem | mesh-controller command tests against a live database |
| The identity provider repairs a refused administrator, verifies, brakes a failed repair, never puts a secret in a command line, and stops asking the server while refused | keycloak module tests with a fake server and a fake container runtime; a live test against a throwaway server when asked for |
| Live: after rollout, `status` shows nothing failing on a healthy mesh; a provider made to fail a consumer for five minutes appears, and disappears on its next success | read through the console after rollout |
## References
- [Issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)
— the failure, twice, and the repair this record makes automatic.
- [Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the scope of the all-well sentence, which this narrows and does not close.
- [ADR 0040](0040-what-a-module-is.md), [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md),
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md), [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md).
- [The module protocol](../03-DESIGN/01-to-be/19-the-module-protocol.md), provisioning, amended alongside.
- mesh-controller `internal/link/standing.go`, `internal/inventory/standing.go` and its migration,
`cmd/mesh-controller/standing.go`; mesh-catalog `modules/postgres` and `modules/keycloak`.
+1
View File
@@ -198,6 +198,7 @@ python3 00-META/checks/index.py fail if stale
- **0219** — [The build queue is controlled through the controller and the build seat](0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md)
- **0221** — [A push sends no build a policy or a plan holds back, except to the machine it names](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)
- **0222** — [A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md)
- **0224** — [A provider that keeps failing a consumer is a problem the controller reports](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
### Its tiers, from the bottom up
+16 -1
View File
@@ -5,8 +5,11 @@ code:
- mesh-sdk src
- mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout)
- mesh-controller internal/link
updated: 2026-09-26
- mesh-controller cmd/mesh-controller (status)
- mesh-catalog modules/postgres, modules/keycloak (the Go provisioner loop)
updated: 2026-10-06
decisions:
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
@@ -232,6 +235,17 @@ A provider ships the provisioner that creates instances of what it offers
same provision, so a grant addressed to a node alone does not name a consumer, and withdrawing one
would take another's away.
### A provider says which consumer it keeps failing
Whether a provision was made is known to the provider's loop alone. A consumer the loop has failed for
five minutes without one success — its create, its periodic check, or reading the secret the mesh
minted for it — is announced as `provisioner.failing`, naming the consumer, its machine, the class of
error and since when, and again every fifteen minutes while it lasts; the first success, a withdrawal,
and the first success after the provider restarts are `provisioner.recovered`. Every module that
receives contributions may publish both, derived and never declared. The controller keeps the newest
failing word per provider, machine and consumer, and `status` names each one until it recovers
([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)).
### Checked, and it agrees (2026-09-16)
This looked like the sharpest disagreement and was not one. The live wire is the contributions file
@@ -255,4 +269,5 @@ two implementations drift from while both believe they conform.
| The envelope is the envelope | An emitted event is compared header by header against a fixture; a missing required header fails, an unknown `x-` header is accepted. |
| Delivery is at-least-once | A fixture delivered twice is handled once. |
| A grant names a consumer | A grant fixture is read by every implementation and yields the same module and the same node. |
| A provider that keeps failing a consumer says so | The loop's tests drive five minutes of failure to one `provisioner.failing` and a success to `provisioner.recovered`; the controller's tests carry it from the bus to `status`. |
| A partial SDK is legitimate | An implementation claiming the floor and events passes those suites and is listed for them; a module using tools in that language is refused at build time with the reason. |
@@ -1,9 +1,9 @@
---
status: resolved
status: located
opened: 2026-10-01
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm)]
fixed-by: done by hand on 2026-10-01 through the server's own bootstrap command — the admin's password set to the value the mesh minted; no code changed
amended-design:
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh, and the same database moved on 2026-10-05), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm), every provider's provisioner loop (a consumer failed for a day said so only in a journal), mesh-controller status (nothing read what a provider could not do)]
fixed-by: twice by hand through the server's own bootstrap command (2026-10-01, 2026-10-05); the safety nets in mesh-controller PR #70, mesh-host PR #28 and mesh-catalog PR #80 (ADR 0224), resolved when they are merged and rolled out
amended-design: 03-DESIGN/01-to-be/19-the-module-protocol.md
---
# 179 — An adopted identity provider's admin never took the secret the mesh minted
@@ -44,6 +44,60 @@ only the download client's: a credential the mesh cannot make is the operator's
*How it was checked:* `keycloak_list_realms` through the console answers; `keycloak_list_clients`
on the realm lists both mesh-named clients; the server's log stops the five-second login error.
## Recurred, 2026-10-05
**The identity provider's database was moved that morning, and the admin's old password came back
with it.** The record the 2026-10-01 repair wrote lived in the database; the database that came up
on the new store was taken from before that repair, so the realm's `admin` again held a password that
predates the mesh, and the server — which applies the minted variable only when it creates its master
realm — did not touch it. From shortly after midnight the provisioner failed every consumer every five
seconds with *401 invalid_grant, Invalid user credentials*: about 31,000 refused logins until it was
repaired by hand, the same way as the first time, at about 23:55. In those twenty-three hours nothing
anywhere said so except the provider's own journal. The machines applied what they were sent, the
module was current, and `status` printed its all-well sentence.
**The repair, as done both times**, inside the server's container and with nothing printed: the
server's `bootstrap-admin user` command creates a temporary administrator, from a password generated
in the container and handed over in an environment variable, on a management port other than the
default (the running server holds that one); the admin client logs in as it and sets `admin`'s password
to the value the mesh minted; the temporary administrator and every temporary file are removed.
## What makes it not happen silently again (ADR 0224)
Twice by hand is a pattern, and the operator's direction was that it never happen again: detected
automatically, repaired automatically where safe, loud where not, and tested.
[ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
records the rule; three safety nets implement it.
1. **The identity provider repairs its own admin.** Its code — ported from TypeScript to Go with this
change — checks that `admin` logs in with the mesh's secret once the server answers, every five
minutes after, and at once whenever the provisioner is refused. A refusal is repaired exactly as
above, by the module, inside the server's container through the container runtime it already
declares, with the secrets passed on standard input and never on a command line; then the login is
checked again and one line and an `admin.repaired` event say what was done and why. A repair that
fails is announced as `admin.unrepaired`, said loudly in the journal with a pointer here, and braked —
ten minutes, doubling to six hours — and while the admin is refused the provisioner stops asking the
server, so a lockout policy is never provoked. `keycloak_admin_check` reports the admin's state and,
with `repair`, repairs it now.
2. **A provider that keeps failing a consumer says so.** The provisioner loop announces a consumer it
has failed for five minutes without a success — create, the periodic check or an unreadable secret —
as `provisioner.failing`, with the class of error, and again every fifteen minutes; its recovery as
`provisioner.recovered`.
3. **`status` names it.** The controller keeps each provider's newest failing word per consumer, and
`status`, its JSON and `node show` list it; it breaks the all-well sentence. On 2026-10-05 the first
line of `status` would have been the identity provider failing both consumers with
`credentials-rejected`, five minutes after midnight.
The first is specific to this module and is the repair; the second and third are for every provider,
and are what would have said it on the day even if the repair had not existed.
*How it is checked:* the module's tests drive the repair against a fake server and a fake container
runtime — repaired and verified on a refusal, never while the server is unreachable, braked after a
failed repair, no secret on any command line, the provisioner short-circuited while refused; the loop's
tests drive the announcements; the controller's tests take a failing word from the bus to `status` and
`node show` and back out on recovery. Live, after rollout: a deliberately wrong admin password on a
throwaway server is repaired within five minutes, and `status` stays all-well while it is.
## Open
The manifest still says the admin password is the mesh's to mint, which is true of a fresh install