diff --git a/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md b/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md index 9334f76..183b8fd 100644 --- a/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md +++ b/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md @@ -1,7 +1,7 @@ --- status: located opened: 2026-10-06 -located-in: [mesh-media-catalog modules/plex, mesh-tools node-tools/internal/launch, mesh-sdk src/stdio, mesh-sdk go] +located-in: [mesh-media-catalog modules/plex, mesh-tools node-tools/internal/launch, mesh-sdk src/stdio, mesh-sdk go, mesh-catalog modules/audit-logger, mesh-catalog modules/model-usage] fixed-by: amended-design: --- @@ -72,3 +72,39 @@ by the SDKs' stdio tests and the runtime's launch test. No other event consumer in the catalogues had a success path that throws. Two consumers do the opposite — catch a failed write and take the event anyway — which the rule now names as wrong; they are listed in the diagnosis and left for their own change. + +## The two consumers that took a failed write + +**Decided by the operator, 2026-10-06: they retry, and never lose an event.** The audit logger and the +usage store each caught a failed write and took the event, so a store that was away lost every event +sent while it was. Under the rule above that is wrong; making them throw alone would trade a silent loss +for a loud one, because after the consumer's maximum deliveries the bus gives the event up. So: + +- **A failed write is thrown**, and the event is offered again. +- **The first failure puts the event in a local spool**, written to disk before the handler throws. + The spool is the module's own data, declared `valuable` ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)). +- **On its last delivery the event is taken**: the spool holds it. The runtime does not tell a handler + which delivery it is, so the spool counts the failed deliveries itself, and keeps the count across a + restart; the maximum is the one the controller gives every module's consumer. +- **A background pass replays the spool** once writing works again. A write that succeeds — redelivered + or replayed — removes the spooled copy. +- **Writing twice writes once.** The trail skips an event id it already holds; the usage store keeps + the reading observed latest, by the event's emit time, so a late replay never overwrites a newer one. +- **Over a bound** — a thousand events waiting, or one waiting half an hour — every event is still + spooled, but the last delivery is no longer taken: the bus gives it up and the controller raises + `max-deliveries` for that module's consumer, while the spool still writes it when it can. A module has + no standing of its own to report today ([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)'s + is a provider's, and permitted only to providers), so the bound borrows the one existing condition + that names a consumer which cannot keep up, rather than inventing a second. Each module's status + tool says what is spooled, held, the oldest and the last error. + +What this does not cover: a write that hangs past the runtime's event timeout is not a failure the +handler sees, so it is counted by the bus and not the spool (the usage store's connection and queries +are bounded well inside that timeout for this reason); and a spool that cannot itself be written — the +disk the trail is on is full — leaves only the bus's deliveries. Checked by each module's tests: a +failed write throws; a redelivery writes once; the last delivery is spooled and taken; the spool +replays; over the bound the last delivery is not taken; an older reading replayed late does not +overwrite a newer, also against a real database. + +Found on the way: the log-only handlers threw a type error on an event with no body. They read it +safely now. diff --git a/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/01-diagnosis.md b/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/01-diagnosis.md index 2f1db91..d8e6c46 100644 --- a/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/01-diagnosis.md +++ b/04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/01-diagnosis.md @@ -34,3 +34,10 @@ event ids; it marks an id as done before the work, and forgets it only when a refresh fails, not when listing the libraries does — an offer after that failure would be taken without a rescan. Noted to its author rather than changed here. +8. **The two consumers that took a failed write** (decided in the report). Whether a handler could tell + its last delivery: the runtime hands the bundle only the envelope — the bus's delivery count is read + by nobody — and the SDKs pass nothing more. A module that must not lose an event therefore counts its + own failed deliveries, and the count lives with the spooled event so a restart does not reset it. + Whether the controller has a place for a module to say it is falling behind: the conditions it raises + are its own observations and the bus's advisories; the one a module emits, a provider's standing, is + permitted only to providers. `max-deliveries` is the existing word for a consumer that cannot keep up.