Issue 276: the audit logger and the usage store retry, and lose no event

The operator decided the two consumers that took a failed write must retry and never lose an event: record how (a thrown write, a spool that takes the last delivery and replays, idempotent writes, a bound that borrows max-deliveries), and why the module counts its own deliveries.
This commit is contained in:
jochen
2026-10-06 18:37:28 +02:00
parent 84707a783a
commit 72fcdc366f
2 changed files with 44 additions and 1 deletions
@@ -1,7 +1,7 @@
---
status: located
opened: 2026-10-06
located-in: [mesh-media-catalog modules/plex, mesh-tools node-tools/internal/launch, mesh-sdk src/stdio, mesh-sdk go]
located-in: [mesh-media-catalog modules/plex, mesh-tools node-tools/internal/launch, mesh-sdk src/stdio, mesh-sdk go, mesh-catalog modules/audit-logger, mesh-catalog modules/model-usage]
fixed-by:
amended-design:
---
@@ -72,3 +72,39 @@ by the SDKs' stdio tests and the runtime's launch test.
No other event consumer in the catalogues had a success path that throws. Two consumers do the
opposite — catch a failed write and take the event anyway — which the rule now names as wrong; they are
listed in the diagnosis and left for their own change.
## The two consumers that took a failed write
**Decided by the operator, 2026-10-06: they retry, and never lose an event.** The audit logger and the
usage store each caught a failed write and took the event, so a store that was away lost every event
sent while it was. Under the rule above that is wrong; making them throw alone would trade a silent loss
for a loud one, because after the consumer's maximum deliveries the bus gives the event up. So:
- **A failed write is thrown**, and the event is offered again.
- **The first failure puts the event in a local spool**, written to disk before the handler throws.
The spool is the module's own data, declared `valuable` ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)).
- **On its last delivery the event is taken**: the spool holds it. The runtime does not tell a handler
which delivery it is, so the spool counts the failed deliveries itself, and keeps the count across a
restart; the maximum is the one the controller gives every module's consumer.
- **A background pass replays the spool** once writing works again. A write that succeeds — redelivered
or replayed — removes the spooled copy.
- **Writing twice writes once.** The trail skips an event id it already holds; the usage store keeps
the reading observed latest, by the event's emit time, so a late replay never overwrites a newer one.
- **Over a bound** — a thousand events waiting, or one waiting half an hour — every event is still
spooled, but the last delivery is no longer taken: the bus gives it up and the controller raises
`max-deliveries` for that module's consumer, while the spool still writes it when it can. A module has
no standing of its own to report today ([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)'s
is a provider's, and permitted only to providers), so the bound borrows the one existing condition
that names a consumer which cannot keep up, rather than inventing a second. Each module's status
tool says what is spooled, held, the oldest and the last error.
What this does not cover: a write that hangs past the runtime's event timeout is not a failure the
handler sees, so it is counted by the bus and not the spool (the usage store's connection and queries
are bounded well inside that timeout for this reason); and a spool that cannot itself be written — the
disk the trail is on is full — leaves only the bus's deliveries. Checked by each module's tests: a
failed write throws; a redelivery writes once; the last delivery is spooled and taken; the spool
replays; over the bound the last delivery is not taken; an older reading replayed late does not
overwrite a newer, also against a real database.
Found on the way: the log-only handlers threw a type error on an event with no body. They read it
safely now.
@@ -34,3 +34,10 @@
event ids; it marks an id as done before the work, and forgets it only when a refresh fails, not when
listing the libraries does — an offer after that failure would be taken without a rescan. Noted to its
author rather than changed here.
8. **The two consumers that took a failed write** (decided in the report). Whether a handler could tell
its last delivery: the runtime hands the bundle only the envelope — the bus's delivery count is read
by nobody — and the SDKs pass nothing more. A module that must not lose an event therefore counts its
own failed deliveries, and the count lives with the spooled event so a restart does not reset it.
Whether the controller has a place for a module to say it is falling behind: the conditions it raises
are its own observations and the bus's advisories; the one a module emits, a provider's standing, is
permitted only to providers. `max-deliveries` is the existing word for a consumer that cannot keep up.