Merge pull request 'Issue 276: the audit logger and the usage store retry, and lose no event' (#145) from issues/276-the-audit-logger-and-usage-store-retry into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on
This commit was merged in pull request #145.
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-media-catalog modules/plex, mesh-tools node-tools/internal/launch, mesh-sdk src/stdio, mesh-sdk go]
|
||||
located-in: [mesh-media-catalog modules/plex, mesh-tools node-tools/internal/launch, mesh-sdk src/stdio, mesh-sdk go, mesh-catalog modules/audit-logger, mesh-catalog modules/model-usage]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
@@ -72,3 +72,39 @@ by the SDKs' stdio tests and the runtime's launch test.
|
||||
No other event consumer in the catalogues had a success path that throws. Two consumers do the
|
||||
opposite — catch a failed write and take the event anyway — which the rule now names as wrong; they are
|
||||
listed in the diagnosis and left for their own change.
|
||||
|
||||
## The two consumers that took a failed write
|
||||
|
||||
**Decided by the operator, 2026-10-06: they retry, and never lose an event.** The audit logger and the
|
||||
usage store each caught a failed write and took the event, so a store that was away lost every event
|
||||
sent while it was. Under the rule above that is wrong; making them throw alone would trade a silent loss
|
||||
for a loud one, because after the consumer's maximum deliveries the bus gives the event up. So:
|
||||
|
||||
- **A failed write is thrown**, and the event is offered again.
|
||||
- **The first failure puts the event in a local spool**, written to disk before the handler throws.
|
||||
The spool is the module's own data, declared `valuable` ([ADR 0233](../../02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)).
|
||||
- **On its last delivery the event is taken**: the spool holds it. The runtime does not tell a handler
|
||||
which delivery it is, so the spool counts the failed deliveries itself, and keeps the count across a
|
||||
restart; the maximum is the one the controller gives every module's consumer.
|
||||
- **A background pass replays the spool** once writing works again. A write that succeeds — redelivered
|
||||
or replayed — removes the spooled copy.
|
||||
- **Writing twice writes once.** The trail skips an event id it already holds; the usage store keeps
|
||||
the reading observed latest, by the event's emit time, so a late replay never overwrites a newer one.
|
||||
- **Over a bound** — a thousand events waiting, or one waiting half an hour — every event is still
|
||||
spooled, but the last delivery is no longer taken: the bus gives it up and the controller raises
|
||||
`max-deliveries` for that module's consumer, while the spool still writes it when it can. A module has
|
||||
no standing of its own to report today ([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)'s
|
||||
is a provider's, and permitted only to providers), so the bound borrows the one existing condition
|
||||
that names a consumer which cannot keep up, rather than inventing a second. Each module's status
|
||||
tool says what is spooled, held, the oldest and the last error.
|
||||
|
||||
What this does not cover: a write that hangs past the runtime's event timeout is not a failure the
|
||||
handler sees, so it is counted by the bus and not the spool (the usage store's connection and queries
|
||||
are bounded well inside that timeout for this reason); and a spool that cannot itself be written — the
|
||||
disk the trail is on is full — leaves only the bus's deliveries. Checked by each module's tests: a
|
||||
failed write throws; a redelivery writes once; the last delivery is spooled and taken; the spool
|
||||
replays; over the bound the last delivery is not taken; an older reading replayed late does not
|
||||
overwrite a newer, also against a real database.
|
||||
|
||||
Found on the way: the log-only handlers threw a type error on an event with no body. They read it
|
||||
safely now.
|
||||
|
||||
@@ -34,3 +34,10 @@
|
||||
event ids; it marks an id as done before the work, and forgets it only when a refresh fails, not when
|
||||
listing the libraries does — an offer after that failure would be taken without a rescan. Noted to its
|
||||
author rather than changed here.
|
||||
8. **The two consumers that took a failed write** (decided in the report). Whether a handler could tell
|
||||
its last delivery: the runtime hands the bundle only the envelope — the bus's delivery count is read
|
||||
by nobody — and the SDKs pass nothing more. A module that must not lose an event therefore counts its
|
||||
own failed deliveries, and the count lives with the spooled event so a restart does not reset it.
|
||||
Whether the controller has a place for a module to say it is falling behind: the conditions it raises
|
||||
are its own observations and the bus's advisories; the one a module emits, a provider's standing, is
|
||||
permitted only to providers. `max-deliveries` is the existing word for a consumer that cannot keep up.
|
||||
|
||||
Reference in New Issue
Block a user