audit-logger, model-usage: retry a failed write and never lose the event (hq issue 276)
Both caught a failed write and took the event, losing it silently; the SDK's rule is to throw when the work was not done. A failed write now throws so the bus offers the event again, and is spooled on disk at once; on its last delivery the spooled event is taken, and a background pass replays the spool once writing works. The runtime does not pass the delivery count, so the spool counts failed deliveries itself, across restarts. Over its bound (1000 events or 30 minutes) the last delivery is no longer taken, so the bus gives it up and the controller raises max-deliveries - the one existing condition that names a consumer which cannot keep up - while the spool still holds it. Writes are idempotent by event: the trail skips an id it already wrote; the usage upsert keeps the reading observed latest (migration 2), so a late replay never overwrites a newer one. Each module has a status tool for the spool, declared as valuable data (ADR 0233). model-usage moves to the bundle shape (ADR 0198) with its schema in a prepare step and numbered migrations; its old container shape had no image. Both on mesh-sdk 0.1.13. The log-only handlers of redis, mssql, mosquitto, mongodb, mesh-vault, showcase and the catalogue no longer throw a TypeError on an event without a body.
This commit is contained in:
@@ -1,38 +1,32 @@
|
||||
// model-usage's entrypoint — the usage context store's consumer (novox/hq ADR 0054). mesh-controller is
|
||||
// a CLI and cannot consume events, so the store that keeps the latest usage reading is a MODULE: it
|
||||
// subscribes to `module.*.usage.*` and upserts each row. Like the audit-logger, the on(...) IS the
|
||||
// whole handshake — the runtime imports this once the broker is bound, and every usage event any
|
||||
// producer emits lands here as well as on the audit trail.
|
||||
// subscribes to `*.usage.*` and upserts each row. Like the audit-logger, the on(...) IS the whole
|
||||
// handshake — the node's runtime launches this once the broker is bound (ADR 0198), and every usage
|
||||
// event any producer emits lands here as well as on the audit trail.
|
||||
//
|
||||
// The producers (the anthropic adapters) emit ALREADY-NORMALISED rows: the vendor→row normalisation
|
||||
// lives in the adapter, not here, so this consumer is vendor-neutral (ADR 0054). A body is
|
||||
// `{ rows: UsageRow[], raw }`; each row is upserted, carrying its own `raw` or the body's as a
|
||||
// fallback. Delivery is at-least-once, so a duplicate is fine — the upsert keeps the latest.
|
||||
// **No event is lost** (novox/hq issue 276). A row the store did not take is thrown, so the bus offers
|
||||
// the event again; the first failure spools it on disk, its last delivery is taken from the spool, and
|
||||
// the spool is replayed into the store once it answers again (spool.ts). The upsert keeps the reading
|
||||
// observed latest, so a redelivery or a late replay never writes an older one over a newer.
|
||||
//
|
||||
// The schema is not brought up here: the prepare step does it before this version starts
|
||||
// (prepare/index.ts; ADR 0135), as for the module catalogue.
|
||||
|
||||
import { on } from "@novox/mesh-sdk/events";
|
||||
import { UsageStore, type UsageRow } from "./store.js";
|
||||
import { UsageStore } from "./store.js";
|
||||
import { usageSpoolPath, writerFor } from "./consume.js";
|
||||
import { replayEvery, Spool, takeOrSpool } from "./spool.js";
|
||||
|
||||
const say = (line: string) => console.error(`[model-usage] ${line}`);
|
||||
const store = UsageStore.fromEnv();
|
||||
const spool = await Spool.open(usageSpoolPath());
|
||||
const write = writerFor(store, say);
|
||||
|
||||
// Create the store's one table before subscribing. The DDL is idempotent (CREATE TABLE IF NOT
|
||||
// EXISTS), so a restart re-runs it harmlessly. This is done here, in the long-lived consumer, rather
|
||||
// than as a gating run-once step: the consumer is a `--restart unless-stopped` service, so if the
|
||||
// provider is not yet reachable — its overlay address comes up as the same push settles — this exits
|
||||
// and is restarted until it can connect, without ever halting the apply. A run-once migrate that had
|
||||
// to reach the provider over the overlay would block the very apply that brings the overlay up.
|
||||
await store.migrate();
|
||||
replayEvery(spool, write, 30_000, say);
|
||||
|
||||
await on("*.usage.*", async (event) => {
|
||||
const body = event.body as { rows?: UsageRow[]; raw?: unknown };
|
||||
for (const row of body.rows ?? []) {
|
||||
try {
|
||||
await store.upsert({ ...row, raw: row.raw ?? body.raw ?? {} });
|
||||
} catch (err) {
|
||||
// A failed upsert is a loud line, never a throw back into the broker that would wedge the
|
||||
// subscription (the audit-logger's discipline).
|
||||
console.error(`[model-usage] could not upsert a row from ${event.type}: ${err}`);
|
||||
}
|
||||
}
|
||||
});
|
||||
await on("*.usage.*", (event) => takeOrSpool(event, write, spool, say));
|
||||
|
||||
console.log("[model-usage] recording model usage to its store");
|
||||
console.log(
|
||||
"[model-usage] recording model usage to its store" +
|
||||
(spool.size > 0 ? `; ${spool.size} spooled event(s) wait in ${spool.dir}` : ""),
|
||||
);
|
||||
|
||||
Reference in New Issue
Block a user