From df2072aed58297113b13406beeac75464b273212 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 5 Oct 2026 17:16:35 +0200 Subject: [PATCH] Issues 248 and 249: the controller's event consumer replayed a week; a new state is refused until a push 248, located: a consumer made with the server's default replays the whole stream, and the controller's held every new merge and build behind a week of old ones. 249, open: a module's new state reaches its bundle before the grants that let it use it. --- .../00-report.md | 58 +++++++++++++++++++ .../01-diagnosis.md | 39 +++++++++++++ .../00-report.md | 37 ++++++++++++ 3 files changed, 134 insertions(+) create mode 100644 04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md create mode 100644 04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/01-diagnosis.md create mode 100644 04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md diff --git a/04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md b/04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md new file mode 100644 index 0000000..0f4d661 --- /dev/null +++ b/04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md @@ -0,0 +1,58 @@ +--- +status: located +opened: 2026-10-05 +located-in: [mesh-controller internal/broker] +fixed-by: +amended-design: +--- + +# 248 — The controller's event consumer replayed a week, and held every new merge behind it + +## What was observed + +A merge to the catalogue at 15:17 never reached the controller. No plan was made and no build was asked +for it. Every machine kept running the build from before it. The merge just before, to the record, was +logged twice. + +The bus showed why. The controller's durable consumer on the event stream was set to deliver +**everything the stream holds**, not from where it last was: + +| | | +|---|---| +| the stream | a week of events, from message 1 514 to message 356 004 | +| the consumer delivered to | message 1 517, later 4 898 | +| acknowledged to | 0, later 1 569 | +| still to deliver | 6 955 merges and build outcomes, a week old | +| allowed outstanding | one at a time ([issue 175](../175-an-announcement-behind-a-long-build-comes-back/00-report.md)) | + +So the controller was working through a week of past merges and build outcomes, one at a time, slowly. Every +new one — a merge, a build asked by hand — waited behind them. A build asked by hand finished on its +machine and was never registered. While this went on, the controller's client dropped messages +("slow consumer") several times, because heartbeats and reports share the loop with these events. From 15:17 +its event loop did nothing more: no line logged, the one delivered event never acknowledged. + +How the consumer came to deliver everything is not certain. The bus and the stores were rebuilt the night +before ([issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)). +A consumer that is missing is made again by the controller's own assertion, and that made it with the +server's default, which is everything. [Issue 207](../207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md) +closed this for a consumer re-made because its type changed. It did not close it for one that is simply +not there. + +## Why it matters + +**Delivery stops, and nothing says so.** Status showed every plan done and no machine behind. The merge +that was missed is not "behind", because no plan was ever made for it. A replay of past merges can also +act on them again. Issue 207 records nine modules re-registered from the past the same way. + +**The way out was a hand on the bus.** No verb resets a consumer. On the operator's explicit word, the +consumer was re-made from now with a one-off program run as the controller, otherwise unchanged. The +controller's loop still held the old event afterwards, so the new consumer's first delivery went +unacknowledged; the loop needs a restart of the controller to let go of it. + +## Noticed alongside, not this issue + +- Each merge to the catalogue planned 99 to 100 modules in two tiers and rebuilt modules it did not touch. + The output was byte-identical, so nothing was redeployed, but it costs minutes of the build machine + per merge. +- Three plans for three merges ran over each other. Each sent the build agent to every machine and asked + for the same builds. Nothing supersedes a plan for an older commit. diff --git a/04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/01-diagnosis.md b/04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/01-diagnosis.md new file mode 100644 index 0000000..c3840ba --- /dev/null +++ b/04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/01-diagnosis.md @@ -0,0 +1,39 @@ +# 248 — Diagnosis + +## 2026-10-05 + +1. A merge was not in the controller's log. The forge's own log showed nothing about delivering it, so the + question moved to the bus. +2. The bus's backlog tool named the controller's event consumer: 6 955 pending, redeliveries, one + unacknowledged. Its configuration was read from the server's monitoring endpoint: deliver policy + *all*, one outstanding, 30 seconds to acknowledge, five deliveries. +3. Sampled three times over a minute it did not move, and over the following hour it crawled forward + through week-old events. The controller's log held only its client's warnings: dropped messages, + and one refused reply. +4. The controller's code makes a missing consumer with the configuration it asserts, which sets no + deliver policy, so the server's default applies: everything. Issue 207's fix sets *from now* only on + the path where an existing consumer's type changes. + +**Unblocked**, on the operator's explicit word: the consumer re-made from now with its configuration +otherwise unchanged (nothing pending afterwards). The controller's loop still held the old event, so +it takes a restart of the controller to let go of it; that restart waits for the operator's word. + +**Fix** (mesh-controller, branch `fix/a-consumer-on-a-history-stream-starts-from-now`): + +- a consumer may say it starts **from now** when it is made. The controller's event consumer does. A + consumer that exists keeps where it is. The server would refuse a changed start anyway. +- `broker consumer-reset ` re-makes a stuck consumer from now, its configuration + otherwise kept, and refuses a work queue, where what is pending is work. It is the person's act, said by + a command, rather than a one-off program. + +Checked by a live test against a throwaway bus: + +- made from now, a consumer holds none of the stream's past and does hold the next announcement; +- asserted again, it keeps its place; +- the default replays all of it; +- a reset leaves nothing pending and keeps every other setting; +- a work queue's consumer is refused. + +**Not fixed here:** the client dropping messages while the loop acts on a long merge. Reports and +heartbeats are redelivered or replaced, so nothing is lost for good, but the loop holding everything while +it builds is the shape issue 175 already describes. diff --git a/04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md b/04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md new file mode 100644 index 0000000..4bcd302 --- /dev/null +++ b/04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md @@ -0,0 +1,37 @@ +--- +status: open +opened: 2026-10-05 +located-in: [] +fixed-by: +amended-design: +--- + +# 249 — A module's new state is refused until a push the merge did not make + +## What was observed + +A merge gave the agent module a new state, a key-value bucket +([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)). +The module's new bundle reached all four machines within a minute of the build, at once. The machines' +bus permissions did not include the new state until a push was made by hand afterwards. + +On three machines the module's watch of the new state was refused for about two minutes: "claude-code keeps +and reads no state called config". It recovered only because the module asks again with a back-off, and +the hand-made pushes issued the permissions. + +## Why it matters + +**The code arrives before the right to use it.** A module that does not retry stays broken until someone +pushes. A module whose first act on start is to read its new state fails its start. Nothing in the plan +says the two must travel together. + +**No machine went first.** The bundle reached every machine at the same moment. The rollout the operator +was told — one machine first, then the rest — could not be followed, because the merge had already +delivered it everywhere. + +## Open questions + +- Should a plan send a machine its membership, the grants that come with a module's new + declarations, in the same push as the bundle, and before it? +- Should a merge that changes a module's declarations (state, events, tools) be delivered to one machine + first, and to the rest only once that one reports it healthy?