Merge pull request 'Issues 248, 249: the controller's event consumer replayed a week; a new state is refused until a push' (#100) from issues/248-249-delivery into main
This commit was merged in pull request #100.
This commit is contained in:
+58
@@ -0,0 +1,58 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller internal/broker]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 248 — The controller's event consumer replayed a week, and held every new merge behind it
|
||||
|
||||
## What was observed
|
||||
|
||||
A merge to the catalogue at 15:17 never reached the controller. No plan was made and no build was asked
|
||||
for it. Every machine kept running the build from before it. The merge just before, to the record, was
|
||||
logged twice.
|
||||
|
||||
The bus showed why. The controller's durable consumer on the event stream was set to deliver
|
||||
**everything the stream holds**, not from where it last was:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| the stream | a week of events, from message 1 514 to message 356 004 |
|
||||
| the consumer delivered to | message 1 517, later 4 898 |
|
||||
| acknowledged to | 0, later 1 569 |
|
||||
| still to deliver | 6 955 merges and build outcomes, a week old |
|
||||
| allowed outstanding | one at a time ([issue 175](../175-an-announcement-behind-a-long-build-comes-back/00-report.md)) |
|
||||
|
||||
So the controller was working through a week of past merges and build outcomes, one at a time, slowly. Every
|
||||
new one — a merge, a build asked by hand — waited behind them. A build asked by hand finished on its
|
||||
machine and was never registered. While this went on, the controller's client dropped messages
|
||||
("slow consumer") several times, because heartbeats and reports share the loop with these events. From 15:17
|
||||
its event loop did nothing more: no line logged, the one delivered event never acknowledged.
|
||||
|
||||
How the consumer came to deliver everything is not certain. The bus and the stores were rebuilt the night
|
||||
before ([issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)).
|
||||
A consumer that is missing is made again by the controller's own assertion, and that made it with the
|
||||
server's default, which is everything. [Issue 207](../207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)
|
||||
closed this for a consumer re-made because its type changed. It did not close it for one that is simply
|
||||
not there.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**Delivery stops, and nothing says so.** Status showed every plan done and no machine behind. The merge
|
||||
that was missed is not "behind", because no plan was ever made for it. A replay of past merges can also
|
||||
act on them again. Issue 207 records nine modules re-registered from the past the same way.
|
||||
|
||||
**The way out was a hand on the bus.** No verb resets a consumer. On the operator's explicit word, the
|
||||
consumer was re-made from now with a one-off program run as the controller, otherwise unchanged. The
|
||||
controller's loop still held the old event afterwards, so the new consumer's first delivery went
|
||||
unacknowledged; the loop needs a restart of the controller to let go of it.
|
||||
|
||||
## Noticed alongside, not this issue
|
||||
|
||||
- Each merge to the catalogue planned 99 to 100 modules in two tiers and rebuilt modules it did not touch.
|
||||
The output was byte-identical, so nothing was redeployed, but it costs minutes of the build machine
|
||||
per merge.
|
||||
- Three plans for three merges ran over each other. Each sent the build agent to every machine and asked
|
||||
for the same builds. Nothing supersedes a plan for an older commit.
|
||||
+39
@@ -0,0 +1,39 @@
|
||||
# 248 — Diagnosis
|
||||
|
||||
## 2026-10-05
|
||||
|
||||
1. A merge was not in the controller's log. The forge's own log showed nothing about delivering it, so the
|
||||
question moved to the bus.
|
||||
2. The bus's backlog tool named the controller's event consumer: 6 955 pending, redeliveries, one
|
||||
unacknowledged. Its configuration was read from the server's monitoring endpoint: deliver policy
|
||||
*all*, one outstanding, 30 seconds to acknowledge, five deliveries.
|
||||
3. Sampled three times over a minute it did not move, and over the following hour it crawled forward
|
||||
through week-old events. The controller's log held only its client's warnings: dropped messages,
|
||||
and one refused reply.
|
||||
4. The controller's code makes a missing consumer with the configuration it asserts, which sets no
|
||||
deliver policy, so the server's default applies: everything. Issue 207's fix sets *from now* only on
|
||||
the path where an existing consumer's type changes.
|
||||
|
||||
**Unblocked**, on the operator's explicit word: the consumer re-made from now with its configuration
|
||||
otherwise unchanged (nothing pending afterwards). The controller's loop still held the old event, so
|
||||
it takes a restart of the controller to let go of it; that restart waits for the operator's word.
|
||||
|
||||
**Fix** (mesh-controller, branch `fix/a-consumer-on-a-history-stream-starts-from-now`):
|
||||
|
||||
- a consumer may say it starts **from now** when it is made. The controller's event consumer does. A
|
||||
consumer that exists keeps where it is. The server would refuse a changed start anyway.
|
||||
- `broker consumer-reset <stream> <consumer>` re-makes a stuck consumer from now, its configuration
|
||||
otherwise kept, and refuses a work queue, where what is pending is work. It is the person's act, said by
|
||||
a command, rather than a one-off program.
|
||||
|
||||
Checked by a live test against a throwaway bus:
|
||||
|
||||
- made from now, a consumer holds none of the stream's past and does hold the next announcement;
|
||||
- asserted again, it keeps its place;
|
||||
- the default replays all of it;
|
||||
- a reset leaves nothing pending and keeps every other setting;
|
||||
- a work queue's consumer is refused.
|
||||
|
||||
**Not fixed here:** the client dropping messages while the loop acts on a long merge. Reports and
|
||||
heartbeats are redelivered or replaced, so nothing is lost for good, but the loop holding everything while
|
||||
it builds is the shape issue 175 already describes.
|
||||
+37
@@ -0,0 +1,37 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 249 — A module's new state is refused until a push the merge did not make
|
||||
|
||||
## What was observed
|
||||
|
||||
A merge gave the agent module a new state, a key-value bucket
|
||||
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)).
|
||||
The module's new bundle reached all four machines within a minute of the build, at once. The machines'
|
||||
bus permissions did not include the new state until a push was made by hand afterwards.
|
||||
|
||||
On three machines the module's watch of the new state was refused for about two minutes: "claude-code keeps
|
||||
and reads no state called config". It recovered only because the module asks again with a back-off, and
|
||||
the hand-made pushes issued the permissions.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**The code arrives before the right to use it.** A module that does not retry stays broken until someone
|
||||
pushes. A module whose first act on start is to read its new state fails its start. Nothing in the plan
|
||||
says the two must travel together.
|
||||
|
||||
**No machine went first.** The bundle reached every machine at the same moment. The rollout the operator
|
||||
was told — one machine first, then the rest — could not be followed, because the merge had already
|
||||
delivered it everywhere.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a plan send a machine its membership, the grants that come with a module's new
|
||||
declarations, in the same push as the bundle, and before it?
|
||||
- Should a merge that changes a module's declarations (state, events, tools) be delivered to one machine
|
||||
first, and to the rest only once that one reports it healthy?
|
||||
Reference in New Issue
Block a user