Files
hq/04-ISSUES/201-a-push-recreated-the-controller-behind-the-row-its-successor-wrote/00-report.md
T

3.1 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-10-02
mesh-controller

201 — A push recreated the controller at a digest older than the seat row its successor had written

What was observed

2026-10-02, two merges a minute apart on the control node: one to the controller, adding a verb to the controller seat's row; one to the host, adding a container field. Each made a plan. The controller's plan built and rolled the new controller, which started, widened its seat row with the new verb, and ran. The host's plan then pushed the control node with the controller digest it had recorded when it was made — the previous build — and recreated the controller container on it. The older binary read the row, found a verb it could not run, and refused to start:

mesh-controller: the mesh-controller seat's row declares "command", which this control plane
cannot run: "command" is not a verb the mesh-controller seat serves

A crash loop followed for ten minutes: nothing answered on the bus, and no build was dispatched, since the controller is what fills the builder's queue. Recovery was the mesh's own binary run once from the newer image, outside the service, to push the control node again; the push sent the newer digest and the controller came up.

Why it matters beyond this instance

The row is the store's and the binary follows it (ADR 0154); a start-up check that refuses a row the binary cannot serve is right, and was built after the outage of 2026-09-27 for exactly this reason. What is wrong is a plan sending a controller older than the one that wrote the row. A plan is made at a moment and sends what it recorded (ADR 0162); for every other module an older digest is a brief regression a later push corrects. For the controller it is the mesh losing its voice, and the correction needs a hand, because the thing that would correct it is the thing that is down. Two plans that overlap will happen again whenever two people merge within a minute.

What a fix would have to do

Either of two, and the first is the smaller:

  • A push never sends a controller digest older than the one the running controller is — the controller knows its own digest and refuses to downgrade itself through a plan, saying so in the plan's words.
  • Or the start-up check tolerates a row wider than the binary while a roll-out is in flight, and serves what it can. Weaker: it makes the row and the binary disagree on purpose, which is what the check exists to refuse.

Until one is built: do not merge a controller change while another plan is rolling, and after merging one, wait for node show on the control node to report the new controller before merging anything else.

References

  • ADR 0154, ADR 0162
  • mesh-controller cmd/mesh-controller/seatverbs.go (seatToolHandlers, the start-up check), cmd/mesh-controller/push.go