5.4 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |
|---|---|---|---|---|---|
| open | 2026-10-02 |
|
02-DECISIONS/0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md |
201 — A push recreated the controller at a digest older than the seat row its successor had written
What was observed
2026-10-02, two merges a minute apart on the control node: one to the controller, adding a verb to the controller seat's row; one to the host, adding a container field. Each made a plan. The controller's plan built and rolled the new controller, which started, widened its seat row with the new verb, and ran. The host's plan then pushed the control node with the controller digest it had recorded when it was made — the previous build — and recreated the controller container on it. The older binary read the row, found a verb it could not run, and refused to start:
mesh-controller: the mesh-controller seat's row declares "command", which this control plane
cannot run: "command" is not a verb the mesh-controller seat serves
A crash loop followed for ten minutes: nothing answered on the bus, and no build was dispatched, since the controller is what fills the builder's queue. Recovery was the mesh's own binary run once from the newer image, outside the service, to push the control node again; the push sent the newer digest and the controller came up.
Why it matters beyond this instance
The row is the store's and the binary follows it (ADR 0154); a start-up check that refuses a row the binary cannot serve is right, and was built after the outage of 2026-09-27 for exactly this reason. What is wrong is a plan sending a controller older than the one that wrote the row. A plan is made at a moment and sends what it recorded (ADR 0162); for every other module an older digest is a brief regression a later push corrects. For the controller it is the mesh losing its voice, and the correction needs a hand, because the thing that would correct it is the thing that is down. Two plans that overlap will happen again whenever two people merge within a minute.
What a fix would have to do
Either of two, and the first is the smaller:
- A push never sends a controller digest older than the one the running controller is — the controller knows its own digest and refuses to downgrade itself through a plan, saying so in the plan's words.
- Or the start-up check tolerates a row wider than the binary while a roll-out is in flight, and serves what it can. Weaker: it makes the row and the binary disagree on purpose, which is what the check exists to refuse.
Until one is built: do not merge a controller change while another plan is rolling, and after merging
one, wait for node show on the control node to report the new controller before merging anything else.
References
- ADR 0154, ADR 0162
- mesh-controller
cmd/mesh-controller/seatverbs.go(seatToolHandlers, the start-up check),cmd/mesh-controller/push.go
Half of it is closed, 2026-10-02
ADR 0185 takes the outage out of it: a control plane behind its seat's row now serves every verb it can run, says which it cannot, and answers the reason when one of those is called. The same race today would cost the verbs the newer build added, for as long as the older binary is in place, and the ordinary "this machine is behind" machinery would put the newer one back without a hand.
The race itself is still open, and this report stays open for it. What was established while closing the other half, so the next reader does not redo it:
- Composing and sending are serialised per machine by a session advisory lock in the store, so two control planes cannot compose one machine's declaration at the same time. The stale content did not come from two concurrent composes.
- A container's image is resolved into the module's manifest when it is built, and a push composes from the catalogue as it is at that moment, under the hold. So a compose that ran after the build was taken in could not have named the older image.
- The declaration's sequence orders arrival and nothing else (the numbering of issue 107); it cannot tell a later send carrying earlier content from a later send carrying later content. The host refuses a declaration numbered below the last it applied, and both of these were above it.
- The machine's own journal shows the two applies ten seconds apart and which replaced what; it does not record which image each declaration named, which is the one fact that would settle it. A host that recorded the digest it was told, per apply, would have answered this in a minute.
So the trigger is not yet pinned, and guessing at the push path is the most expensive place in the mesh to guess. The fix the report first suggested — a push never sending a control plane a digest older than the one that machine reports running — closes the class without needing the trigger, and is now a correctness nicety rather than the difference between a working mesh and a dead one.