Files
hq/04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md
T
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00

4.9 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-29
mesh-controller internal/broker/jetstream.go (EnsureConsumer)
mesh-controller e7da39d (PR 148)

156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running

What was observed

The control plane crash-looped, every restart ending the same way:

mesh-controller: asserting how novox hears its declaration:
  bringing consumer novox on NODES to match: nats: consumer name already in use

It came up on the first build of the controller in eight hours. The machines kept running what they already held — this stops the mesh being changed, not the services it placed — and nothing could be pushed, no report was consumed and no enrolment answered, for as long as it lasted.

Why

Issue 146 put the stream into a push consumer's delivery subject, because one process holding two consumers of the same name on two streams was given one subject and acted on every message twice.

The server will not move a push consumer's delivery subject while a subscriber is bound to it. It refuses with consumer name already in use — a message about the name, for a conflict about the subject, which is why the trail starts in the wrong place.

A node is bound to its declaration consumer the whole time it is up. That is a node listening for what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring to match — and the assertion happens before the controller serves, so it never serves.

The controller's own two consumers moved without trouble, and are on the new subject in the live mesh. It asserts them before it subscribes, so nothing was bound.

Why nothing caught it

The change was exercised on a mesh being raised, where every consumer is created rather than updated and nothing is bound to any of them. On that path the code is correct. The test that would have caught it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then the assertion.

Reproduced exactly that way before the fix — same server version, same stream shape, same consumer — and it fails with the same words as the machine did. An earlier version of the same test unsubscribed first and passed against the code that was crash-looping on the control node.

What it is not

  • Not the change it shipped beside. The merge that triggered this build carried issue 152's fix and three other commits that had never been deployed; this one is 146's.
  • Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.

How it was fixed

The consumer that works is kept, and the assertion says so instead of failing.

Not deleted and re-made. Re-making moves the subject, and a holder may not yet be allowed to subscribe to the new one: the wider grant travels in the bus's user list, which the control plane composes and a machine applies minutes later. On the live mesh the nodes are granted _DELIVER.<node> and not _DELIVER.<node>.> — re-making would have silenced every machine in the mesh, which is worse than the collision it was fixing and far harder to undo. That was the first fix written here, and the permission is the reason it was not shipped.

Not fatal, which is what 146's change intended and did not do: the bare subject still delivers, and collides only where one holder has two consumers of one name. A node has one.

How the fix is checked

Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported, and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies where the collision actually was.

What is left

The node consumers stay on the bare subject, which is correct and not tidy. Nothing is wrong while they do not move: one consumer per name per stream cannot collide with itself.

The controller reports each one it kept, and did, on the start that fixed this — four node consumers and the build machine's worker, which is bound the same way and was not anticipated here:

consumer novox on NODES still delivers to "_DELIVER.novox" and not "_DELIVER.novox.NODES":
  nats: consumer name already in use. It keeps working; the subject moves on an assertion
  made while nothing is bound to it

The wider grant has since landed (2026-09-30, measured on the mesh's own broker config): every node is now allowed _DELIVER.<node>.> as well as the bare subject. That was the thing missing when this was diagnosed, and it is why re-making the consumers then would have silenced every machine. What remains is only the second half — an assertion made while each node is detached from its consumer — and that is its own piece of work, not a side effect of a restart.