Issue 156: moving a consumer's delivery subject stops a running mesh #193

Merged
jschoubben merged 1 commits from issue/156-a-consumer-that-works-is-not-replaced into main 2026-09-29 21:57:00 +00:00
@@ -0,0 +1,85 @@
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller internal/broker/jetstream.go (EnsureConsumer)
fixed-by: mesh-controller fix/156-a-consumer-that-works-is-not-replaced
amended-design:
---
# 156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running
## What was observed
The control plane crash-looped, every restart ending the same way:
```
mesh-controller: asserting how novox hears its declaration:
bringing consumer novox on NODES to match: nats: consumer name already in use
```
It came up on the first build of the controller in eight hours. The machines kept running what they
already held — this stops the mesh being *changed*, not the services it placed — and nothing could be
pushed, no report was consumed and no enrolment answered, for as long as it lasted.
## Why
[Issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md) put the
stream into a push consumer's delivery subject, because one process holding two consumers of the same
name on two streams was given one subject and acted on every message twice.
**The server will not move a push consumer's delivery subject while a subscriber is bound to it.** It
refuses with `consumer name already in use` — a message about the name, for a conflict about the
subject, which is why the trail starts in the wrong place.
A node is bound to its declaration consumer the whole time it is up. That *is* a node listening for
what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring
to match — and the assertion happens before the controller serves, so it never serves.
The controller's own two consumers moved without trouble, and are on the new subject in the live mesh.
It asserts them before it subscribes, so nothing was bound.
## Why nothing caught it
The change was exercised on a mesh being raised, where every consumer is created rather than updated
and nothing is bound to any of them. On that path the code is correct. The test that would have caught
it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then
the assertion.
Reproduced exactly that way before the fix — same server version, same stream shape, same consumer —
and it fails with the same words as the machine did. An earlier version of the same test unsubscribed
first and passed against the code that was crash-looping on the control node.
## What it is not
- Not the change it shipped beside. The merge that triggered this build carried
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md)'s fix and three
other commits that had never been deployed; this one is 146's.
- Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.
## How it was fixed
The consumer that works is kept, and the assertion says so instead of failing.
**Not deleted and re-made.** Re-making moves the subject, and a holder may not yet be allowed to
subscribe to the new one: the wider grant travels in the bus's user list, which the control plane
composes and a machine applies minutes later. On the live mesh the nodes are granted `_DELIVER.<node>`
and not `_DELIVER.<node>.>` — re-making would have silenced every machine in the mesh, which is worse
than the collision it was fixing and far harder to undo. That was the first fix written here, and the
permission is the reason it was not shipped.
**Not fatal**, which is what 146's change intended and did not do: the bare subject still delivers, and
collides only where one holder has two consumers of one name. A node has one.
## How the fix is checked
Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported,
and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies
where the collision actually was.
## What is left
The node consumers stay on the bare subject, which is correct and not tidy. Moving them needs the
wider grant to reach every machine first, and then an assertion made while each node is detached —
which is its own piece of work, not a side effect of a restart. Nothing is wrong while they do not
move: one consumer per name per stream cannot collide with itself.