Regenerate the decision index after merging main
This commit is contained in:
+35
@@ -0,0 +1,35 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-02
|
||||
located-in: [mesh-tools src/client.ts (toolsOn left node-scoped seats out of the listing), mesh-tools src/mcp.ts (the call resolved the key against that listing)]
|
||||
fixed-by: mesh-tools 27 — node-scoped seats are listed with their scope, the verb requires the machine, and the call carries it in the subject
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 199 — A node-scoped seat's verb could not be called through the console
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-02, the first time a node-scoped seat declared verbs
|
||||
([ADR 0170](../../02-DECISIONS/0170-the-firewall-seat-serves-its-verbs.md)). The packet filter's
|
||||
holder on every machine served `rules`, `reload` and `remove` on the seat's per-machine subjects, and
|
||||
the bus admitted them. The console answered every call with *nothing serves
|
||||
node-packet-filter.rules@<machine>*.
|
||||
|
||||
[Design 33 §4](../../03-DESIGN/01-to-be/33-the-tools-the-mesh-answers.md) says a node-scoped seat's
|
||||
tool carries the node it is asked of, as `<seat>.<verb>@<node>`. The console's listing left
|
||||
node-scoped seats out — the comment said they *wait for a caller naming the node* — but the roles map
|
||||
the console resolves a name against is built from that same listing. So the name never resolved as a
|
||||
seat's verb, fell through to a module's subject nobody served, and the refusal named the wrong cause.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
A stated behaviour that did not happen, with a refusal that pointed elsewhere: the seat's verbs were
|
||||
live on four machines and unreachable from the one surface a person uses. It could only be found by a
|
||||
node-scoped seat declaring verbs, which none had.
|
||||
|
||||
## Resolved, 2026-10-02
|
||||
|
||||
mesh-tools 27: node-scoped seats are listed with their scope, their verbs take a required `node`, the
|
||||
call carries it in the subject, and a call without one is refused in words. Tested with a round trip
|
||||
asking one machine's holder and being refused without a machine.
|
||||
+36
@@ -0,0 +1,36 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-02
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 200 — The controller's answer to a long console call is refused by the bus
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-02. A `push` of the control node asked through the console came back as *mesh-controller.push
|
||||
did not answer in time. Something is serving it, so this is the tool being slow rather than absent.*
|
||||
The push had run; the machine applied. The bus's log on the control node, at the same moment:
|
||||
|
||||
```
|
||||
[ERR] 10.10.0.1:56030 - cid:4015 - Publish Violation - User "controller",
|
||||
Subject "_INBOX.shanks.mesh-console.WRNO5V9IJ1FX35AU5NRBEB.WRNO5V9IJ1FX35AU5PN1UW"
|
||||
```
|
||||
|
||||
The controller's reply to the console's request was refused: the controller's bus account may not
|
||||
publish to the console's reply inbox. Shorter calls (`status`, `node`, `nodes`) answer; the long ones
|
||||
(`push` of a large machine, `issue`) time out on the console's side although they succeed.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
A call that succeeds and is reported as a timeout sends a person to retry what already happened — a
|
||||
second push, a second issue — and reads as the mesh being slow when it is the mesh refusing itself.
|
||||
Whether the inbox prefix the console uses is the one the controller's account is allowed to answer, or
|
||||
the request outlives the inbox subscription, is what a diagnosis has to tell apart.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Which reply inboxes may the controller's account publish to, and which does the console request on?
|
||||
- Does a reply after the requester's timeout count as a violation, or is the prefix itself refused?
|
||||
+88
@@ -0,0 +1,88 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-02
|
||||
located-in:
|
||||
- mesh-controller
|
||||
fixed-by: 02-DECISIONS/0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 201 — A push recreated the controller at a digest older than the seat row its successor had written
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-02, two merges a minute apart on the control node: one to the controller, adding a verb to the
|
||||
controller seat's row; one to the host, adding a container field. Each made a plan. The controller's plan
|
||||
built and rolled the new controller, which started, widened its seat row with the new verb, and ran. The
|
||||
host's plan then pushed the control node with the controller digest it had recorded when it was made —
|
||||
the previous build — and recreated the controller container on it. The older binary read the row, found
|
||||
a verb it could not run, and refused to start:
|
||||
|
||||
```
|
||||
mesh-controller: the mesh-controller seat's row declares "command", which this control plane
|
||||
cannot run: "command" is not a verb the mesh-controller seat serves
|
||||
```
|
||||
|
||||
A crash loop followed for ten minutes: nothing answered on the bus, and no build was dispatched, since
|
||||
the controller is what fills the builder's queue. Recovery was the mesh's own binary run once from the
|
||||
newer image, outside the service, to push the control node again; the push sent the newer digest and
|
||||
the controller came up.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
The row is the store's and the binary follows it ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md));
|
||||
a start-up check that refuses a row the binary cannot serve is right, and was built after the outage of
|
||||
2026-09-27 for exactly this reason. What is wrong is a plan sending a controller older than the one that
|
||||
wrote the row. A plan is made at a moment and sends what it recorded ([ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md));
|
||||
for every other module an older digest is a brief regression a later push corrects. For the controller
|
||||
it is the mesh losing its voice, and the correction needs a hand, because the thing that would correct
|
||||
it is the thing that is down. Two plans that overlap will happen again whenever two people merge within
|
||||
a minute.
|
||||
|
||||
## What a fix would have to do
|
||||
|
||||
Either of two, and the first is the smaller:
|
||||
|
||||
- A push never sends a controller digest older than the one the running controller is — the controller
|
||||
knows its own digest and refuses to downgrade itself through a plan, saying so in the plan's words.
|
||||
- Or the start-up check tolerates a row wider than the binary while a roll-out is in flight, and serves
|
||||
what it can. Weaker: it makes the row and the binary disagree on purpose, which is what the check
|
||||
exists to refuse.
|
||||
|
||||
Until one is built: do not merge a controller change while another plan is rolling, and after merging
|
||||
one, wait for `node show` on the control node to report the new controller before merging anything else.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md), [ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md)
|
||||
- mesh-controller `cmd/mesh-controller/seatverbs.go` (`seatToolHandlers`, the start-up check), `cmd/mesh-controller/push.go`
|
||||
|
||||
## Half of it is closed, 2026-10-02
|
||||
|
||||
[ADR 0185](../../02-DECISIONS/0185-a-control-plane-behind-its-seats-row-serves-what-it-can.md) takes
|
||||
the outage out of it: a control plane behind its seat's row now serves every verb it can run, says
|
||||
which it cannot, and answers the reason when one of those is called. The same race today would cost
|
||||
the verbs the newer build added, for as long as the older binary is in place, and the ordinary
|
||||
"this machine is behind" machinery would put the newer one back without a hand.
|
||||
|
||||
**The race itself is still open**, and this report stays open for it. What was established while
|
||||
closing the other half, so the next reader does not redo it:
|
||||
|
||||
- Composing and sending are serialised per machine by a session advisory lock in the store, so two
|
||||
control planes cannot compose one machine's declaration at the same time. The stale content did
|
||||
not come from two concurrent composes.
|
||||
- A container's image is resolved into the module's manifest when it is *built*, and a push composes
|
||||
from the catalogue as it is at that moment, under the hold. So a compose that ran after the build
|
||||
was taken in could not have named the older image.
|
||||
- The declaration's sequence orders arrival and nothing else (the numbering of
|
||||
[issue 107](../107-a-declaration-carries-no-order/00-report.md)); it cannot tell a later send
|
||||
carrying earlier content from a later send carrying later content. The host refuses a declaration
|
||||
numbered below the last it applied, and both of these were above it.
|
||||
- The machine's own journal shows the two applies ten seconds apart and which replaced what; it does
|
||||
not record which image each declaration named, which is the one fact that would settle it. A host
|
||||
that recorded the digest it was told, per apply, would have answered this in a minute.
|
||||
|
||||
So the trigger is not yet pinned, and guessing at the push path is the most expensive place in the
|
||||
mesh to guess. The fix the report first suggested — a push never sending a control plane a digest
|
||||
older than the one that machine reports running — closes the class without needing the trigger, and
|
||||
is now a correctness nicety rather than the difference between a working mesh and a dead one.
|
||||
Reference in New Issue
Block a user