Issue 204: diagnosed and located — a declaration was numbered after composing, and a dying sender recorded nothing

This commit is contained in:
jochen
2026-10-03 04:03:51 +02:00
parent 4891cdeac5
commit bf3338898e
2 changed files with 38 additions and 1 deletions
@@ -1,5 +1,5 @@
---
status: open
status: located
opened: 2026-10-02
located-in:
- mesh-controller
@@ -0,0 +1,37 @@
# Diagnosis — 204
**2026-10-03, in the controller's code.**
Ruled out: a start-up re-send (a starting controller sends nothing; it asserts the bus, resumes plans,
follows events); a plan sending a recorded declaration (a plan records artifact digests and module
states, and its rollout composes fresh at send time); a cache (every compose reads the store); a path
that does not record its send (push, the cascade and the rollout all record after sending; only the raw
`declare <node> <file>` command did not); the host applying an older sequence (it already refuses a
declaration numbered below the one it kept).
Found, two faults that together produce the evidence:
1. **The sequence number went on at send time, after composing.** Every sending path composed first and
numbered each declaration as it was sent; a multi-machine send composes every machine before sending
any. So a declaration composed *before* an assignment changed and sent *after* a fresher one carried
the higher number — and the host, refusing only lower numbers, applied the older content as the mesh's
newest word. The stale declaration was accepted, so its number was higher, so it was composed earlier
and sent later.
2. **The send record was written on the sender's own context, after the send.** A controller being
replaced in that second has its context cancelled between telling the machine and writing the record;
the machine was told, the record never written, and status showed only the person's earlier send.
The likely sender, consistent with both and with the timing: the outgoing controller's reaction to a
catalogue registration during the build round, which re-sends the machines running the registered
module (one ran on exactly the two machines affected), composed under its hold before the person's
assignments, numbered and sent at 21:29:19–20 as the controller was being replaced. The old container's
log is gone, so the sender is inferred from code and timing; the mechanism is not.
**Owner:** mesh-controller. **Fix direction:** number a declaration before composing it, in every path,
so what was composed earlier is numbered lower whatever order the sends happen in and the host's
existing refusal does its job; record a send on a context that outlives the sender; the raw `declare`
command records too.
Two questions left for HQ: whether status should show the sequence a machine was last sent beside the
digest, and whether a sender's hold should also cover the assignment verbs, which today run between a
hold's compose and its send without waiting.