Merge pull request 'Issue 264: a self-updating engine lost the report of the apply that delivered it' (#120) from issues/264-self-update-lost-report into main

This commit was merged in pull request #120.
This commit is contained in:
2026-10-05 22:31:47 +00:00
@@ -0,0 +1,72 @@
---
status: located
opened: 2026-10-06
located-in: [mesh-host]
fixed-by:
amended-design:
---
# 264. A self-updating engine lost the report of the apply that delivered it
## Symptom
One declaration for the anchor carried both a new controller and a new version of the host. The
host applied it and said, in its journal:
- `host <version> is delivered; standing aside so the launcher runs it`
- `applied, and could not tell the mesh: reporting: context canceled`
The new host started and said `in the mesh, hearing what this node should be`, and nothing more
about that declaration. The controller's release plan waits for the first machine to report the
exact digest of the declaration it was sent. That report never came, so the plan waited until a
person pushed again by hand.
This is the same family as [issue 257](../257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md)
and [issue 261](../261-a-module-was-applied-and-given-back-half-a-minute-later/00-report.md): the
machine did what it was sent, and the report the plan waits on did not say so.
## Cause
Standing aside for a successor ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md)) is
decided inside the apply, once the apply has finished, and it cancels the link's context. The link
then published the apply's report on that same cancelled context, and the bus refused it before it
left. The declaration was then acknowledged anyway, so the mesh did not deliver it again, and the
new host had nothing that told it a report was owed.
A crash or power cut between the apply and the report has the same shape. The declaration is
delivered again in that case only because it was not yet acknowledged, and nothing on the machine
records that a report is owed.
## Fix
Two parts, in the host:
1. The report of an apply, and of a declaration set aside for a newer one, is published on a context
that standing aside does not cancel. It is still limited by the host's timeout, and it is sent
before the connection is closed. After an apply that stood aside, nothing further is applied.
2. The report of the last apply of a declaration from the mesh is kept beside the node's state until
the broker has taken it. When a host links, it first sends any report still kept, before it applies
anything newly delivered, so the re-sent report cannot arrive after a newer one. What is re-sent
is exactly the report the apply made, with the same digest and the same outcome. It is never a new
description of the machine, so it cannot claim that something was applied when it was not. A
report that names no declaration, such as a refused one, clears what was kept. A node that never
applied anything sends nothing.
## How it is checked
These tests in the host fail without the fix:
- `TestAnApplyThatEndsTheLinkIsStillReported`: an apply that cancels the link still has its report
delivered, and its declaration is settled after the report.
- `TestAReportLostAfterTheApplyIsSaidOnTheNextLink`: a report the broker refused is sent again on the
next link, unchanged.
These tests guard the edges:
- `TestNothingIsSaidAgainWhenNothingWasLost`: nothing is re-sent when nothing was lost.
- `TestTheKeptReportSurvivesTheHost` and `TestNothingKeptWhenNothingWasApplied`: the kept report
survives a restart exactly as it was made, and it is cleared only by a report about the same
declaration.
Once the fix is rolled out, the live check is the next declaration that carries a new host version.
The first machine's report should reach the plan without anyone pushing by hand.