Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -0,0 +1,65 @@
|
|||||||
|
---
|
||||||
|
status: fixed
|
||||||
|
opened: 2026-08-30
|
||||||
|
located-in: [mesh-host]
|
||||||
|
fixed-by: mesh-host — apply attempts every resource and reports every failure
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 011 — One broken module stops every module after it, for ever
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A machine was assigned a module declaring a package that does not exist. Every later push to that
|
||||||
|
machine applied **nothing at all**, and kept doing so.
|
||||||
|
|
||||||
|
Found while proving something else. A test assigned a deliberately-impossible module to a machine
|
||||||
|
to check that the mesh reports a failure — which it does. A later test on the same machine then
|
||||||
|
failed, and the evidence said why:
|
||||||
|
|
||||||
|
```
|
||||||
|
applied 0 and failed: applying "impossible.nothing":
|
||||||
|
installing a-package-that-does-not-exist: target not found
|
||||||
|
0 resource(s) were applied before this and remain
|
||||||
|
```
|
||||||
|
|
||||||
|
The broker's queues were **empty**, so the declaration had been delivered and read. The machine
|
||||||
|
simply stopped at the first failing resource and never reached the rest.
|
||||||
|
|
||||||
|
## Why it matters more than one machine
|
||||||
|
|
||||||
|
- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports
|
||||||
|
"failed" without saying that the rest were never attempted.
|
||||||
|
- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the
|
||||||
|
obvious next feature — would retry a permanent failure for ever and make no progress on
|
||||||
|
everything else.
|
||||||
|
- **The order is not the operator's.** Which module is "first" is an accident of resolution, so
|
||||||
|
which nine modules a broken one blocks is unpredictable.
|
||||||
|
|
||||||
|
## What was there, and what it rested on
|
||||||
|
|
||||||
|
The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites
|
||||||
|
[ADR 0010](../../02-DECISIONS/0010-delivery.md).
|
||||||
|
|
||||||
|
**That record does not decide this.** What it says is that a failed *job* stops and names its step
|
||||||
|
while a reconciler retries forever, as an argument about pipelines against reconcilers. It says
|
||||||
|
nothing about whether one resource failing should prevent the next from being attempted. The
|
||||||
|
citation was doing more work than the record supports.
|
||||||
|
|
||||||
|
## The fix, and the argument that was on the other side
|
||||||
|
|
||||||
|
**Everything is attempted, and every failure is reported.**
|
||||||
|
|
||||||
|
The case for stopping is that a resource may depend on an earlier one — a service on the file it
|
||||||
|
reads. That is real, and it survives: such a service fails its own check and is reported. This host
|
||||||
|
reads back after every write precisely so a thing that did not work is caught rather than assumed,
|
||||||
|
so attempting it produces *more* information than skipping it.
|
||||||
|
|
||||||
|
**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is
|
||||||
|
applied. That is a different thing — *this machine could not do it* against *this was never a
|
||||||
|
declaration* — and they are fixed in different places.
|
||||||
|
|
||||||
|
## How this is checked
|
||||||
|
|
||||||
|
`internal/apply`: an apply with a resource that cannot succeed still applies the ones after it,
|
||||||
|
reports every failure, and says how many. Confirmed to fail if the loop stops at the first.
|
||||||
Reference in New Issue
Block a user