Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: fixed
|
||||
opened: 2026-08-30
|
||||
located-in: [mesh-host]
|
||||
fixed-by: mesh-host — apply attempts every resource and reports every failure
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 011 — One broken module stops every module after it, for ever
|
||||
|
||||
## Symptom
|
||||
|
||||
A machine was assigned a module declaring a package that does not exist. Every later push to that
|
||||
machine applied **nothing at all**, and kept doing so.
|
||||
|
||||
Found while proving something else. A test assigned a deliberately-impossible module to a machine
|
||||
to check that the mesh reports a failure — which it does. A later test on the same machine then
|
||||
failed, and the evidence said why:
|
||||
|
||||
```
|
||||
applied 0 and failed: applying "impossible.nothing":
|
||||
installing a-package-that-does-not-exist: target not found
|
||||
0 resource(s) were applied before this and remain
|
||||
```
|
||||
|
||||
The broker's queues were **empty**, so the declaration had been delivered and read. The machine
|
||||
simply stopped at the first failing resource and never reached the rest.
|
||||
|
||||
## Why it matters more than one machine
|
||||
|
||||
- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports
|
||||
"failed" without saying that the rest were never attempted.
|
||||
- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the
|
||||
obvious next feature — would retry a permanent failure for ever and make no progress on
|
||||
everything else.
|
||||
- **The order is not the operator's.** Which module is "first" is an accident of resolution, so
|
||||
which nine modules a broken one blocks is unpredictable.
|
||||
|
||||
## What was there, and what it rested on
|
||||
|
||||
The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites
|
||||
[ADR 0010](../../02-DECISIONS/0010-delivery.md).
|
||||
|
||||
**That record does not decide this.** What it says is that a failed *job* stops and names its step
|
||||
while a reconciler retries forever, as an argument about pipelines against reconcilers. It says
|
||||
nothing about whether one resource failing should prevent the next from being attempted. The
|
||||
citation was doing more work than the record supports.
|
||||
|
||||
## The fix, and the argument that was on the other side
|
||||
|
||||
**Everything is attempted, and every failure is reported.**
|
||||
|
||||
The case for stopping is that a resource may depend on an earlier one — a service on the file it
|
||||
reads. That is real, and it survives: such a service fails its own check and is reported. This host
|
||||
reads back after every write precisely so a thing that did not work is caught rather than assumed,
|
||||
so attempting it produces *more* information than skipping it.
|
||||
|
||||
**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is
|
||||
applied. That is a different thing — *this machine could not do it* against *this was never a
|
||||
declaration* — and they are fixed in different places.
|
||||
|
||||
## How this is checked
|
||||
|
||||
`internal/apply`: an apply with a resource that cannot succeed still applies the ones after it,
|
||||
reports every failure, and says how many. Confirmed to fail if the loop stops at the first.
|
||||
Reference in New Issue
Block a user