Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
Showing only changes of commit 4a601ceb02 - Show all commits
@@ -0,0 +1,65 @@
---
status: fixed
opened: 2026-08-30
located-in: [mesh-host]
fixed-by: mesh-host — apply attempts every resource and reports every failure
amended-design:
---
# 011 — One broken module stops every module after it, for ever
## Symptom
A machine was assigned a module declaring a package that does not exist. Every later push to that
machine applied **nothing at all**, and kept doing so.
Found while proving something else. A test assigned a deliberately-impossible module to a machine
to check that the mesh reports a failure — which it does. A later test on the same machine then
failed, and the evidence said why:
```
applied 0 and failed: applying "impossible.nothing":
installing a-package-that-does-not-exist: target not found
0 resource(s) were applied before this and remain
```
The broker's queues were **empty**, so the declaration had been delivered and read. The machine
simply stopped at the first failing resource and never reached the rest.
## Why it matters more than one machine
- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports
"failed" without saying that the rest were never attempted.
- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the
obvious next feature — would retry a permanent failure for ever and make no progress on
everything else.
- **The order is not the operator's.** Which module is "first" is an accident of resolution, so
which nine modules a broken one blocks is unpredictable.
## What was there, and what it rested on
The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites
[ADR 0010](../../02-DECISIONS/0010-delivery.md).
**That record does not decide this.** What it says is that a failed *job* stops and names its step
while a reconciler retries forever, as an argument about pipelines against reconcilers. It says
nothing about whether one resource failing should prevent the next from being attempted. The
citation was doing more work than the record supports.
## The fix, and the argument that was on the other side
**Everything is attempted, and every failure is reported.**
The case for stopping is that a resource may depend on an earlier one — a service on the file it
reads. That is real, and it survives: such a service fails its own check and is reported. This host
reads back after every write precisely so a thing that did not work is caught rather than assumed,
so attempting it produces *more* information than skipping it.
**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is
applied. That is a different thing — *this machine could not do it* against *this was never a
declaration* — and they are fixed in different places.
## How this is checked
`internal/apply`: an apply with a resource that cannot succeed still applies the ones after it,
reports every failure, and says how many. Confirmed to fail if the loop stops at the first.