diff --git a/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md b/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md new file mode 100644 index 0000000..7b95a33 --- /dev/null +++ b/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md @@ -0,0 +1,65 @@ +--- +status: fixed +opened: 2026-08-30 +located-in: [mesh-host] +fixed-by: mesh-host — apply attempts every resource and reports every failure +amended-design: +--- + +# 011 — One broken module stops every module after it, for ever + +## Symptom + +A machine was assigned a module declaring a package that does not exist. Every later push to that +machine applied **nothing at all**, and kept doing so. + +Found while proving something else. A test assigned a deliberately-impossible module to a machine +to check that the mesh reports a failure — which it does. A later test on the same machine then +failed, and the evidence said why: + +``` +applied 0 and failed: applying "impossible.nothing": + installing a-package-that-does-not-exist: target not found +0 resource(s) were applied before this and remain +``` + +The broker's queues were **empty**, so the declaration had been delivered and read. The machine +simply stopped at the first failing resource and never reached the rest. + +## Why it matters more than one machine + +- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports + "failed" without saying that the rest were never attempted. +- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the + obvious next feature — would retry a permanent failure for ever and make no progress on + everything else. +- **The order is not the operator's.** Which module is "first" is an accident of resolution, so + which nine modules a broken one blocks is unpredictable. + +## What was there, and what it rested on + +The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites +[ADR 0010](../../02-DECISIONS/0010-delivery.md). + +**That record does not decide this.** What it says is that a failed *job* stops and names its step +while a reconciler retries forever, as an argument about pipelines against reconcilers. It says +nothing about whether one resource failing should prevent the next from being attempted. The +citation was doing more work than the record supports. + +## The fix, and the argument that was on the other side + +**Everything is attempted, and every failure is reported.** + +The case for stopping is that a resource may depend on an earlier one — a service on the file it +reads. That is real, and it survives: such a service fails its own check and is reported. This host +reads back after every write precisely so a thing that did not work is caught rather than assumed, +so attempting it produces *more* information than skipping it. + +**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is +applied. That is a different thing — *this machine could not do it* against *this was never a +declaration* — and they are fixed in different places. + +## How this is checked + +`internal/apply`: an apply with a resource that cannot succeed still applies the ones after it, +reports every failure, and says how many. Confirmed to fail if the loop stops at the first.