Found in the lab. A machine with one impossible module applied nothing at all on every later push, and the mesh said "failed" without saying the rest was never attempted. Recorded with the evidence, including that the behaviour's test cited a record which does not decide it: ADR 0010 argues about pipelines against reconcilers and says nothing about whether one resource failing should stop the next being attempted. Fixed in mesh-host: everything is attempted, every failure reported.
2.9 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |
|---|---|---|---|---|---|
| fixed | 2026-08-30 |
|
mesh-host — apply attempts every resource and reports every failure |
011 — One broken module stops every module after it, for ever
Symptom
A machine was assigned a module declaring a package that does not exist. Every later push to that machine applied nothing at all, and kept doing so.
Found while proving something else. A test assigned a deliberately-impossible module to a machine to check that the mesh reports a failure — which it does. A later test on the same machine then failed, and the evidence said why:
applied 0 and failed: applying "impossible.nothing":
installing a-package-that-does-not-exist: target not found
0 resource(s) were applied before this and remain
The broker's queues were empty, so the declaration had been delivered and read. The machine simply stopped at the first failing resource and never reached the rest.
Why it matters more than one machine
- A machine with one bad module and nine good ones runs none of the nine, and the mesh reports "failed" without saying that the rest were never attempted.
- It cannot be recovered by retrying. Anything that re-pushes to machines that are behind — the obvious next feature — would retry a permanent failure for ever and make no progress on everything else.
- The order is not the operator's. Which module is "first" is an accident of resolution, so which nine modules a broken one blocks is unpredictable.
What was there, and what it rested on
The behaviour had a test asserting it: nothing after the failure ran. Its comment cites ADR 0010.
That record does not decide this. What it says is that a failed job stops and names its step while a reconciler retries forever, as an argument about pipelines against reconcilers. It says nothing about whether one resource failing should prevent the next from being attempted. The citation was doing more work than the record supports.
The fix, and the argument that was on the other side
Everything is attempted, and every failure is reported.
The case for stopping is that a resource may depend on an earlier one — a service on the file it reads. That is real, and it survives: such a service fails its own check and is reported. This host reads back after every write precisely so a thing that did not work is caught rather than assumed, so attempting it produces more information than skipping it.
What is unchanged: a declaration that cannot be parsed is still refused whole, and nothing is applied. That is a different thing — this machine could not do it against this was never a declaration — and they are fixed in different places.
How this is checked
internal/apply: an apply with a resource that cannot succeed still applies the ones after it,
reports every failure, and says how many. Confirmed to fail if the loop stops at the first.