Files
hq/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md
T
jschoubben 4a601ceb02 Issue 011 — one broken module stops every module after it
Found in the lab. A machine with one impossible module applied nothing at
all on every later push, and the mesh said "failed" without saying the
rest was never attempted.

Recorded with the evidence, including that the behaviour's test cited a
record which does not decide it: ADR 0010 argues about pipelines against
reconcilers and says nothing about whether one resource failing should
stop the next being attempted.

Fixed in mesh-host: everything is attempted, every failure reported.
2026-08-30 19:26:03 +02:00

2.9 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
fixed 2026-08-30
mesh-host
mesh-host — apply attempts every resource and reports every failure

011 — One broken module stops every module after it, for ever

Symptom

A machine was assigned a module declaring a package that does not exist. Every later push to that machine applied nothing at all, and kept doing so.

Found while proving something else. A test assigned a deliberately-impossible module to a machine to check that the mesh reports a failure — which it does. A later test on the same machine then failed, and the evidence said why:

applied 0 and failed: applying "impossible.nothing":
  installing a-package-that-does-not-exist: target not found
0 resource(s) were applied before this and remain

The broker's queues were empty, so the declaration had been delivered and read. The machine simply stopped at the first failing resource and never reached the rest.

Why it matters more than one machine

  • A machine with one bad module and nine good ones runs none of the nine, and the mesh reports "failed" without saying that the rest were never attempted.
  • It cannot be recovered by retrying. Anything that re-pushes to machines that are behind — the obvious next feature — would retry a permanent failure for ever and make no progress on everything else.
  • The order is not the operator's. Which module is "first" is an accident of resolution, so which nine modules a broken one blocks is unpredictable.

What was there, and what it rested on

The behaviour had a test asserting it: nothing after the failure ran. Its comment cites ADR 0010.

That record does not decide this. What it says is that a failed job stops and names its step while a reconciler retries forever, as an argument about pipelines against reconcilers. It says nothing about whether one resource failing should prevent the next from being attempted. The citation was doing more work than the record supports.

The fix, and the argument that was on the other side

Everything is attempted, and every failure is reported.

The case for stopping is that a resource may depend on an earlier one — a service on the file it reads. That is real, and it survives: such a service fails its own check and is reported. This host reads back after every write precisely so a thing that did not work is caught rather than assumed, so attempting it produces more information than skipping it.

What is unchanged: a declaration that cannot be parsed is still refused whole, and nothing is applied. That is a different thing — this machine could not do it against this was never a declaration — and they are fixed in different places.

How this is checked

internal/apply: an apply with a resource that cannot succeed still applies the ones after it, reports every failure, and says how many. Confirmed to fail if the loop stops at the first.