Files
hq/04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md
T
jschoubben 36d9b38a0d An issue is open, diagnosing, located, resolved or wontfix — nothing else
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
2026-09-21 17:38:17 +02:00

4.3 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-08-30
mesh-host
mesh-host — apply attempts every resource and reports every failure

011 — One broken module stops every module after it, for ever

Symptom

A machine was assigned a module declaring a package that does not exist. Every later push to that machine applied nothing at all, and kept doing so.

Found while proving something else. A test assigned a deliberately-impossible module to a machine to check that the mesh reports a failure — which it does. A later test on the same machine then failed, and the evidence said why:

applied 0 and failed: applying "impossible.nothing":
  installing a-package-that-does-not-exist: target not found
0 resource(s) were applied before this and remain

The broker's queues were empty, so the declaration had been delivered and read. The machine simply stopped at the first failing resource and never reached the rest.

Why it matters more than one machine

  • A machine with one bad module and nine good ones runs none of the nine, and the mesh reports "failed" without saying that the rest were never attempted.
  • It cannot be recovered by retrying. Anything that re-pushes to machines that are behind — the obvious next feature — would retry a permanent failure for ever and make no progress on everything else.
  • The order is not the operator's. Which module is "first" is an accident of resolution, so which nine modules a broken one blocks is unpredictable.

What was there, and what it rested on

The behaviour had a test asserting it: nothing after the failure ran. Its comment cites ADR 0010.

That record does not decide this. What it says is that a failed job stops and names its step while a reconciler retries forever, as an argument about pipelines against reconcilers. It says nothing about whether one resource failing should prevent the next from being attempted. The citation was doing more work than the record supports.

The fix, and the argument that was on the other side

Everything is attempted, and every failure is reported.

The case for stopping is that a resource may depend on an earlier one — a service on the file it reads. That is real, and it survives: such a service fails its own check and is reported. This host reads back after every write precisely so a thing that did not work is caught rather than assumed, so attempting it produces more information than skipping it.

What is unchanged: a declaration that cannot be parsed is still refused whole, and nothing is applied. That is a different thing — this machine could not do it against this was never a declaration — and they are fixed in different places.

And one shape still stops what follows, which the first fix got wrong

An action does. The first version of this fix continued past everything, and the next lab run failed at the bootstrap: the store did not answer in three minutes and then said the database system is shutting down. Carrying on past the readiness gate had started the broker and the control plane against a machine that was not ready, and on a small machine that is how a database still initialising has its memory taken away.

An action is the only shape whose purpose is to make something true before the next thing needs it — which is why it is the only one with a verify. The bootstrap is a row of them: the store answers, then its databases exist, then their schemas, then the broker. Everything else is independent state: a package that will not install has nothing to do with a file on the other side of the declaration.

So the rule is: a failed action stops what follows; nothing else does. Both faults are fixed by it, and the report says which happened — these things failed and these things failed and the rest was never tried are different machines.

How this is checked

internal/apply, three tests: an apply with a resource that cannot succeed still applies the ones after it; every failure is counted, not just the first; and a failed action stops what follows and says so. Each confirmed to fail when its behaviour is removed — including the last, which fails if actions stop being treated as gates and if everything is treated as one.