The playbook, the README and the status skill knew five statuses; the cycle check knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks with their fixes shipped, one step short of closed. They are resolved; the check refuses the word from now on and accepts the one the playbook allows.
86 lines
4.3 KiB
Markdown
86 lines
4.3 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-08-30
|
|
located-in: [mesh-host]
|
|
fixed-by: mesh-host — apply attempts every resource and reports every failure
|
|
amended-design:
|
|
---
|
|
|
|
# 011 — One broken module stops every module after it, for ever
|
|
|
|
## Symptom
|
|
|
|
A machine was assigned a module declaring a package that does not exist. Every later push to that
|
|
machine applied **nothing at all**, and kept doing so.
|
|
|
|
Found while proving something else. A test assigned a deliberately-impossible module to a machine
|
|
to check that the mesh reports a failure — which it does. A later test on the same machine then
|
|
failed, and the evidence said why:
|
|
|
|
```
|
|
applied 0 and failed: applying "impossible.nothing":
|
|
installing a-package-that-does-not-exist: target not found
|
|
0 resource(s) were applied before this and remain
|
|
```
|
|
|
|
The broker's queues were **empty**, so the declaration had been delivered and read. The machine
|
|
simply stopped at the first failing resource and never reached the rest.
|
|
|
|
## Why it matters more than one machine
|
|
|
|
- **A machine with one bad module and nine good ones runs none of the nine**, and the mesh reports
|
|
"failed" without saying that the rest were never attempted.
|
|
- **It cannot be recovered by retrying.** Anything that re-pushes to machines that are behind — the
|
|
obvious next feature — would retry a permanent failure for ever and make no progress on
|
|
everything else.
|
|
- **The order is not the operator's.** Which module is "first" is an accident of resolution, so
|
|
which nine modules a broken one blocks is unpredictable.
|
|
|
|
## What was there, and what it rested on
|
|
|
|
The behaviour had a test asserting it: *nothing after the failure ran*. Its comment cites
|
|
[ADR 0010](../../02-DECISIONS/0010-delivery.md).
|
|
|
|
**That record does not decide this.** What it says is that a failed *job* stops and names its step
|
|
while a reconciler retries forever, as an argument about pipelines against reconcilers. It says
|
|
nothing about whether one resource failing should prevent the next from being attempted. The
|
|
citation was doing more work than the record supports.
|
|
|
|
## The fix, and the argument that was on the other side
|
|
|
|
**Everything is attempted, and every failure is reported.**
|
|
|
|
The case for stopping is that a resource may depend on an earlier one — a service on the file it
|
|
reads. That is real, and it survives: such a service fails its own check and is reported. This host
|
|
reads back after every write precisely so a thing that did not work is caught rather than assumed,
|
|
so attempting it produces *more* information than skipping it.
|
|
|
|
**What is unchanged:** a declaration that cannot be parsed is still refused whole, and nothing is
|
|
applied. That is a different thing — *this machine could not do it* against *this was never a
|
|
declaration* — and they are fixed in different places.
|
|
|
|
### And one shape still stops what follows, which the first fix got wrong
|
|
|
|
**An action does.** The first version of this fix continued past everything, and the next lab run
|
|
failed at the bootstrap: the store did not answer in three minutes and then said *the database
|
|
system is shutting down*. Carrying on past the readiness gate had started the broker and the
|
|
control plane against a machine that was not ready, and on a small machine that is how a database
|
|
still initialising has its memory taken away.
|
|
|
|
**An action is the only shape whose purpose is to make something true *before* the next thing needs
|
|
it** — which is why it is the only one with a `verify`. The bootstrap is a row of them: the store
|
|
answers, then its databases exist, then their schemas, then the broker. Everything else is
|
|
independent state: a package that will not install has nothing to do with a file on the other side
|
|
of the declaration.
|
|
|
|
So the rule is: **a failed action stops what follows; nothing else does.** Both faults are fixed by
|
|
it, and the report says which happened — *these things failed* and *these things failed and the
|
|
rest was never tried* are different machines.
|
|
|
|
## How this is checked
|
|
|
|
`internal/apply`, three tests: an apply with a resource that cannot succeed still applies the ones
|
|
after it; every failure is counted, not just the first; and a failed **action** stops what follows
|
|
and says so. Each confirmed to fail when its behaviour is removed — including the last, which fails
|
|
if actions stop being treated as gates *and* if everything is treated as one.
|