--- status: accepted date: 2026-06-05 deciders: jochen reconstructed: true --- # 8. A step that fails must fail the job > Reconstructed after the fact from the evidence cited below. ## Context The mesh's expensive faults are not crashes. They are the operations that reported success and did nothing: an artifact that partially downloaded and was extracted anyway, a package that 404ed from every mirror while the job went green, a hook that never ran because it was named for a feature the module does not declare, a deploy that reported the transport succeeded rather than that the effect happened. Each of these was found long after it happened, by someone investigating an unrelated symptom. The cost is not the failure; it is the interval between the failure and anyone learning of it, during which decisions are made on the assumption that the thing worked. ## Considered options 1. **Continue on error and report at the end.** Rejected — it is largely what existed. A summary nobody reads is not a report, and later steps run against the state the failed step should have produced. 2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a precise, located failure into a vague one discovered elsewhere, and requires a health check for every possible partial state. 3. **Fail the step, fail the job, say which step.** Chosen. ## Decision A step that fails stops the sequence it is part of, and the failure is surfaced where the work was requested — not only in a log. Concretely, and these are the forms it takes: - A scripted sequence gates each step on the previous one. A directory change that fails must stop the commands that assumed it. - An artifact that does not fully download is not extracted. - A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must mean the thing is running, not that a command returned. - A template that cannot resolve a variable is not written half-rendered. **Prefer failing to lying.** A green result that is not true costs more than a red one. ## Consequences - Failures are noisier and land earlier, on the person who caused them. - Some jobs that used to complete now stop. In every case examined so far, that job was producing a partial result that something downstream trusted. - This is a rule the mesh has adopted repeatedly rather than once, because each instance is written in a different place — a shell hook, a download path, a deploy stage. It is not enforced by a mechanism, and cannot currently be checked in general. New instances are still being found; the package-install case remains open as [`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md). ## References - `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05. - `A flavor template with an unresolved variable is written to disk instead of failing` (#710), 2026-08-08. - Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`, `troubleshooting/service-started-is-not-ready`, `troubleshooting/green-pipeline-means-transport-not-effect`, `troubleshooting/silent-failures-and-stale-state`. - The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must be loud."