Jochen: a normal application has 3-5 ADRs, maybe 10 for a large one, and we are at 65. Fair, and the cause is mine -- I recorded every FINDING as a decision rather than every fork in the road. Two merges, both cases where one decision had been split across many records because it was taken over several days rather than at once. 0019 absorbs ten records about how this repository works: what it is and that it is public, the folder flow, the two design layers, the issue front door, status in frontmatter, playbooks, the naming rule, the product name. Those were never ten decisions -- they were one, seen from ten angles as the repository took shape. 0016 absorbs the five about the lab: a node is a virtual machine, a router is scenery, a scenario declares the underlay, a scenario is a closed address space, and the two scenario classes. Same pattern -- one design, split by the order it was worked out in. The consolidated 0019 also raises the bar for what earns a record, since that is what produced 65: a record is warranted when there is a genuine fork -- a direction reversed, an alternative that will be proposed again, something contested. A finding is not a decision, and a bug is certainly not. Everything else belongs in the design document where the reasoning is actually read. The checker earned its place here. Deleting nine records left 13 dangling links across the repository and it named every one, including in AGENTS.md. Nothing was found by reading. Remaining clusters worth the same treatment: the host (8 records), delivery (5), modules (6), connectivity (4), substrate and control plane (4). That would be 52 down to roughly 30.
1.5 KiB
1.5 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design |
|---|---|---|---|---|
| open | 2026-08-22 |
001 — A failed package install does not fail the job
Symptom
A module declared a package. The install produced, from every mirror:
error: failed retrieving file … 404
followed by:
-> error installing repo packages
The prepare job then reported success. The package is absent; the pipeline is green.
Why this matters more than one missing package
The first thing the lab work asked the mesh to install demonstrated the exact fault the lab exists to catch — a step that failed, reported success, and left the next step to run against state that was never produced.
It is also a direct violation of a decision already taken and recorded: ADR 0008 says a step that fails must fail the job. That record notes the rule is applied instance by instance and enforced by no mechanism. This is an instance where it was never applied.
Evidence
- Observed 2026-08-22 while declaring the virtualisation package required by ADR 0016.
- A fix is written and open as a pull request, unmerged since 2026-08-20.
Open questions
- Why is the failure swallowed — is the exit status discarded, or never checked?
- Is this specific to package installation, or does the surrounding stage swallow every non-zero exit?
- The fix has been open for two days. What is the review path for a change of this class?