A host change merged yesterday reached no machine without a person copying a
file. The host is not a build target, no declaration delivers it, and the half
that recovers from a bad host — noticing the executable changed, a known-good
record, a launcher that rolls back — is written, tested and called by nothing.
All four machines run a byte-identical hand-copied binary that no package owns
and no record names, so nothing can say a machine is behind.
Found because ADR 0140 needs the machine to report a new fact, and merging that
could not roll it out.
The front door for "something is wrong" at the level of the mesh's design or governance.
Diagnosis happens here, where the whole mesh is in view; the fix lands in the owning code
repository.
What belongs here
Belongs here
Belongs in the knowledge base
The design permits a failure to be silent
How to fix one occurrence of it
A documented rule is enforced by nothing
A command that works around it
A stated invariant is false in practice
A node-specific quirk
The owner is unknown and finding it needs the whole mesh in view
Symptom → fix, once the answer is known
The knowledge base already holds the operational record and is indexed on symptoms. This
folder is not a second copy of it. An issue here is a question HQ must answer; an entry
there is an incident someone must clear. An issue whose answer is a general lesson belongs in
both.
Structure
NNN-short-name/
00-report.md the symptom as observed, with the evidence; status in frontmatter
01-diagnosis.md the investigation trail, dated, including what was ruled out
Frontmatter, on 00-report.md
---status:open | diagnosing | located | resolved | wontfixopened:YYYY-MM-DDlocated-in:[]# owning repo(s) or module(s), filled by diagnosisfixed-by:# pull request or commit reference, filled at resolutionamended-design:# design doc path, when the root cause was a design gap---
Rules
Anyone may open an issue. No localisation is required to report one.