The cause was one line. Machines get a systemd-networkd unit with a static address, so networkd finishes and reports the link configured. The registry ran `ip addr add` inline, which leaves networkd waiting to configure something it was never told about — and systemd-networkd-wait-online has an infinite timeout. So network-online.target was never reached and everything ordered after it never started. On these machines that is Docker, so `docker load` blocked on a socket whose daemon was queued behind a target that would never come, and three bounded timeouts stacked to thirty-five minutes. These machines have no DHCP by design, so that wait was never going to end. The hypothesis in this record was wrong and the record now says so. Stocking had just been changed, so stocking looked guilty; stocking takes 34 seconds and always did, timed directly before changing anything. Fixed with two things that made it cost hours instead of minutes: an image placement now waits for the runtime and refuses after 120s naming what systemd is waiting on, and the end-to-end test passes onProgress — the raise reported every step and the test discarded it, which is why thirty-five minutes and four minutes of silence looked the same. The suite then ran to completion, 23 of 24, the one failure a check of its own flagging a path as a credential because `/` is in the base64 alphabet. Also recorded: a redirected log lags, because Node block-buffers stdout to a file. Read as a stall twice, the second time right after the real fix — where a buffering artifact argues the fix did not work.
04-ISSUES
The front door for "something is wrong" at the level of the mesh's design or governance. Diagnosis happens here, where the whole mesh is in view; the fix lands in the owning code repository.
What belongs here
| Belongs here | Belongs in the knowledge base |
|---|---|
| The design permits a failure to be silent | How to fix one occurrence of it |
| A documented rule is enforced by nothing | A command that works around it |
| A stated invariant is false in practice | A node-specific quirk |
| The owner is unknown and finding it needs the whole mesh in view | Symptom → fix, once the answer is known |
The knowledge base already holds the operational record and is indexed on symptoms. This folder is not a second copy of it. An issue here is a question HQ must answer; an entry there is an incident someone must clear. An issue whose answer is a general lesson belongs in both.
Structure
NNN-short-name/
00-report.md the symptom as observed, with the evidence; status in frontmatter
01-diagnosis.md the investigation trail, dated, including what was ruled out
Frontmatter, on 00-report.md
---
status: open | diagnosing | located | resolved | wontfix
opened: YYYY-MM-DD
located-in: [] # owning repo(s) or module(s), filled by diagnosis
fixed-by: # pull request or commit reference, filled at resolution
amended-design: # design doc path, when the root cause was a design gap
---
Rules
- Anyone may open an issue. No localisation is required to report one.
- The full flow is playbook
00-META/process/03-issues.md. - Closed issues are never deleted — they are the mesh's symptom-to-component memory.
wontfixis legitimate and requires a sentence saying why.