Files
hq/00-META/process/03-issues.md
T
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00

3.1 KiB

Playbook 03 — Issues

Trigger. Something is wrong at the level of the mesh's design or governance — a rule that turns out to be unenforced, a stated behaviour that does not happen, a silent failure the design permits.

Who runs it. Anyone may open an issue. No localisation is required to report one.

What belongs here, and what does not

Belongs in 04-ISSUES Belongs in the knowledge base
The design permits a failure to be silent How to fix one occurrence of it
A documented rule is enforced by nothing A command that works around it
A stated invariant is false in practice A node-specific quirk
The owning component is unknown and finding it needs the whole mesh in view Symptom → fix, once the answer is known

The knowledge base already holds the operational record and is indexed on symptoms. This folder is not a second copy of it. An issue here is a question HQ must answer, not an incident someone must clear.

Steps

  1. Take the next free number — across main and every open pull request, not main alone. Work sits on unmerged branches for days, so two people both reading main allocate the same number; it happened twice in one hour between two machines, and the second collision reached main with every check passing (issue 155). cycle.py now refuses two records sharing a number, which catches a collision but does not prevent one. Create 04-ISSUES/NNN-short-name/00-report.md:

    ---
    status: open
    opened: YYYY-MM-DD
    located-in: []       # owning repo(s)/module(s), filled by diagnosis
    fixed-by:            # PR or commit reference, filled at resolution
    amended-design:      # design doc path, when the root cause was a design gap
    ---
    

    Then the symptom as observed, in plain terms, with the evidence that it happened.

  2. Investigate in 01-diagnosis.md in the same folder — the trail, dated, including what was ruled out. Move status: to diagnosing, then located once the owner is known.

  3. Resolve. Set status: resolved, fill fixed-by:, and if the root cause was a design gap, run playbook 02 and fill amended-design:.

Rules

  • Closed issues are never deleted — they are the mesh's symptom-to-component memory.
  • fixed-by: names something that will still exist: a commit or a pull request, never a branch. A branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by the time anybody follows it.
  • A fix that turns out to have broken something else is written back into the record that asked for it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it is must not have to already know there was a sequel.
  • Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose author is still pushing only moves the race.
  • An issue whose answer is a general lesson should also be written to the knowledge base, so the next person searching a symptom finds it. Both, not either.
  • status: wontfix is legitimate and requires a sentence saying why.