Research 031: open the effort on a core that cannot fail silently

The operator's mandate: races, ignored commands and silence keep recurring.
Classifies 48 recent issues by class, proposes nine checkable principles and
a phased roadmap, for review before anything graduates.
This commit is contained in:
jochen
2026-10-06 02:09:55 +02:00
parent 202f2aa144
commit 97ff655e34
4 changed files with 732 additions and 0 deletions
@@ -0,0 +1,62 @@
---
status: active
initiated: 2026-10-06
touches:
- 00-META/how-we-build.md
- 01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md
- 01-RESEARCH/028-the-meshs-output-channel/00-overview.md
- 01-RESEARCH/019-a-warm-twin-of-the-running-mesh/00-overview.md
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
- 03-DESIGN/01-to-be/06-the-controller.md
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
- 03-DESIGN/01-to-be/25-the-bus-on-nats.md
- 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
- 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md
became: []
---
# 031 — A core that cannot fail silently
**What.** The principles the mesh's core must hold — the controller, the machine host, the bus, the
console and the path a change takes through them — so that it is fully diagnosable, monitors itself,
heals what it knows how to heal, and upgrades itself without a person standing by. And the mechanisms
and the order in which to build them.
**Why.** The operator, 2026-10-06: *"I still notice a lot of race issues, and commands being ignored, or
no feedback, no logs, no monitoring. Our mesh core setup must be fully diagnosable, with active
monitoring, self-healing, self-upgradeable, self-monitoring. The core principles must be very sturdy, no
ambiguities, clear plan of execution, fail-proof setup."*
The record bears it out. In the six days to 2026-10-06, 92 issue reports were opened. Of the 48 read here
as core failures, **every one was noticed because a person or an agent looked**, and **none was raised by
the mesh unasked**. Four faults came back through a different door after their first fix, because each
fix closed an instance and left its class open. One merge was skipped by the bus, and twenty-three over three
days have no matching action; a provider failed for twenty-three hours with only its own
journal saying so; seven databases were dropped on one unreadable file.
**What it touches.** The controller (its verbs, `status`, plans, a lease), the host (its apply and
report), the bus (its advisories and its upgrade), the console, the build path, the output channel of
research 028, and the self-healing intent of research 017, which this effort extends from the loops that
converge modules to the core that runs those loops. 017 deferred heartbeats, conditions and advisories
until the bus was NATS; it is now.
**Documents.**
- [01 — The evidence](01-evidence.md): 48 issues classified by class of failure (races, dropped
commands, no feedback, logs only, two writers, manual repair, self-upgrade, CI-versus-live, third
party), with time to detect and how each was noticed.
- [02 — Principles](02-principles.md): nine, each with what exists, what is missing and **how it is
checked**; and the candidates weighed and not kept.
- [03 — Mechanisms and roadmap](03-mechanisms-and-roadmap.md): conditions, watchdogs from a signals table,
a self-check (`doctor`), the output channel, healers with a hand-act log, a controller lease and report
sequences, staged core upgrades with rollback, a facts snapshot for merge checks, a lab replay of every
incident; five phases, ordered by risk removed; the five largest risks today.
**Next.** The operator reviews the principles and the order of the phases. Graduation would be one
record for the principles (likely amending how-we-build's non-negotiables), and to-be designs for the
condition store with its signals table, and for staged core upgrades. Two measurements are owed before
then: the bounds in the signals table, measured on the live mesh; and a week of the hand-act log, as the
baseline for the self-healing phase.
@@ -0,0 +1,202 @@
# 01 — The evidence, classified by class of failure
Every issue report opened between 2026-09-30 and 2026-10-06 that bears on the mesh's core — the
controller, the machine host, the bus, the console, the build path — read in full and classified by
**the class of failure**, not by the component it was found in. A component view says "fix the host";
a class view says "the same thing is wrong in four places", which is what a principle is for.
## The count
| | |
|---|---|
| issue reports opened 2026-10-01 to 2026-10-06 | **92** (about fifteen a day) |
| of those (and a few from the days before), classified below as core failures | **48** distinct issues |
| classified in more than one class | 17 of 48 |
| a fault that **came back** after a fix of the same symptom | 4 chains: 200 → 265, 230 → 264, 257 → 261 → 267, 175 → 184 → 248 |
| noticed because a person or an agent looked — at a stalled plan, a wrong outcome, a journal, a test run by hand, a review | **48 of 48** |
| of those, the mesh's own answer carried the fact for whoever asked (a refusal, a `status` line, a push's output) | 4 (233, 244, 259, 263) |
| raised by the mesh to anyone, unasked | **0 of 48** |
The recurrences matter most. Each fix was correct for its instance and left the class standing, so the
same symptom came back through a different door days later. That is the measurement behind the
operator's mandate: point fixes are converging on the instances, not on the class.
## The classes
Nine classes, as the mandate frames them. The table under each is the evidence; *detected* is the time
from the fault's start to the moment anyone knew; *noticed by* is how.
### (a) Races: concurrent actors without an ordering
Two actors act on the same thing, and the order of arrival — not an explicit order — decides the outcome.
| Issue | The two actors | Detected | Noticed by |
|---|---|---|---|
| 204 | an outgoing and an incoming controller both sent declarations | 2 min | a person saw a module undone |
| 201 | a plan's push carried a controller digest older than its successor had written | 10 min crash loop | a person, nothing answered |
| 214 | the controller rebuilding itself; the outcome reached the old one or neither | 27 min | a person asked the plan twice |
| 219 | an older build finishing later replaced a newer one | hours | a person reading builds |
| 234 | seven declarations arrived during an eleven-minute apply; the newest was composed from a stale view and undeclared four modules | 8 min of removals | a person, the modules were gone |
| 254 | three plans for three merges ran at once, each asking the same builds | hours | a person, one plan stuck "building" |
| 256 | a first machine's report landed between a module's two sends and read as stale | 7 min | a person |
| 257, 261 | the machine's five-minute reconcile and a delivery took the apply lock in the wrong order | 5 min (257), 29 s visible undo (261) | a person |
| 267 | a reconcile's report queued behind a delivery's apply overtook it at the controller | until a hand push | a person |
| 265 | a push reloaded the bus's permissions while its own answer was still owed | 54 of 103 pushes over two days | a person, "did not answer in time" |
**What they share.** Every one is a receiver that kept *the last thing written* rather than *the
newest thing by an explicit order*. Declarations carry a sequence (issue 107); **reports do not**
(267 says so: "the report does not carry the declaration's sequence number, so the digest decides").
Plans are ordered by when they were made only since ADR 0218. Builds are ordered since issue 219.
Controllers have no epoch, so two instances can both act (204). Ordering was added one message kind at
a time, each after a race in it was seen.
### (b) Commands silently ignored, arguments dropped
| Issue | What was dropped | Effect |
|---|---|---|
| 244 | the console removed `node` from every mesh-seat verb's schema and call | `plan` could only refuse; **`push <one machine>` arrived empty and pushed every machine** |
| 259 | a named push's flush sent every machine a build a policy held back | a fault met on every machine at once, not one |
| 202 | a module whose setting was unset was *left out* of the machine | the resolver vanished from a declaration, no error |
| 188 | a refusal inside "who is on the network" dropped a machine | 40 min, every symptom pointed elsewhere |
| 231 | a misspelled placeholder written to a file as literal text | passes every check |
| 241 | an unreadable contributions file read as "nobody asks" | **seven databases dropped and recreated empty** |
| 255 | the journal verb read nothing and said "-- No entries --" | a refusal that reads as a quiet service |
| 246 | a runtime that answered late was treated as absent | the console said modules "run nowhere" |
**What they share.** A receiver that could not tell *nothing was asked* from *something was lost on the
way*, and chose a default. In 241 and 244 the default was the most destructive reading available.
### (c) Outcomes not fed back to the caller
| Issue | What the caller was told | What happened |
|---|---|---|
| 200, 265 | "did not answer in time" | the push ran; the answer was refused by the bus |
| 176 | the console's build tool neither waits nor registers | — |
| 229 | `plans` answers once in prose; nothing waits for a plan | an agent went round the mesh with `curl` |
| 230, 264 | a host stood aside for its successor and its report was cancelled | the plan waited for ever, reading `late: false` |
| 186 | the build machine dropped 26 of 43 asks; nothing counts asks against outcomes | inferred two hours later |
| 237 | `assign` answered "held" and "does not resolve for lack of it" in one breath | a person or agent would loop |
Since 2026-10-06 the controller answers within ten seconds and keeps every call's outcome under an id
(`calls`, issue 265). Read live the same night: **that log holds the last hundred calls in the
controller's memory**, so a controller restart — which every merge to the controller's own repository
causes — forgets every outcome it held. And `status`, a read-only verb, took **18 seconds** to answer,
twice in a row, so even the health question is answered only through the "still running, ask `calls`"
path.
### (d) Failures visible only as log lines
| Issue | Where it was said | For how long |
|---|---|---|
| 179 (recurred) | the identity provider's journal, every five seconds | **23 hours**, about 31 000 refused logins |
| 184 | the controller's log: slow consumer, heartbeats dropped | 24 min deaf |
| 187 | five faults in one day, each found by reading a container's log hours later | hours each |
| 183, 217, 265 | a `Permissions Violation` line from the bus client library | days |
| 233 | `status` said `refused`, correctly; nothing said it had lasted | 1.5 h |
| 243 | nothing: machines silently ignored lower licence generations | until a login waited three minutes |
| 248 | the controller's event loop stopped logging at 15:17 | hours; "status showed every plan done" |
| 266 | nothing: a merge was skipped by the bus | **23 unmatched merges over three days** |
| 238 | a ban list of 400 entries | the operator's own address banned for four weeks |
`status` printed its all-well sentence through 179, 248 and 266. ADR 0224 made the first of those break
it. The other two have no signal that `status` reads.
### (e) State that two writers own
| Issue | The two writers |
|---|---|
| 190, 222 | the runtime's configuration written by modules that are not the runtime, and by the controller |
| 201 | the controller seat's row written by a successor, read by a predecessor pushed back in |
| 239 | two definitions, in two repositories, held one module name |
| 245 | "behind" answered by a commit comparison beside the plan that already knows |
| 257, 261, 267 | the machine's state written by both the delivery and the five-minute reconcile |
| 250 | a merge announced by the forge's tool and by its poll |
| 179 | the identity provider's admin password: the mesh minted one, the database kept another |
The operator's direction on 245 is the principle in their own words: *"a second answer to the same
question is how the two came to disagree."*
### (f) Manual repair needed
Counted from the reports' own "what unblocked it" sections:
| Repair by hand | Issues | Times |
|---|---|---|
| a push by hand to unstick a plan waiting on a report | 230, 257, 264, 267 | at least 4 |
| a controller restart to recreate a missing object or let go of a held message | 208, 248 | 2 |
| a one-off program run as the controller, outside the service | 201, 248 | 2 |
| the identity provider's admin reset through its bootstrap command | 179 | 2 |
| a kept file restored on a machine by hand | 233 | 1 |
| a stuck plan closed by hand | 214, 254 | 2 |
| a ban lifted by hand | 238 | 1 |
| a consumer remade from now | 248 | 1 |
ADR 0224 (*detected automatically, repaired where safe, loud where not*) is the first rule that turns
one of these into a mechanism. 248's `broker consumer-reset` and 254's `plans close` turned two into
verbs a person runs. Every other row is still a hand on a machine.
### (g) Self-upgrade fragility
The core updates itself: the controller rebuilds and replaces itself, the host delivers its own
successor (ADR 0141), the runtime and the console are modules, and the bus is a module on the control
node.
| Issue | What the self-upgrade broke |
|---|---|
| 201 | the controller pushed back to an older build than its own successor's row |
| 204 | two controllers both sending during a handover |
| 213, 223 | the controller ran as a container, and a new mesh installed it so |
| 214 | the plan that rebuilds the controller lost track of it |
| 230, 264 | the host that stands aside loses the report of the apply that delivered it (fixed twice) |
| 245 | a rebuild of everything replaced the bus's container: **every runtime lost the bus for a minute** |
| 248, 266 | a controller restart is where merges go missing: most of 266's 23 lie in such windows |
| 217 | a refused announcement crash-looped nine runtimes, closing the path that would merge the fix |
A machine's first-in-line rollout (ADR 0218) protects modules. It does not protect the core from itself:
the health a plan waits for is "reported applied", which a controller that cannot plan, a host that
cannot report or a bus that drops messages can each satisfy. Nothing rolls back. The recovery in 201
was the mesh's own binary run by hand from the newer image.
### (h) Checks that pass in CI and fail live
| Issue | The environmental fact the check did not have |
|---|---|
| 177 | the store-backed tests are skipped by the quick check, and the mesh runs none of a module's tests |
| 236 | the host's declaration validation is not run by the catalogue check |
| 262 | musl takes an NXDOMAIN for IPv6 as final; glibc does not |
| 263 | the real machine names make a consumer's identity 23–26 characters against a bound of 20 |
| 202 | the controller's test against the real catalogue, run by nobody until that day |
| 228 | the host's removal has no case for a `user` — found by a review, not a test |
Each check was right about the world it was given. None of them was given the mesh's world: its machine
names, its catalogue, its host's validation, its C libraries.
### (i) Third-party bugs
| Issue | |
|---|---|
| 266 | the bus server's 2.10 release skips messages on a consumer with several filter subjects |
| 265 | a reload of the bus's authorization forgets every reply permission already granted (documented server behaviour, read from its source) |
| 262 | a resolver answering NXDOMAIN where NODATA is correct, met by musl's stricter reading |
The lesson of 266 is not "upgrade the bus" — it is that nothing compared *what was announced* with
*what was acted on*, so a dependency's bug was silent for three days. A defence in depth (watch the
outcome, not the transport) would have caught it whichever layer was wrong.
## What would have prevented or caught each class
| Class | Would have been prevented by | Would have been caught by |
|---|---|---|
| (a) races | every message ordered by its writer, stale refused by every receiver | a lab replay of the interleaving |
| (b) dropped | refusing an unknown or unreadable input by name | a schema walk over every verb |
| (c) no feedback | answer at once with an id; outcome kept durably | a watchdog on "asked and never finished" |
| (d) logs only | — | a condition in `status` and a notification |
| (e) two writers | one writer per piece of state | a registry of writers checked at composition |
| (f) manual repair | a healer for every repair done twice | a counter of hand acts |
| (g) self-upgrade | one machine first, health-gated, rolled back | a lab upgrade with a deliberately broken build |
| (h) CI vs live | checks fed the real mesh's facts | the same, before merge |
| (i) third party | pinning and testing the version that runs | an end-to-end count of announced vs acted |
The two columns are the principles of [02](02-principles.md). Read by count, **the "caught by" column
is the cheapest and widest**: a watchdog and a condition would have shortened most of the 48 from "a
person noticed" to minutes, whatever the cause.
@@ -0,0 +1,247 @@
# 02 — Principles for the core, each with how it is checked
Nine principles. Each is stated as a rule a reviewer can refuse a change against, carries the classes of
[01](01-evidence.md) it answers, says what exists already, and says **how it is checked** — the
repository's own rule ([how-we-build §5](../../00-META/how-we-build.md)): a rule that states no check is
indistinguishable from a wrong one.
They extend, not replace, the six of [research 017](../017-a-mesh-that-heals-itself/01-the-intended-behaviour.md)
(a loop compares with what is; healing is the ordinary path again; a repair never destroys; nothing fails
silently; what cannot be fixed goes to an agent; correctness, not only liveness). Those are about the
loops that converge modules. These are about **the core that runs those loops**: the controller, the
host, the bus, the console and the path a change takes through them. 017 deferred heartbeats, conditions
and advisories until the bus was NATS. It is now, so that deferral has expired.
**The core**, for this document: the controller, the machine host and its launcher, the bus server, the
tool runtime and the console, the build seat, and the forge's announcer of merges. Everything a change
passes through before a module's own code runs.
---
## P1 — One writer per piece of state
Every piece of state the mesh keeps has exactly one writer, named. Anyone else who wants it changed asks
that writer; nobody writes beside it, and nobody computes a second answer to a question it already
answers.
- **Answers:** (e), most of (a). Issues 190, 201, 204, 222, 239, 245, 250, 257/261/267.
- **Exists:** the controller is the only writer of stream definitions (to-be 25); ADR 0222 §3 (the
controller writes no file a seat's holder owns); the collision check at composition (no two modules
declare one path, unit, name or package); the operator's direction on 245.
- **Missing:** a single writer for a *machine's applied state* (the delivery and the five-minute
reconcile both apply and both report — three issues in two days); a single *controller* (two instances
can both act during a handover, 204: nothing holds a lease); a single announcer per event kind (250
was found by counting duplicates by hand).
- **How it is checked:**
1. A **writers table** — state kind, its writer, where it is kept — is part of the to-be design, and a
test in each core repository asserts that the code paths that write each kind are the one named
(by a lint over the store's write calls and the bus subjects each component publishes on, the
latter read from the grants the controller composes: a subject two components may publish on is
refused at composition unless the table says it is shared).
2. **Live:** the controller holds a **lease** (a key-value entry with a revision) and every message it
sends carries the lease's epoch; a host refuses a declaration from an older epoch and says so. A
probe ([03](03-mechanisms-and-roadmap.md), the self-check) asserts one lease holder and no message
from a stale epoch in the last interval.
## P2 — Everything that changes state carries its writer's order, and every receiver refuses what is older
Declarations, reports, plans, builds, calls and announcements each carry `(writer, epoch, sequence)`.
Every receiver keeps the highest it has accepted per writer, and **refuses** — with a line in the mesh's
own words and a counter — anything older. Arrival order never decides.
- **Answers:** (a). Issues 201, 204, 214, 219, 234, 256, 257, 261, 264, 267.
- **Exists:** a declaration carries a sequence and a host keeps the newest (issue 107, to-be 25 §3); a
newer build wins over an older one finishing later (219); a newer merge's plan supersedes an older
one (ADR 0218 §3); a report about a declaration the mesh has moved past no longer replaces the stored
account (267).
- **Missing:** a **report carries no sequence** — the digest last recorded as sent decides, which is why
each of 256, 257 and 267 needed its own rule. No epoch on the controller. A plan's state carries no
revision, so two instances can both advance it (214).
- **How it is checked:**
1. A contract test per message kind, in the receiver's repository: deliver `n`, then `n−1`; the state
names `n` and a refusal is counted. Deliver from epoch `e−1` after `e`: refused. A new message kind
without such a test fails a check that lists every subject the component consumes against the
tests that name it.
2. **Live:** the refusals counter is a signal (P5): zero is normal, a burst is a condition naming the
writer that sent stale.
## P3 — Every command is acknowledged at once, and its outcome is kept where it can be read later
A call is answered within a bound the caller can rely on — with its result, or with an id. Its outcome is
kept **durably**, outlives the process that ran it, and can be read by id or waited on. Nothing is fired
and forgotten, and no answer depends on what the command does to the transport carrying it.
- **Answers:** (c). Issues 176, 186, 200, 229, 230, 237, 264, 265.
- **Exists:** since 265, a seat's call answers within ten seconds or says "running" with an id, `push`
answers before it acts, and `calls` keeps the last hundred calls and their answers; the build path
says "asked, not waited for" and ADR 0219 §3 makes every cancel/kill leave an outcome.
- **Missing:** `calls` lives in the controller's memory, so **a controller restart forgets every
outcome** — and the controller restarts on every merge to its own repository. Nothing waits on a plan
(229). Observed the night this effort began: the read-only `status` took 18 s, so the health question
itself is answered through the "still running" path.
- **How it is checked:**
1. A test that walks every verb the controller announces: each answers within `AnswerWithin`, with a
result or an id (already partly built for 265).
2. A test that restarts the controller between a call and the read of its outcome: the outcome is
still there.
3. **Live:** a probe calls `status` and asserts it answers *in full* within the bound; a call
`running` for longer than its verb's declared bound is a condition (P5).
## P4 — Nothing is dropped silently: an input that is unknown, unreadable or unmet is refused by name
A receiver that cannot read, place or honour an input refuses it and says which, where and why. It never
substitutes a default — above all never "empty" — for "I could not tell". An unknown argument, key,
placeholder, seat or subject is refused, naming it.
- **Answers:** (b). Issues 188, 202, 231, 241, 244, 246, 255, 259.
- **Exists:** the host's parser refuses unknown keys (ADR 0007, 0045); an undeclared setting or endpoint
is refused (ADR 0164, 0138); a verb takes only its declared arguments (ADR 0154) and, since 244, the
controller and the console refuse an undeclared one by name; a seat the mesh does not answer is
refused by name (ADR 0222 §1); ADR 0219 §3, "nothing dropped is silent".
- **Missing:** the rule is in a dozen records and in no principle, so each new reader re-meets it. The
destructive form — **an unreadable input read as "nothing asked", then acted on** (241) — has no
general guard.
- **How it is checked:**
1. Per component, a test that feeds each input reader an unreadable, malformed and foreign input and
asserts a refusal, never an empty result. A reader whose error path returns an empty value fails a
lint that the core repositories run (the shape is mechanical: an error branch that returns the
zero value of a collection).
2. Schema walks for every verb (244's tests) and every placeholder namespace (231).
3. **Destructive deltas are braked**: a reconcile that would withdraw more than a bound of what it
holds (one consumer, one module, a fraction set per provider) stops and raises a condition instead
(P7). Checked by a test that empties the input and asserts nothing is withdrawn.
## P5 — Every expected signal has a watchdog: absence is itself a condition
Every signal the core expects on a cadence or after an act — a heartbeat, a report after a send, a
plan's progress, a build's outcome after its ask, the controller's event loop taking something in, a
provider's standing, an announcement turning into an action — has a declared bound. Silence past the
bound is raised as a condition, naming what was expected, from whom, since when.
- **Answers:** (c), (d), (i). Issues 179, 184, 186, 187, 230, 243, 248, 257, 264, 266, 267.
- **Exists:** a machine's "last heard from — out of touch N m" in `node show`; ADR 0090's stuck machine
(three identical reports); ADR 0224's provider standing, shown with its silence after thirty minutes;
266's catch-up of merges not acted on after ten minutes (the first true *announced-versus-acted*
watchdog); ADR 0162's bound on a plan, which 230 found never fires (`late: false` for ever).
- **Missing:** a list of the signals, their bounds and their owners; the plan bound working; the event
loop's last-taken age (187); asks counted against outcomes (186); the bus's own advisories (slow
consumer, maximum deliveries, permission violations) read as observations instead of log lines.
- **How it is checked:**
1. A **signals table** — signal, emitter, cadence or trigger, bound, condition raised — kept in the
to-be design, and compiled into the controller. A test generated from it suppresses each signal in
turn and asserts the named condition is raised within its bound and cleared when the signal
returns.
2. **Live:** the self-check (P6) reports, for every row, the age of the newest signal, so a row that
never fires is itself visible.
## P6 — The mesh checks itself continuously, against live facts, and says what it found outward
The invariants the design states are probed **against the running mesh** on a schedule, not only in unit
tests. A violation is a **condition** — durable, with since-when, evidence and who can resolve it —
shown in `status` and **sent to the operator** through the output channel. The checker's own heartbeat is
watched from somewhere it does not run.
- **Answers:** (d), (h). Issues 177, 187, 238, 245, 253, 262.
- **Exists:** `status` itself (assembled from reports, as how-we-build §5 requires); 017's *condition*
shape; 028's output seat, researched only; individual live checks done by hand (262: "the live check
stays by hand, on each resolver"); 253's measurement before the collector's first run — the one time
in the window a check ran before the damage.
- **Missing:** a scheduled runner, a condition store, a channel out, and a watcher's watcher.
- **How it is checked:**
1. Every invariant in the to-be design that names a live check is a **probe** in the self-check's
registry; a check over the design documents counts invariants with a stated live probe against
those without, and the number may only go down.
2. The self-check publishes a heartbeat; a second machine's watcher raises "the self-check is silent"
through a channel that does not depend on the control node.
3. **Live, once:** a deliberately broken invariant on a lab mesh appears in `status` and as a
notification within one probe interval.
## P7 — A known failure heals itself, under a brake, and every repair is said
A failure that has been repaired by hand twice is a failure the mesh must repair itself: by running the
ordinary path again (017 P2), never by destroying (017 P3), with a budget and a back-off, and with one
line and one event saying what it did and why. When the budget is spent, or the only repair destroys,
it is a condition and a notification, not a retry.
- **Answers:** (f). Issues 179, 208, 214, 230, 233, 248, 254, 257, 264, 267.
- **Exists:** ADR 0224 §5, *detected automatically, repaired where safe, loud where not*, applied to the
identity provider's admin; the host's launcher rolls back once to known-good and then halts (ADR
0141); the provisioner's `holds` re-provisions what a backend lost (017/02); `broker consumer-reset`
and `plans close` as verbs.
- **Missing:** the general mechanism. The most frequent hand act in the window — **a push by hand to
unstick a plan waiting on a report** — has no healer: the controller could ask the machine to report
again (the machine knows what it applied) before waiting longer.
- **How it is checked:**
1. Every hand act on the core is done through a verb that records it (who, what, why) — the **hand-act
log**. Its count per week is a reported number; an act recorded twice for the same cause is a
condition asking for a healer.
2. Each healer ships with a test that induces its failure, asserts the repair and the event, and
asserts the brake after the budget.
## P8 — The core upgrades itself one machine at a time, health-gated, and rolls back on its own
A new controller, host, runtime or bus reaches one machine first; it is **healthy** only when the
self-check's probes for that component pass there (not merely when it "reported applied"); the rest
follow only then. A component that does not become healthy within its bound is rolled back to the last
known good **by something other than itself**, and the rollback is said. The component being replaced is
never the only witness of its successor's success.
- **Answers:** (g). Issues 201, 204, 213, 214, 217, 230, 245, 248, 264, 266.
- **Exists:** ADR 0218's one machine first, for modules, with "applied and current" as the gate; ADR
0141/0142's side-by-side host versions, known-good and the launcher's single rollback; ADR 0185's
controller serving what it can when it is behind its seat's row; 264's report kept on disk across the
hand-over.
- **Missing:** a health definition per core component; a gate stronger than "reported"; rollback for
the controller, the runtime and the bus; a lease hand-over between controllers (P1); a planned,
rehearsed path for the bus, which is still one process on one machine and whose next upgrade (266:
2.10 → 2.11) is one-way.
- **How it is checked:**
1. **Lab:** a deliberately broken build of each core component (one that starts and does nothing; one
that crashes; one that cannot reach the bus) is merged on a lab mesh. Each is rolled back without a
hand, the mesh ends on the previous build, and a condition and a notification say so.
2. **Live:** every core rollout leaves a record — first machine, health verdict, time to verdict,
rolled back or not — readable through `plans`.
## P9 — A check is fed the real mesh's facts before a change is merged
A check whose verdict depends on the environment — names and their lengths, the machines that exist,
the catalogue as it is, the host's validation, the C library, the server versions — runs against **the
mesh's real facts**, exported and anonymised, before merge. A dependency's version that the mesh runs is
the version its tests run.
- **Answers:** (h), (i). Issues 177, 202, 228, 236, 262, 263, 266.
- **Exists:** 266's `TestTheImageIsTheServerTestedHere` (the bus image's release equals the tested
server's); ADR 0223's composition test that renders the resolver's machine list; 202's test against the real
catalogue (run by hand).
- **Missing:** the export of facts; a merge gate that composes every real machine's declaration with the
change and runs the host's validation over it (which would have refused 236, 263 and 202 in their own
pull requests); a resolver test under musl as well as glibc.
- **How it is checked:**
1. The controller exports a **facts snapshot** (machines, names, assignments, seats, catalogue
commit — no secrets, no addresses) daily; the core repositories' merge check composes every machine
from it with the change applied and runs the host's validation; a pull request that makes any
machine fail to compose or validate fails its check, naming the machine's role and the module.
2. The snapshot's age is a signal (P5).
---
## Candidates weighed and not kept as principles
- **"Every invariant has a live probe, not only a unit test"** — merged into P6; it is how P6 is built.
- **"Environment-dependent checks run against the real mesh's facts"** — kept as P9; the third-party
case (i) folded into it, because pinning and testing the version that runs is the same act.
- **"Self-healing is the default"** — kept as P7 but narrowed to *known* failures, those repaired by
hand twice. A default of healing everything heals what is not understood, which is how a repair
destroys (241's reconcile was, in its own terms, healing).
- **"No loop blocks on long work"** (184, 248, 175) — not a separate principle: a blocked loop is a
signal gone silent (P5, the loop's last-taken age) and a design defect each owner fixes; stating it
as a principle adds a rule with no mesh-wide check.
- **"Clear plan of execution"** from the mandate — not a principle about the mesh; it is the roadmap in
[03](03-mechanisms-and-roadmap.md), and each phase there states its own verification.
## How the principles relate
P2 and P1 **prevent** the races. P4 **prevents** the silent drops. P3, P5 and P6 **catch** whatever the
first three miss, at the cost of minutes, not hours. P7 and P8 **repair**. P9 **moves** the catching
before merge. The order of the roadmap follows from that: catching first, because it is cheapest and
covers every class, including the ones nobody has met yet.
@@ -0,0 +1,221 @@
# 03 — Mechanisms and a phased roadmap
The principles of [02](02-principles.md) need few new things. Most of the parts exist in some form;
what is missing is the connective tissue that makes a fact the mesh already has reach someone without
being asked. This document names the mechanisms, then orders them into phases **by risk removed per unit
of effort**, each phase with deliverables and a verification that says it is done.
Effort is given in **focused working days** of one agent-and-operator pair, at the pace the record shows
(a located issue to a merged fix in under a day is common). The figures are for ordering, not promises.
## The mechanisms
### M1 — Conditions (P5, P6)
017's *condition*, built: a durable fact about something the mesh owns — what is wrong, since when, the
evidence, what was tried, who can resolve it, and whether it may clear itself. Kept in a key-value
bucket the controller writes (one writer, P1), keyed by subject (`plan/<id>`, `machine/<role>`,
`provider/<module>/<consumer>`, `core/<component>`). Raised and cleared by observation only; a person
can **silence** one for a stated time, never resolve it. `status` becomes, first, the list of open
conditions; the all-well sentence is "no open conditions". ADR 0224's provider standing is the first
condition kind and moves into it unchanged.
### M2 — Watchdogs from a signals table (P5)
One table, compiled into the controller, of every signal the core expects. Its first rows, each from an
issue in [01](01-evidence.md):
| Signal | Bound (to be measured, then set) | Condition raised | Issue |
|---|---|---|---|
| machine heartbeat | 3 × interval | machine silent (asleep is a declared state, ADR 0211) | 187 |
| report after a send | the machine's last apply duration × 3, at least 2 min | sent, not reported — then **ask the machine to report again** (M5) | 230, 257, 264, 267 |
| plan tier progress | per tier, from build and apply durations | plan stalled at tier N, waiting on X | 214, 230, 254 |
| controller event loop took something | 2 min while the stream has pending | controller deaf | 184, 248 |
| merge announced → plan made or "nothing reads it" | 10 min (exists, 266) | merge never acted on | 248, 266 |
| build asked → outcome | build's own declared timeout | ask lost | 186 |
| call `running` → finished | the verb's declared bound | call hung | 265 |
| provider standing repeated | 30 min (exists, ADR 0224) | provider silent | 179 |
| bus advisories: slow consumer, maximum deliveries, permission violation | any | bus refused or dropped X for Y | 183, 187, 217, 265 |
| self-check heartbeat | 2 × its interval, watched from a second machine | the watcher is silent | — |
| facts snapshot age | 2 days | merge checks run on stale facts | 263 |
The bus advisories are the cheapest row: the server already publishes them on its system subjects, and
the controller only has to subscribe (read-only) and translate each into the mesh's words, naming the
call or the consumer, as 265 now does for a refused reply.
### M3 — The self-check: `doctor` (P6)
A controller verb, `doctor`, and the same code run every few minutes by the controller itself. Each run
executes the **probe registry** — the live form of the design's invariants — and raises or clears
conditions. First probes, each an invariant that a person checked by hand in the window:
- every machine's declaration composes, and every host would accept it (236, 263);
- every holder of the mesh's resolver answers a machine name for IPv4 and NODATA for IPv6 (262);
- every seat on record has a live holder that answers (208, 218);
- every kept archive is held by a manifest (253 — the controller's collection command already reports it);
- exactly one controller holds the lease (P1);
- every durable consumer's position is near its stream's head (248's replay, 266's skip);
- no address the mesh owns is in a ban list (238);
- `status` answers in full within its bound (P3).
`doctor` with no argument answers the last run's verdict at once (P3) and, with `run`, runs now under an
id. Its own heartbeat is a signal (M2), watched from a second machine.
### M4 — The output channel (P6)
[Research 028](../028-the-meshs-output-channel/00-overview.md)'s seat, built minimally first: **one**
channel the operator chose (028 records it), plus the desktop notifier where the operator is. A condition
is sent when raised, once more if it lasts past a bound, and when it clears. Deduplicated by the
condition's key. The watcher's watcher (028's open question) is the second-machine watchdog of M3,
sending through a channel that does not pass through the control node.
### M5 — Healers (P7)
A healer is a registered response to one condition kind: its repair (the ordinary path again), its
budget, its brake, and the event it emits. First healers, all from hand acts in [01](01-evidence.md) §(f):
| Condition | Repair | Brake |
|---|---|---|
| sent, not reported | ask the machine to report what it last applied (it keeps it since 264); if that names another declaration, send again | twice, then condition |
| plan stalled on a superseded or finished wait | close the plan with its note (`plans close`, done by the mesh) | once |
| seat holder without its worker | raise the seat's objects again (208) | once per holder |
| consumer far behind on a history stream | `broker consumer-reset` (248) — **only** for the consumers the table marks resettable | once, then condition |
| provider admin refuses the mesh's secret | ADR 0224 §5 (exists) | exists |
And the **hand-act log**: every repair a person makes on the core goes through a verb that records who,
what and why. Its weekly count is the measure of P7.
### M6 — Order and epochs (P1, P2)
- The controller takes a **lease** in a key-value bucket before it acts and renews it; its revision is
the **epoch** every declaration and plan write carries. A starting controller waits for the lease;
the outgoing one stops sending when it loses it. That closes 204 and makes 201/214 detectable.
- A **report carries the sequence** of the declaration it is about; the controller keeps the highest
per machine and refuses older accounts by sequence, not by a digest lookup (256, 257, 267 become one
rule).
- **The host has one apply queue.** A delivery and the reconcile are two reasons to enqueue the same
act; the queue applies the newest declaration once, and makes one report (257/261/267 become
impossible rather than handled).
- Every receiver's refusal of something stale is a counted line (P2's live check).
### M7 — Staged, reversible core upgrades (P8)
- **A health definition per core component**, written as probes in M3's registry: the controller
answers `status` in bound and holds the lease; the host has reported its current declaration; the
runtime has announced and answers a PING; the bus has every stream and every durable consumer and
passes a request/reply round trip.
- **The gate:** ADR 0218's first machine is judged by those probes, not by "applied".
- **Rollback by a witness that is not the new build:** the host's launcher for the host (exists); for
the controller, the previous controller's process kept installed beside it, re-started by the host
when the new one does not take the lease in bound; for the runtime, the host's known-good the same
way. Each rollback is a condition, so it is said.
- **The bus is planned, not rolled.** A bus upgrade is a declared maintenance step: streams snapshotted,
the step announced as a condition while it runs, every consumer's position checked after (M3). The
one-way 2.10 → 2.11 upgrade of 266 is the first. Whether the bus should become a cluster of three so
that it can be upgraded live is a question for its own effort.
### M8 — The facts snapshot and the merge gate (P9)
The controller exports a facts snapshot (machines by role and name length, assignments, seats, catalogue
commit) to a place the build seat reads. The core repositories' and the catalogue's merge checks compose
every machine with the change and run the host's validation. The resolver module's tests run under musl
and glibc. A bus, store or library version the mesh runs is the one its tests run (266's pattern,
generalised).
### M9 — The lab replay (all)
[Research 019](../019-a-warm-twin-of-the-running-mesh/00-overview.md)'s warm twin, or a throwaway lab of
three containers, running **scripted replays of each incident** in the window: a reconcile due during a
push; a host self-update during its report; a bus reload during a call; a consumer with several filters
under mixed traffic; a missing consumer made with the server's default; an unreadable contributions
file; a controller rebuilding itself mid-plan; two controllers at once. Each replay asserts the
principle's outcome (refused stale, condition raised, healed, rolled back). They run on every merge to a
core repository.
---
## The roadmap
Ordered by **risk removed per day**. Detection comes first because it covers every class at once,
including the ones not met yet; prevention second; repair and staging third; the pre-merge and lab work
last because they are larger and pay off over months.
### Phase 0 — Finish what is in flight (2–3 days)
- Land and roll out the located fixes: 244, 264, 265, 266 (including the bus's planned 2.11 upgrade, done
as M7's first planned bus step), 267, 257/261.
- Make `calls` durable (a key-value bucket, bounded by count and age), and bring `status` inside its own
answer bound.
- Start the **hand-act log** now, before anything else, so every later phase is measured against a
baseline.
- **Done when:** the four located core issues resolve with their live checks; a controller restart
keeps `calls`; `status` answers in full within ten seconds.
### Phase 1 — The mesh says when it is wrong (5–8 days)
- M1 conditions (ADR 0224's standing moved into them); M2 watchdogs for the first eight rows; the bus
advisories subscribed and translated; M3 `doctor` with the first probes; M4 with one channel and the
second-machine watcher.
- **Done when:** on a lab mesh, suppressing each signal in the table raises its condition within its
bound and sends a notification; clearing it clears both. Live: a week of conditions read back, every
one either real or a bound corrected.
- **Risk removed:** every class in [01](01-evidence.md) moves from "noticed by a person, hours later" to
"said by the mesh, minutes later".
### Phase 2 — Order and one writer (5–7 days)
- M6: the controller's lease and epoch; a report's sequence; the host's single apply queue; stale
refusals counted.
- The writers table and the signals table written into a to-be design, with their compile-time checks.
- P4's lint for empty-on-error readers in the core repositories, and the destructive-delta brake in the
provisioner harness.
- **Done when:** the lab replays of 204, 257/261/267 and 241 end with "refused stale", "one report" and
"withdrawal braked" respectively; the contract test for every consumed subject exists.
- **Risk removed:** class (a), the largest by count, and the destructive half of (b).
### Phase 3 — Healers (3–5 days)
- M5's first healers and the rule that a repair done by hand twice asks for one (from the hand-act log).
- **Done when:** a lab mesh recovers from each induced failure in the healers table with no hand, says so,
and brakes after its budget. Live: a week in which the hand-act log has no repeat.
### Phase 4 — Core upgrades that roll back (8–12 days)
- M7: health definitions; the gate; rollback for the controller and the runtime; the bus as a planned
step.
- **Done when:** on a lab mesh, a deliberately broken build of the controller, the host and the runtime
is each rolled back without a hand, the mesh ends on the previous build, and the rollback is a
condition and a notification. Live: the next three core rollouts each record a health verdict.
- **Risk removed:** class (g) — the failures that take the control path itself down.
### Phase 5 — Checks before merge, and the replay suite (8–12 days, then ongoing)
- M8: the facts snapshot and the compose-and-validate merge gate; the libc matrix; versions tested as
run.
- M9: every incident in the window as a scripted replay, run on every core merge; each new core issue
adds its replay as its "how it is checked".
- **Done when:** the replays of 236, 262, 263 and 266 fail on the commit before their fix and pass after;
a new core issue cannot resolve without a replay or a stated reason why none is possible.
**Total, roughly six to eight weeks of focused work**, with the first visible change — the mesh saying
when it is wrong — inside the first two.
## The five largest risks today
Ranked by likelihood × damage, from the window's evidence and the mesh as read the night this effort
began:
1. **A stall nobody is told about.** A plan, a merge or a report that stops is noticed only by someone
looking (230, 248, 257, 264, 266, 267; issue 187 still open). Every release goes through a plan.
2. **A destructive act on a misread input.** 241 dropped seven databases on one unreadable file; 234
undeclared four modules from a stale composition; 245's advice rebuilt everything and cut the bus;
253's collector would have deleted every archive. There is no general brake on a large withdrawal.
3. **The core replacing itself with nothing to roll it back.** A controller build that starts but cannot
plan (201, 214) halts every later change, because the controller is what plans the fix. Only the
host has a launcher rollback.
4. **The bus as a single, un-upgradeable process.** It runs a release that skips messages (266), its
reload drops owed replies (265), replacing it cuts every machine off (245), and its next upgrade is
one-way and needs a restart.
5. **Concurrent actors with no lease or report order.** Several agent sessions, plans and reconciles act
at once; on the night this began, four named pushes, one per machine, started within two seconds. The digests
catch most of it now, one rule per message kind; the next message kind will not have one.