From 97ff655e3442a363e51f437b975d5dfaed4aa555 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 6 Oct 2026 02:09:55 +0200 Subject: [PATCH] Research 031: open the effort on a core that cannot fail silently The operator's mandate: races, ignored commands and silence keep recurring. Classifies 48 recent issues by class, proposes nine checkable principles and a phased roadmap, for review before anything graduates. --- .../00-overview.md | 62 +++++ .../01-evidence.md | 202 ++++++++++++++ .../02-principles.md | 247 ++++++++++++++++++ .../03-mechanisms-and-roadmap.md | 221 ++++++++++++++++ 4 files changed, 732 insertions(+) create mode 100644 01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md create mode 100644 01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md create mode 100644 01-RESEARCH/031-a-core-that-cannot-fail-silently/02-principles.md create mode 100644 01-RESEARCH/031-a-core-that-cannot-fail-silently/03-mechanisms-and-roadmap.md diff --git a/01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md b/01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md new file mode 100644 index 0000000..5bd3977 --- /dev/null +++ b/01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md @@ -0,0 +1,62 @@ +--- +status: active +initiated: 2026-10-06 +touches: + - 00-META/how-we-build.md + - 01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md + - 01-RESEARCH/028-the-meshs-output-channel/00-overview.md + - 01-RESEARCH/019-a-warm-twin-of-the-running-mesh/00-overview.md + - 02-DECISIONS/0141-the-host-delivers-its-own-successor.md + - 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md + - 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md + - 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md + - 03-DESIGN/01-to-be/06-the-controller.md + - 03-DESIGN/01-to-be/09-the-node-lifecycle.md + - 03-DESIGN/01-to-be/25-the-bus-on-nats.md + - 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md + - 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md +became: [] +--- + +# 031 — A core that cannot fail silently + +**What.** The principles the mesh's core must hold — the controller, the machine host, the bus, the +console and the path a change takes through them — so that it is fully diagnosable, monitors itself, +heals what it knows how to heal, and upgrades itself without a person standing by. And the mechanisms +and the order in which to build them. + +**Why.** The operator, 2026-10-06: *"I still notice a lot of race issues, and commands being ignored, or +no feedback, no logs, no monitoring. Our mesh core setup must be fully diagnosable, with active +monitoring, self-healing, self-upgradeable, self-monitoring. The core principles must be very sturdy, no +ambiguities, clear plan of execution, fail-proof setup."* + +The record bears it out. In the six days to 2026-10-06, 92 issue reports were opened. Of the 48 read here +as core failures, **every one was noticed because a person or an agent looked**, and **none was raised by +the mesh unasked**. Four faults came back through a different door after their first fix, because each +fix closed an instance and left its class open. One merge was skipped by the bus, and twenty-three over three +days have no matching action; a provider failed for twenty-three hours with only its own +journal saying so; seven databases were dropped on one unreadable file. + +**What it touches.** The controller (its verbs, `status`, plans, a lease), the host (its apply and +report), the bus (its advisories and its upgrade), the console, the build path, the output channel of +research 028, and the self-healing intent of research 017, which this effort extends from the loops that +converge modules to the core that runs those loops. 017 deferred heartbeats, conditions and advisories +until the bus was NATS; it is now. + +**Documents.** + +- [01 — The evidence](01-evidence.md): 48 issues classified by class of failure (races, dropped + commands, no feedback, logs only, two writers, manual repair, self-upgrade, CI-versus-live, third + party), with time to detect and how each was noticed. +- [02 — Principles](02-principles.md): nine, each with what exists, what is missing and **how it is + checked**; and the candidates weighed and not kept. +- [03 — Mechanisms and roadmap](03-mechanisms-and-roadmap.md): conditions, watchdogs from a signals table, + a self-check (`doctor`), the output channel, healers with a hand-act log, a controller lease and report + sequences, staged core upgrades with rollback, a facts snapshot for merge checks, a lab replay of every + incident; five phases, ordered by risk removed; the five largest risks today. + +**Next.** The operator reviews the principles and the order of the phases. Graduation would be one +record for the principles (likely amending how-we-build's non-negotiables), and to-be designs for the +condition store with its signals table, and for staged core upgrades. Two measurements are owed before +then: the bounds in the signals table, measured on the live mesh; and a week of the hand-act log, as the +baseline for the self-healing phase. diff --git a/01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md b/01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md new file mode 100644 index 0000000..bdf09e6 --- /dev/null +++ b/01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md @@ -0,0 +1,202 @@ +# 01 — The evidence, classified by class of failure + +Every issue report opened between 2026-09-30 and 2026-10-06 that bears on the mesh's core — the +controller, the machine host, the bus, the console, the build path — read in full and classified by +**the class of failure**, not by the component it was found in. A component view says "fix the host"; +a class view says "the same thing is wrong in four places", which is what a principle is for. + +## The count + +| | | +|---|---| +| issue reports opened 2026-10-01 to 2026-10-06 | **92** (about fifteen a day) | +| of those (and a few from the days before), classified below as core failures | **48** distinct issues | +| classified in more than one class | 17 of 48 | +| a fault that **came back** after a fix of the same symptom | 4 chains: 200 → 265, 230 → 264, 257 → 261 → 267, 175 → 184 → 248 | +| noticed because a person or an agent looked — at a stalled plan, a wrong outcome, a journal, a test run by hand, a review | **48 of 48** | +| of those, the mesh's own answer carried the fact for whoever asked (a refusal, a `status` line, a push's output) | 4 (233, 244, 259, 263) | +| raised by the mesh to anyone, unasked | **0 of 48** | + +The recurrences matter most. Each fix was correct for its instance and left the class standing, so the +same symptom came back through a different door days later. That is the measurement behind the +operator's mandate: point fixes are converging on the instances, not on the class. + +## The classes + +Nine classes, as the mandate frames them. The table under each is the evidence; *detected* is the time +from the fault's start to the moment anyone knew; *noticed by* is how. + +### (a) Races: concurrent actors without an ordering + +Two actors act on the same thing, and the order of arrival — not an explicit order — decides the outcome. + +| Issue | The two actors | Detected | Noticed by | +|---|---|---|---| +| 204 | an outgoing and an incoming controller both sent declarations | 2 min | a person saw a module undone | +| 201 | a plan's push carried a controller digest older than its successor had written | 10 min crash loop | a person, nothing answered | +| 214 | the controller rebuilding itself; the outcome reached the old one or neither | 27 min | a person asked the plan twice | +| 219 | an older build finishing later replaced a newer one | hours | a person reading builds | +| 234 | seven declarations arrived during an eleven-minute apply; the newest was composed from a stale view and undeclared four modules | 8 min of removals | a person, the modules were gone | +| 254 | three plans for three merges ran at once, each asking the same builds | hours | a person, one plan stuck "building" | +| 256 | a first machine's report landed between a module's two sends and read as stale | 7 min | a person | +| 257, 261 | the machine's five-minute reconcile and a delivery took the apply lock in the wrong order | 5 min (257), 29 s visible undo (261) | a person | +| 267 | a reconcile's report queued behind a delivery's apply overtook it at the controller | until a hand push | a person | +| 265 | a push reloaded the bus's permissions while its own answer was still owed | 54 of 103 pushes over two days | a person, "did not answer in time" | + +**What they share.** Every one is a receiver that kept *the last thing written* rather than *the +newest thing by an explicit order*. Declarations carry a sequence (issue 107); **reports do not** +(267 says so: "the report does not carry the declaration's sequence number, so the digest decides"). +Plans are ordered by when they were made only since ADR 0218. Builds are ordered since issue 219. +Controllers have no epoch, so two instances can both act (204). Ordering was added one message kind at +a time, each after a race in it was seen. + +### (b) Commands silently ignored, arguments dropped + +| Issue | What was dropped | Effect | +|---|---|---| +| 244 | the console removed `node` from every mesh-seat verb's schema and call | `plan` could only refuse; **`push ` arrived empty and pushed every machine** | +| 259 | a named push's flush sent every machine a build a policy held back | a fault met on every machine at once, not one | +| 202 | a module whose setting was unset was *left out* of the machine | the resolver vanished from a declaration, no error | +| 188 | a refusal inside "who is on the network" dropped a machine | 40 min, every symptom pointed elsewhere | +| 231 | a misspelled placeholder written to a file as literal text | passes every check | +| 241 | an unreadable contributions file read as "nobody asks" | **seven databases dropped and recreated empty** | +| 255 | the journal verb read nothing and said "-- No entries --" | a refusal that reads as a quiet service | +| 246 | a runtime that answered late was treated as absent | the console said modules "run nowhere" | + +**What they share.** A receiver that could not tell *nothing was asked* from *something was lost on the +way*, and chose a default. In 241 and 244 the default was the most destructive reading available. + +### (c) Outcomes not fed back to the caller + +| Issue | What the caller was told | What happened | +|---|---|---| +| 200, 265 | "did not answer in time" | the push ran; the answer was refused by the bus | +| 176 | the console's build tool neither waits nor registers | — | +| 229 | `plans` answers once in prose; nothing waits for a plan | an agent went round the mesh with `curl` | +| 230, 264 | a host stood aside for its successor and its report was cancelled | the plan waited for ever, reading `late: false` | +| 186 | the build machine dropped 26 of 43 asks; nothing counts asks against outcomes | inferred two hours later | +| 237 | `assign` answered "held" and "does not resolve for lack of it" in one breath | a person or agent would loop | + +Since 2026-10-06 the controller answers within ten seconds and keeps every call's outcome under an id +(`calls`, issue 265). Read live the same night: **that log holds the last hundred calls in the +controller's memory**, so a controller restart — which every merge to the controller's own repository +causes — forgets every outcome it held. And `status`, a read-only verb, took **18 seconds** to answer, +twice in a row, so even the health question is answered only through the "still running, ask `calls`" +path. + +### (d) Failures visible only as log lines + +| Issue | Where it was said | For how long | +|---|---|---| +| 179 (recurred) | the identity provider's journal, every five seconds | **23 hours**, about 31 000 refused logins | +| 184 | the controller's log: slow consumer, heartbeats dropped | 24 min deaf | +| 187 | five faults in one day, each found by reading a container's log hours later | hours each | +| 183, 217, 265 | a `Permissions Violation` line from the bus client library | days | +| 233 | `status` said `refused`, correctly; nothing said it had lasted | 1.5 h | +| 243 | nothing: machines silently ignored lower licence generations | until a login waited three minutes | +| 248 | the controller's event loop stopped logging at 15:17 | hours; "status showed every plan done" | +| 266 | nothing: a merge was skipped by the bus | **23 unmatched merges over three days** | +| 238 | a ban list of 400 entries | the operator's own address banned for four weeks | + +`status` printed its all-well sentence through 179, 248 and 266. ADR 0224 made the first of those break +it. The other two have no signal that `status` reads. + +### (e) State that two writers own + +| Issue | The two writers | +|---|---| +| 190, 222 | the runtime's configuration written by modules that are not the runtime, and by the controller | +| 201 | the controller seat's row written by a successor, read by a predecessor pushed back in | +| 239 | two definitions, in two repositories, held one module name | +| 245 | "behind" answered by a commit comparison beside the plan that already knows | +| 257, 261, 267 | the machine's state written by both the delivery and the five-minute reconcile | +| 250 | a merge announced by the forge's tool and by its poll | +| 179 | the identity provider's admin password: the mesh minted one, the database kept another | + +The operator's direction on 245 is the principle in their own words: *"a second answer to the same +question is how the two came to disagree."* + +### (f) Manual repair needed + +Counted from the reports' own "what unblocked it" sections: + +| Repair by hand | Issues | Times | +|---|---|---| +| a push by hand to unstick a plan waiting on a report | 230, 257, 264, 267 | at least 4 | +| a controller restart to recreate a missing object or let go of a held message | 208, 248 | 2 | +| a one-off program run as the controller, outside the service | 201, 248 | 2 | +| the identity provider's admin reset through its bootstrap command | 179 | 2 | +| a kept file restored on a machine by hand | 233 | 1 | +| a stuck plan closed by hand | 214, 254 | 2 | +| a ban lifted by hand | 238 | 1 | +| a consumer remade from now | 248 | 1 | + +ADR 0224 (*detected automatically, repaired where safe, loud where not*) is the first rule that turns +one of these into a mechanism. 248's `broker consumer-reset` and 254's `plans close` turned two into +verbs a person runs. Every other row is still a hand on a machine. + +### (g) Self-upgrade fragility + +The core updates itself: the controller rebuilds and replaces itself, the host delivers its own +successor (ADR 0141), the runtime and the console are modules, and the bus is a module on the control +node. + +| Issue | What the self-upgrade broke | +|---|---| +| 201 | the controller pushed back to an older build than its own successor's row | +| 204 | two controllers both sending during a handover | +| 213, 223 | the controller ran as a container, and a new mesh installed it so | +| 214 | the plan that rebuilds the controller lost track of it | +| 230, 264 | the host that stands aside loses the report of the apply that delivered it (fixed twice) | +| 245 | a rebuild of everything replaced the bus's container: **every runtime lost the bus for a minute** | +| 248, 266 | a controller restart is where merges go missing: most of 266's 23 lie in such windows | +| 217 | a refused announcement crash-looped nine runtimes, closing the path that would merge the fix | + +A machine's first-in-line rollout (ADR 0218) protects modules. It does not protect the core from itself: +the health a plan waits for is "reported applied", which a controller that cannot plan, a host that +cannot report or a bus that drops messages can each satisfy. Nothing rolls back. The recovery in 201 +was the mesh's own binary run by hand from the newer image. + +### (h) Checks that pass in CI and fail live + +| Issue | The environmental fact the check did not have | +|---|---| +| 177 | the store-backed tests are skipped by the quick check, and the mesh runs none of a module's tests | +| 236 | the host's declaration validation is not run by the catalogue check | +| 262 | musl takes an NXDOMAIN for IPv6 as final; glibc does not | +| 263 | the real machine names make a consumer's identity 23–26 characters against a bound of 20 | +| 202 | the controller's test against the real catalogue, run by nobody until that day | +| 228 | the host's removal has no case for a `user` — found by a review, not a test | + +Each check was right about the world it was given. None of them was given the mesh's world: its machine +names, its catalogue, its host's validation, its C libraries. + +### (i) Third-party bugs + +| Issue | | +|---|---| +| 266 | the bus server's 2.10 release skips messages on a consumer with several filter subjects | +| 265 | a reload of the bus's authorization forgets every reply permission already granted (documented server behaviour, read from its source) | +| 262 | a resolver answering NXDOMAIN where NODATA is correct, met by musl's stricter reading | + +The lesson of 266 is not "upgrade the bus" — it is that nothing compared *what was announced* with +*what was acted on*, so a dependency's bug was silent for three days. A defence in depth (watch the +outcome, not the transport) would have caught it whichever layer was wrong. + +## What would have prevented or caught each class + +| Class | Would have been prevented by | Would have been caught by | +|---|---|---| +| (a) races | every message ordered by its writer, stale refused by every receiver | a lab replay of the interleaving | +| (b) dropped | refusing an unknown or unreadable input by name | a schema walk over every verb | +| (c) no feedback | answer at once with an id; outcome kept durably | a watchdog on "asked and never finished" | +| (d) logs only | — | a condition in `status` and a notification | +| (e) two writers | one writer per piece of state | a registry of writers checked at composition | +| (f) manual repair | a healer for every repair done twice | a counter of hand acts | +| (g) self-upgrade | one machine first, health-gated, rolled back | a lab upgrade with a deliberately broken build | +| (h) CI vs live | checks fed the real mesh's facts | the same, before merge | +| (i) third party | pinning and testing the version that runs | an end-to-end count of announced vs acted | + +The two columns are the principles of [02](02-principles.md). Read by count, **the "caught by" column +is the cheapest and widest**: a watchdog and a condition would have shortened most of the 48 from "a +person noticed" to minutes, whatever the cause. diff --git a/01-RESEARCH/031-a-core-that-cannot-fail-silently/02-principles.md b/01-RESEARCH/031-a-core-that-cannot-fail-silently/02-principles.md new file mode 100644 index 0000000..92e6613 --- /dev/null +++ b/01-RESEARCH/031-a-core-that-cannot-fail-silently/02-principles.md @@ -0,0 +1,247 @@ +# 02 — Principles for the core, each with how it is checked + +Nine principles. Each is stated as a rule a reviewer can refuse a change against, carries the classes of +[01](01-evidence.md) it answers, says what exists already, and says **how it is checked** — the +repository's own rule ([how-we-build §5](../../00-META/how-we-build.md)): a rule that states no check is +indistinguishable from a wrong one. + +They extend, not replace, the six of [research 017](../017-a-mesh-that-heals-itself/01-the-intended-behaviour.md) +(a loop compares with what is; healing is the ordinary path again; a repair never destroys; nothing fails +silently; what cannot be fixed goes to an agent; correctness, not only liveness). Those are about the +loops that converge modules. These are about **the core that runs those loops**: the controller, the +host, the bus, the console and the path a change takes through them. 017 deferred heartbeats, conditions +and advisories until the bus was NATS. It is now, so that deferral has expired. + +**The core**, for this document: the controller, the machine host and its launcher, the bus server, the +tool runtime and the console, the build seat, and the forge's announcer of merges. Everything a change +passes through before a module's own code runs. + +--- + +## P1 — One writer per piece of state + +Every piece of state the mesh keeps has exactly one writer, named. Anyone else who wants it changed asks +that writer; nobody writes beside it, and nobody computes a second answer to a question it already +answers. + +- **Answers:** (e), most of (a). Issues 190, 201, 204, 222, 239, 245, 250, 257/261/267. +- **Exists:** the controller is the only writer of stream definitions (to-be 25); ADR 0222 §3 (the + controller writes no file a seat's holder owns); the collision check at composition (no two modules + declare one path, unit, name or package); the operator's direction on 245. +- **Missing:** a single writer for a *machine's applied state* (the delivery and the five-minute + reconcile both apply and both report — three issues in two days); a single *controller* (two instances + can both act during a handover, 204: nothing holds a lease); a single announcer per event kind (250 + was found by counting duplicates by hand). +- **How it is checked:** + 1. A **writers table** — state kind, its writer, where it is kept — is part of the to-be design, and a + test in each core repository asserts that the code paths that write each kind are the one named + (by a lint over the store's write calls and the bus subjects each component publishes on, the + latter read from the grants the controller composes: a subject two components may publish on is + refused at composition unless the table says it is shared). + 2. **Live:** the controller holds a **lease** (a key-value entry with a revision) and every message it + sends carries the lease's epoch; a host refuses a declaration from an older epoch and says so. A + probe ([03](03-mechanisms-and-roadmap.md), the self-check) asserts one lease holder and no message + from a stale epoch in the last interval. + +## P2 — Everything that changes state carries its writer's order, and every receiver refuses what is older + +Declarations, reports, plans, builds, calls and announcements each carry `(writer, epoch, sequence)`. +Every receiver keeps the highest it has accepted per writer, and **refuses** — with a line in the mesh's +own words and a counter — anything older. Arrival order never decides. + +- **Answers:** (a). Issues 201, 204, 214, 219, 234, 256, 257, 261, 264, 267. +- **Exists:** a declaration carries a sequence and a host keeps the newest (issue 107, to-be 25 §3); a + newer build wins over an older one finishing later (219); a newer merge's plan supersedes an older + one (ADR 0218 §3); a report about a declaration the mesh has moved past no longer replaces the stored + account (267). +- **Missing:** a **report carries no sequence** — the digest last recorded as sent decides, which is why + each of 256, 257 and 267 needed its own rule. No epoch on the controller. A plan's state carries no + revision, so two instances can both advance it (214). +- **How it is checked:** + 1. A contract test per message kind, in the receiver's repository: deliver `n`, then `n−1`; the state + names `n` and a refusal is counted. Deliver from epoch `e−1` after `e`: refused. A new message kind + without such a test fails a check that lists every subject the component consumes against the + tests that name it. + 2. **Live:** the refusals counter is a signal (P5): zero is normal, a burst is a condition naming the + writer that sent stale. + +## P3 — Every command is acknowledged at once, and its outcome is kept where it can be read later + +A call is answered within a bound the caller can rely on — with its result, or with an id. Its outcome is +kept **durably**, outlives the process that ran it, and can be read by id or waited on. Nothing is fired +and forgotten, and no answer depends on what the command does to the transport carrying it. + +- **Answers:** (c). Issues 176, 186, 200, 229, 230, 237, 264, 265. +- **Exists:** since 265, a seat's call answers within ten seconds or says "running" with an id, `push` + answers before it acts, and `calls` keeps the last hundred calls and their answers; the build path + says "asked, not waited for" and ADR 0219 §3 makes every cancel/kill leave an outcome. +- **Missing:** `calls` lives in the controller's memory, so **a controller restart forgets every + outcome** — and the controller restarts on every merge to its own repository. Nothing waits on a plan + (229). Observed the night this effort began: the read-only `status` took 18 s, so the health question + itself is answered through the "still running" path. +- **How it is checked:** + 1. A test that walks every verb the controller announces: each answers within `AnswerWithin`, with a + result or an id (already partly built for 265). + 2. A test that restarts the controller between a call and the read of its outcome: the outcome is + still there. + 3. **Live:** a probe calls `status` and asserts it answers *in full* within the bound; a call + `running` for longer than its verb's declared bound is a condition (P5). + +## P4 — Nothing is dropped silently: an input that is unknown, unreadable or unmet is refused by name + +A receiver that cannot read, place or honour an input refuses it and says which, where and why. It never +substitutes a default — above all never "empty" — for "I could not tell". An unknown argument, key, +placeholder, seat or subject is refused, naming it. + +- **Answers:** (b). Issues 188, 202, 231, 241, 244, 246, 255, 259. +- **Exists:** the host's parser refuses unknown keys (ADR 0007, 0045); an undeclared setting or endpoint + is refused (ADR 0164, 0138); a verb takes only its declared arguments (ADR 0154) and, since 244, the + controller and the console refuse an undeclared one by name; a seat the mesh does not answer is + refused by name (ADR 0222 §1); ADR 0219 §3, "nothing dropped is silent". +- **Missing:** the rule is in a dozen records and in no principle, so each new reader re-meets it. The + destructive form — **an unreadable input read as "nothing asked", then acted on** (241) — has no + general guard. +- **How it is checked:** + 1. Per component, a test that feeds each input reader an unreadable, malformed and foreign input and + asserts a refusal, never an empty result. A reader whose error path returns an empty value fails a + lint that the core repositories run (the shape is mechanical: an error branch that returns the + zero value of a collection). + 2. Schema walks for every verb (244's tests) and every placeholder namespace (231). + 3. **Destructive deltas are braked**: a reconcile that would withdraw more than a bound of what it + holds (one consumer, one module, a fraction set per provider) stops and raises a condition instead + (P7). Checked by a test that empties the input and asserts nothing is withdrawn. + +## P5 — Every expected signal has a watchdog: absence is itself a condition + +Every signal the core expects on a cadence or after an act — a heartbeat, a report after a send, a +plan's progress, a build's outcome after its ask, the controller's event loop taking something in, a +provider's standing, an announcement turning into an action — has a declared bound. Silence past the +bound is raised as a condition, naming what was expected, from whom, since when. + +- **Answers:** (c), (d), (i). Issues 179, 184, 186, 187, 230, 243, 248, 257, 264, 266, 267. +- **Exists:** a machine's "last heard from — out of touch N m" in `node show`; ADR 0090's stuck machine + (three identical reports); ADR 0224's provider standing, shown with its silence after thirty minutes; + 266's catch-up of merges not acted on after ten minutes (the first true *announced-versus-acted* + watchdog); ADR 0162's bound on a plan, which 230 found never fires (`late: false` for ever). +- **Missing:** a list of the signals, their bounds and their owners; the plan bound working; the event + loop's last-taken age (187); asks counted against outcomes (186); the bus's own advisories (slow + consumer, maximum deliveries, permission violations) read as observations instead of log lines. +- **How it is checked:** + 1. A **signals table** — signal, emitter, cadence or trigger, bound, condition raised — kept in the + to-be design, and compiled into the controller. A test generated from it suppresses each signal in + turn and asserts the named condition is raised within its bound and cleared when the signal + returns. + 2. **Live:** the self-check (P6) reports, for every row, the age of the newest signal, so a row that + never fires is itself visible. + +## P6 — The mesh checks itself continuously, against live facts, and says what it found outward + +The invariants the design states are probed **against the running mesh** on a schedule, not only in unit +tests. A violation is a **condition** — durable, with since-when, evidence and who can resolve it — +shown in `status` and **sent to the operator** through the output channel. The checker's own heartbeat is +watched from somewhere it does not run. + +- **Answers:** (d), (h). Issues 177, 187, 238, 245, 253, 262. +- **Exists:** `status` itself (assembled from reports, as how-we-build §5 requires); 017's *condition* + shape; 028's output seat, researched only; individual live checks done by hand (262: "the live check + stays by hand, on each resolver"); 253's measurement before the collector's first run — the one time + in the window a check ran before the damage. +- **Missing:** a scheduled runner, a condition store, a channel out, and a watcher's watcher. +- **How it is checked:** + 1. Every invariant in the to-be design that names a live check is a **probe** in the self-check's + registry; a check over the design documents counts invariants with a stated live probe against + those without, and the number may only go down. + 2. The self-check publishes a heartbeat; a second machine's watcher raises "the self-check is silent" + through a channel that does not depend on the control node. + 3. **Live, once:** a deliberately broken invariant on a lab mesh appears in `status` and as a + notification within one probe interval. + +## P7 — A known failure heals itself, under a brake, and every repair is said + +A failure that has been repaired by hand twice is a failure the mesh must repair itself: by running the +ordinary path again (017 P2), never by destroying (017 P3), with a budget and a back-off, and with one +line and one event saying what it did and why. When the budget is spent, or the only repair destroys, +it is a condition and a notification, not a retry. + +- **Answers:** (f). Issues 179, 208, 214, 230, 233, 248, 254, 257, 264, 267. +- **Exists:** ADR 0224 §5, *detected automatically, repaired where safe, loud where not*, applied to the + identity provider's admin; the host's launcher rolls back once to known-good and then halts (ADR + 0141); the provisioner's `holds` re-provisions what a backend lost (017/02); `broker consumer-reset` + and `plans close` as verbs. +- **Missing:** the general mechanism. The most frequent hand act in the window — **a push by hand to + unstick a plan waiting on a report** — has no healer: the controller could ask the machine to report + again (the machine knows what it applied) before waiting longer. +- **How it is checked:** + 1. Every hand act on the core is done through a verb that records it (who, what, why) — the **hand-act + log**. Its count per week is a reported number; an act recorded twice for the same cause is a + condition asking for a healer. + 2. Each healer ships with a test that induces its failure, asserts the repair and the event, and + asserts the brake after the budget. + +## P8 — The core upgrades itself one machine at a time, health-gated, and rolls back on its own + +A new controller, host, runtime or bus reaches one machine first; it is **healthy** only when the +self-check's probes for that component pass there (not merely when it "reported applied"); the rest +follow only then. A component that does not become healthy within its bound is rolled back to the last +known good **by something other than itself**, and the rollback is said. The component being replaced is +never the only witness of its successor's success. + +- **Answers:** (g). Issues 201, 204, 213, 214, 217, 230, 245, 248, 264, 266. +- **Exists:** ADR 0218's one machine first, for modules, with "applied and current" as the gate; ADR + 0141/0142's side-by-side host versions, known-good and the launcher's single rollback; ADR 0185's + controller serving what it can when it is behind its seat's row; 264's report kept on disk across the + hand-over. +- **Missing:** a health definition per core component; a gate stronger than "reported"; rollback for + the controller, the runtime and the bus; a lease hand-over between controllers (P1); a planned, + rehearsed path for the bus, which is still one process on one machine and whose next upgrade (266: + 2.10 → 2.11) is one-way. +- **How it is checked:** + 1. **Lab:** a deliberately broken build of each core component (one that starts and does nothing; one + that crashes; one that cannot reach the bus) is merged on a lab mesh. Each is rolled back without a + hand, the mesh ends on the previous build, and a condition and a notification say so. + 2. **Live:** every core rollout leaves a record — first machine, health verdict, time to verdict, + rolled back or not — readable through `plans`. + +## P9 — A check is fed the real mesh's facts before a change is merged + +A check whose verdict depends on the environment — names and their lengths, the machines that exist, +the catalogue as it is, the host's validation, the C library, the server versions — runs against **the +mesh's real facts**, exported and anonymised, before merge. A dependency's version that the mesh runs is +the version its tests run. + +- **Answers:** (h), (i). Issues 177, 202, 228, 236, 262, 263, 266. +- **Exists:** 266's `TestTheImageIsTheServerTestedHere` (the bus image's release equals the tested + server's); ADR 0223's composition test that renders the resolver's machine list; 202's test against the real + catalogue (run by hand). +- **Missing:** the export of facts; a merge gate that composes every real machine's declaration with the + change and runs the host's validation over it (which would have refused 236, 263 and 202 in their own + pull requests); a resolver test under musl as well as glibc. +- **How it is checked:** + 1. The controller exports a **facts snapshot** (machines, names, assignments, seats, catalogue + commit — no secrets, no addresses) daily; the core repositories' merge check composes every machine + from it with the change applied and runs the host's validation; a pull request that makes any + machine fail to compose or validate fails its check, naming the machine's role and the module. + 2. The snapshot's age is a signal (P5). + +--- + +## Candidates weighed and not kept as principles + +- **"Every invariant has a live probe, not only a unit test"** — merged into P6; it is how P6 is built. +- **"Environment-dependent checks run against the real mesh's facts"** — kept as P9; the third-party + case (i) folded into it, because pinning and testing the version that runs is the same act. +- **"Self-healing is the default"** — kept as P7 but narrowed to *known* failures, those repaired by + hand twice. A default of healing everything heals what is not understood, which is how a repair + destroys (241's reconcile was, in its own terms, healing). +- **"No loop blocks on long work"** (184, 248, 175) — not a separate principle: a blocked loop is a + signal gone silent (P5, the loop's last-taken age) and a design defect each owner fixes; stating it + as a principle adds a rule with no mesh-wide check. +- **"Clear plan of execution"** from the mandate — not a principle about the mesh; it is the roadmap in + [03](03-mechanisms-and-roadmap.md), and each phase there states its own verification. + +## How the principles relate + +P2 and P1 **prevent** the races. P4 **prevents** the silent drops. P3, P5 and P6 **catch** whatever the +first three miss, at the cost of minutes, not hours. P7 and P8 **repair**. P9 **moves** the catching +before merge. The order of the roadmap follows from that: catching first, because it is cheapest and +covers every class, including the ones nobody has met yet. diff --git a/01-RESEARCH/031-a-core-that-cannot-fail-silently/03-mechanisms-and-roadmap.md b/01-RESEARCH/031-a-core-that-cannot-fail-silently/03-mechanisms-and-roadmap.md new file mode 100644 index 0000000..ebc3f58 --- /dev/null +++ b/01-RESEARCH/031-a-core-that-cannot-fail-silently/03-mechanisms-and-roadmap.md @@ -0,0 +1,221 @@ +# 03 — Mechanisms and a phased roadmap + +The principles of [02](02-principles.md) need few new things. Most of the parts exist in some form; +what is missing is the connective tissue that makes a fact the mesh already has reach someone without +being asked. This document names the mechanisms, then orders them into phases **by risk removed per unit +of effort**, each phase with deliverables and a verification that says it is done. + +Effort is given in **focused working days** of one agent-and-operator pair, at the pace the record shows +(a located issue to a merged fix in under a day is common). The figures are for ordering, not promises. + +## The mechanisms + +### M1 — Conditions (P5, P6) + +017's *condition*, built: a durable fact about something the mesh owns — what is wrong, since when, the +evidence, what was tried, who can resolve it, and whether it may clear itself. Kept in a key-value +bucket the controller writes (one writer, P1), keyed by subject (`plan/`, `machine/`, +`provider//`, `core/`). Raised and cleared by observation only; a person +can **silence** one for a stated time, never resolve it. `status` becomes, first, the list of open +conditions; the all-well sentence is "no open conditions". ADR 0224's provider standing is the first +condition kind and moves into it unchanged. + +### M2 — Watchdogs from a signals table (P5) + +One table, compiled into the controller, of every signal the core expects. Its first rows, each from an +issue in [01](01-evidence.md): + +| Signal | Bound (to be measured, then set) | Condition raised | Issue | +|---|---|---|---| +| machine heartbeat | 3 × interval | machine silent (asleep is a declared state, ADR 0211) | 187 | +| report after a send | the machine's last apply duration × 3, at least 2 min | sent, not reported — then **ask the machine to report again** (M5) | 230, 257, 264, 267 | +| plan tier progress | per tier, from build and apply durations | plan stalled at tier N, waiting on X | 214, 230, 254 | +| controller event loop took something | 2 min while the stream has pending | controller deaf | 184, 248 | +| merge announced → plan made or "nothing reads it" | 10 min (exists, 266) | merge never acted on | 248, 266 | +| build asked → outcome | build's own declared timeout | ask lost | 186 | +| call `running` → finished | the verb's declared bound | call hung | 265 | +| provider standing repeated | 30 min (exists, ADR 0224) | provider silent | 179 | +| bus advisories: slow consumer, maximum deliveries, permission violation | any | bus refused or dropped X for Y | 183, 187, 217, 265 | +| self-check heartbeat | 2 × its interval, watched from a second machine | the watcher is silent | — | +| facts snapshot age | 2 days | merge checks run on stale facts | 263 | + +The bus advisories are the cheapest row: the server already publishes them on its system subjects, and +the controller only has to subscribe (read-only) and translate each into the mesh's words, naming the +call or the consumer, as 265 now does for a refused reply. + +### M3 — The self-check: `doctor` (P6) + +A controller verb, `doctor`, and the same code run every few minutes by the controller itself. Each run +executes the **probe registry** — the live form of the design's invariants — and raises or clears +conditions. First probes, each an invariant that a person checked by hand in the window: + +- every machine's declaration composes, and every host would accept it (236, 263); +- every holder of the mesh's resolver answers a machine name for IPv4 and NODATA for IPv6 (262); +- every seat on record has a live holder that answers (208, 218); +- every kept archive is held by a manifest (253 — the controller's collection command already reports it); +- exactly one controller holds the lease (P1); +- every durable consumer's position is near its stream's head (248's replay, 266's skip); +- no address the mesh owns is in a ban list (238); +- `status` answers in full within its bound (P3). + +`doctor` with no argument answers the last run's verdict at once (P3) and, with `run`, runs now under an +id. Its own heartbeat is a signal (M2), watched from a second machine. + +### M4 — The output channel (P6) + +[Research 028](../028-the-meshs-output-channel/00-overview.md)'s seat, built minimally first: **one** +channel the operator chose (028 records it), plus the desktop notifier where the operator is. A condition +is sent when raised, once more if it lasts past a bound, and when it clears. Deduplicated by the +condition's key. The watcher's watcher (028's open question) is the second-machine watchdog of M3, +sending through a channel that does not pass through the control node. + +### M5 — Healers (P7) + +A healer is a registered response to one condition kind: its repair (the ordinary path again), its +budget, its brake, and the event it emits. First healers, all from hand acts in [01](01-evidence.md) §(f): + +| Condition | Repair | Brake | +|---|---|---| +| sent, not reported | ask the machine to report what it last applied (it keeps it since 264); if that names another declaration, send again | twice, then condition | +| plan stalled on a superseded or finished wait | close the plan with its note (`plans close`, done by the mesh) | once | +| seat holder without its worker | raise the seat's objects again (208) | once per holder | +| consumer far behind on a history stream | `broker consumer-reset` (248) — **only** for the consumers the table marks resettable | once, then condition | +| provider admin refuses the mesh's secret | ADR 0224 §5 (exists) | exists | + +And the **hand-act log**: every repair a person makes on the core goes through a verb that records who, +what and why. Its weekly count is the measure of P7. + +### M6 — Order and epochs (P1, P2) + +- The controller takes a **lease** in a key-value bucket before it acts and renews it; its revision is + the **epoch** every declaration and plan write carries. A starting controller waits for the lease; + the outgoing one stops sending when it loses it. That closes 204 and makes 201/214 detectable. +- A **report carries the sequence** of the declaration it is about; the controller keeps the highest + per machine and refuses older accounts by sequence, not by a digest lookup (256, 257, 267 become one + rule). +- **The host has one apply queue.** A delivery and the reconcile are two reasons to enqueue the same + act; the queue applies the newest declaration once, and makes one report (257/261/267 become + impossible rather than handled). +- Every receiver's refusal of something stale is a counted line (P2's live check). + +### M7 — Staged, reversible core upgrades (P8) + +- **A health definition per core component**, written as probes in M3's registry: the controller + answers `status` in bound and holds the lease; the host has reported its current declaration; the + runtime has announced and answers a PING; the bus has every stream and every durable consumer and + passes a request/reply round trip. +- **The gate:** ADR 0218's first machine is judged by those probes, not by "applied". +- **Rollback by a witness that is not the new build:** the host's launcher for the host (exists); for + the controller, the previous controller's process kept installed beside it, re-started by the host + when the new one does not take the lease in bound; for the runtime, the host's known-good the same + way. Each rollback is a condition, so it is said. +- **The bus is planned, not rolled.** A bus upgrade is a declared maintenance step: streams snapshotted, + the step announced as a condition while it runs, every consumer's position checked after (M3). The + one-way 2.10 → 2.11 upgrade of 266 is the first. Whether the bus should become a cluster of three so + that it can be upgraded live is a question for its own effort. + +### M8 — The facts snapshot and the merge gate (P9) + +The controller exports a facts snapshot (machines by role and name length, assignments, seats, catalogue +commit) to a place the build seat reads. The core repositories' and the catalogue's merge checks compose +every machine with the change and run the host's validation. The resolver module's tests run under musl +and glibc. A bus, store or library version the mesh runs is the one its tests run (266's pattern, +generalised). + +### M9 — The lab replay (all) + +[Research 019](../019-a-warm-twin-of-the-running-mesh/00-overview.md)'s warm twin, or a throwaway lab of +three containers, running **scripted replays of each incident** in the window: a reconcile due during a +push; a host self-update during its report; a bus reload during a call; a consumer with several filters +under mixed traffic; a missing consumer made with the server's default; an unreadable contributions +file; a controller rebuilding itself mid-plan; two controllers at once. Each replay asserts the +principle's outcome (refused stale, condition raised, healed, rolled back). They run on every merge to a +core repository. + +--- + +## The roadmap + +Ordered by **risk removed per day**. Detection comes first because it covers every class at once, +including the ones not met yet; prevention second; repair and staging third; the pre-merge and lab work +last because they are larger and pay off over months. + +### Phase 0 — Finish what is in flight (2–3 days) + +- Land and roll out the located fixes: 244, 264, 265, 266 (including the bus's planned 2.11 upgrade, done + as M7's first planned bus step), 267, 257/261. +- Make `calls` durable (a key-value bucket, bounded by count and age), and bring `status` inside its own + answer bound. +- Start the **hand-act log** now, before anything else, so every later phase is measured against a + baseline. +- **Done when:** the four located core issues resolve with their live checks; a controller restart + keeps `calls`; `status` answers in full within ten seconds. + +### Phase 1 — The mesh says when it is wrong (5–8 days) + +- M1 conditions (ADR 0224's standing moved into them); M2 watchdogs for the first eight rows; the bus + advisories subscribed and translated; M3 `doctor` with the first probes; M4 with one channel and the + second-machine watcher. +- **Done when:** on a lab mesh, suppressing each signal in the table raises its condition within its + bound and sends a notification; clearing it clears both. Live: a week of conditions read back, every + one either real or a bound corrected. +- **Risk removed:** every class in [01](01-evidence.md) moves from "noticed by a person, hours later" to + "said by the mesh, minutes later". + +### Phase 2 — Order and one writer (5–7 days) + +- M6: the controller's lease and epoch; a report's sequence; the host's single apply queue; stale + refusals counted. +- The writers table and the signals table written into a to-be design, with their compile-time checks. +- P4's lint for empty-on-error readers in the core repositories, and the destructive-delta brake in the + provisioner harness. +- **Done when:** the lab replays of 204, 257/261/267 and 241 end with "refused stale", "one report" and + "withdrawal braked" respectively; the contract test for every consumed subject exists. +- **Risk removed:** class (a), the largest by count, and the destructive half of (b). + +### Phase 3 — Healers (3–5 days) + +- M5's first healers and the rule that a repair done by hand twice asks for one (from the hand-act log). +- **Done when:** a lab mesh recovers from each induced failure in the healers table with no hand, says so, + and brakes after its budget. Live: a week in which the hand-act log has no repeat. + +### Phase 4 — Core upgrades that roll back (8–12 days) + +- M7: health definitions; the gate; rollback for the controller and the runtime; the bus as a planned + step. +- **Done when:** on a lab mesh, a deliberately broken build of the controller, the host and the runtime + is each rolled back without a hand, the mesh ends on the previous build, and the rollback is a + condition and a notification. Live: the next three core rollouts each record a health verdict. +- **Risk removed:** class (g) — the failures that take the control path itself down. + +### Phase 5 — Checks before merge, and the replay suite (8–12 days, then ongoing) + +- M8: the facts snapshot and the compose-and-validate merge gate; the libc matrix; versions tested as + run. +- M9: every incident in the window as a scripted replay, run on every core merge; each new core issue + adds its replay as its "how it is checked". +- **Done when:** the replays of 236, 262, 263 and 266 fail on the commit before their fix and pass after; + a new core issue cannot resolve without a replay or a stated reason why none is possible. + +**Total, roughly six to eight weeks of focused work**, with the first visible change — the mesh saying +when it is wrong — inside the first two. + +## The five largest risks today + +Ranked by likelihood × damage, from the window's evidence and the mesh as read the night this effort +began: + +1. **A stall nobody is told about.** A plan, a merge or a report that stops is noticed only by someone + looking (230, 248, 257, 264, 266, 267; issue 187 still open). Every release goes through a plan. +2. **A destructive act on a misread input.** 241 dropped seven databases on one unreadable file; 234 + undeclared four modules from a stale composition; 245's advice rebuilt everything and cut the bus; + 253's collector would have deleted every archive. There is no general brake on a large withdrawal. +3. **The core replacing itself with nothing to roll it back.** A controller build that starts but cannot + plan (201, 214) halts every later change, because the controller is what plans the fix. Only the + host has a launcher rollback. +4. **The bus as a single, un-upgradeable process.** It runs a release that skips messages (266), its + reload drops owed replies (265), replacing it cuts every machine off (245), and its next upgrade is + one-way and needs a restart. +5. **Concurrent actors with no lease or report order.** Several agent sessions, plans and reconciles act + at once; on the night this began, four named pushes, one per machine, started within two seconds. The digests + catch most of it now, one rule per message kind; the next message kind will not have one.