From dea72dc4fc8cb206792790f7fe074092b0d26cc7 Mon Sep 17 00:00:00 2001 From: jochen Date: Wed, 7 Oct 2026 02:03:47 +0200 Subject: [PATCH] To-be 48 Phase A is being built: name its code and say what the build chose --- .../48-a-module-says-how-it-is-healthy.md | 41 ++++++++++++++++++- 1 file changed, 39 insertions(+), 2 deletions(-) diff --git a/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md b/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md index 81b8e3f8..690056dc 100644 --- a/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md +++ b/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md @@ -1,7 +1,7 @@ --- layer: to-be -status: designed -code: [] +status: in-progress +code: [mesh-host, mesh-controller, mesh-lab] updated: 2026-10-07 decisions: - 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md @@ -232,6 +232,43 @@ stays up, and the gate's points as they stand. Phases A and B may be built together; A is live first, because it judges without a single declaration. +## As built — Phase A + +Built on one feature branch in each of mesh-host, mesh-controller and mesh-lab; not yet merged or rolled out. +What the build chose where this design left it open: + +- **The look.** The node-engine looks every 15 seconds: one inspect of every container it runs for a module, + one show of every unit per service manager, and nothing else — no execution inside a container, no event + history. A runtime that does not answer is `unknown` for every container, never `down`. What it counts is + kept in a file beside the node's state. +- **What a restart is.** The runtime's own count is read only to see it move. A count that moved is the + runtime restarting what exited, counted after the grace; a container made again (its identity changed with + no count moved) or started again by somebody (its start time moved with no count moved) is a new start, + from a fresh grace, its counted restarts kept. A container a maintenance window holds is `held`, and the + window ending is a start. A resource not running after its grace is `unhealthy` — `restarting` while the + runtime or the manager is restarting it, `down` otherwise — before any restart is counted. +- **What is judged.** A module's containers that stay up, its services stated running and its processes + that stay up. What the mesh declares in its own right (no module) and what an adopted machine holds as it + was found are not, in this phase. +- **The statement.** In every report the engine makes, judged right after the apply, so a resource the + apply started says `starting`; as an event on the machine's own subject (inside the grant every host + already has — no grant changed, and the writers table's row for the machine's report names the subject + beside the report's) on each change, again every minute while one is not healthy, and every five minutes + anyway, so a controller started again knows a healthy machine's state without waiting for its next apply. + Ordered by when the engine looked; an older statement is refused. +- **The controller** keeps the newest statement per machine in its store, with each module's run of + unhealthy statements, so the gate, `node show` and a restarted controller read the same word. The + condition joins the scopes as `module`. The gate reads a statement heard since the send; a resource + starting, unhealthy, held or unknown makes the judging not a pass — the build fails at the bound, as + ADR 0236 §3 says, rather than at once. A machine whose engine states nothing is judged as before. +- **The proof.** mesh-lab's replay `R-crashloop` raises a container whose program exits at start, has the + node-engine at its commit judge it through the runtime, and the controller at its commit judge the gate + from what the engine said. Proved: it fails on the trunk before this build, and passes on it. + +Still to do for Phase A's "done when": every machine's report carrying a state for every long-running +resource, and a module stopped on purpose raised and cleared, are read on the live mesh once both are rolled +out — the node-engine first, then the controller. + ## What is not decided here - Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).