Merge pull request 'To-be 48 Phase A is being built: in-progress, its code named, what the build chose' (#162) from feat/a-module-says-how-it-is-healthy into main

This commit was merged in pull request #162.
This commit is contained in:
2026-10-07 00:39:41 +00:00
@@ -1,7 +1,7 @@
---
layer: to-be
status: designed
code: []
status: in-progress
code: [mesh-host, mesh-controller, mesh-lab]
updated: 2026-10-07
decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
@@ -232,6 +232,43 @@ stays up, and the gate's points as they stand.
Phases A and B may be built together; A is live first, because it judges without a single declaration.
## As built — Phase A
Built on one feature branch in each of mesh-host, mesh-controller and mesh-lab; not yet merged or rolled out.
What the build chose where this design left it open:
- **The look.** The node-engine looks every 15 seconds: one inspect of every container it runs for a module,
one show of every unit per service manager, and nothing else — no execution inside a container, no event
history. A runtime that does not answer is `unknown` for every container, never `down`. What it counts is
kept in a file beside the node's state.
- **What a restart is.** The runtime's own count is read only to see it move. A count that moved is the
runtime restarting what exited, counted after the grace; a container made again (its identity changed with
no count moved) or started again by somebody (its start time moved with no count moved) is a new start,
from a fresh grace, its counted restarts kept. A container a maintenance window holds is `held`, and the
window ending is a start. A resource not running after its grace is `unhealthy` — `restarting` while the
runtime or the manager is restarting it, `down` otherwise — before any restart is counted.
- **What is judged.** A module's containers that stay up, its services stated running and its processes
that stay up. What the mesh declares in its own right (no module) and what an adopted machine holds as it
was found are not, in this phase.
- **The statement.** In every report the engine makes, judged right after the apply, so a resource the
apply started says `starting`; as an event on the machine's own subject (inside the grant every host
already has — no grant changed, and the writers table's row for the machine's report names the subject
beside the report's) on each change, again every minute while one is not healthy, and every five minutes
anyway, so a controller started again knows a healthy machine's state without waiting for its next apply.
Ordered by when the engine looked; an older statement is refused.
- **The controller** keeps the newest statement per machine in its store, with each module's run of
unhealthy statements, so the gate, `node show` and a restarted controller read the same word. The
condition joins the scopes as `module`. The gate reads a statement heard since the send; a resource
starting, unhealthy, held or unknown makes the judging not a pass — the build fails at the bound, as
ADR 0236 §3 says, rather than at once. A machine whose engine states nothing is judged as before.
- **The proof.** mesh-lab's replay `R-crashloop` raises a container whose program exits at start, has the
node-engine at its commit judge it through the runtime, and the controller at its commit judge the gate
from what the engine said. Proved: it fails on the trunk before this build, and passes on it.
Still to do for Phase A's "done when": every machine's report carrying a state for every long-running
resource, and a module stopped on purpose raised and cleared, are read on the live mesh once both are rolled
out — the node-engine first, then the controller.
## What is not decided here
- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).