To-be 48 Phase A is being built: name its code and say what the build chose
This commit is contained in:
@@ -1,7 +1,7 @@
|
|||||||
---
|
---
|
||||||
layer: to-be
|
layer: to-be
|
||||||
status: designed
|
status: in-progress
|
||||||
code: []
|
code: [mesh-host, mesh-controller, mesh-lab]
|
||||||
updated: 2026-10-07
|
updated: 2026-10-07
|
||||||
decisions:
|
decisions:
|
||||||
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
|
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
|
||||||
@@ -232,6 +232,43 @@ stays up, and the gate's points as they stand.
|
|||||||
|
|
||||||
Phases A and B may be built together; A is live first, because it judges without a single declaration.
|
Phases A and B may be built together; A is live first, because it judges without a single declaration.
|
||||||
|
|
||||||
|
## As built — Phase A
|
||||||
|
|
||||||
|
Built on one feature branch in each of mesh-host, mesh-controller and mesh-lab; not yet merged or rolled out.
|
||||||
|
What the build chose where this design left it open:
|
||||||
|
|
||||||
|
- **The look.** The node-engine looks every 15 seconds: one inspect of every container it runs for a module,
|
||||||
|
one show of every unit per service manager, and nothing else — no execution inside a container, no event
|
||||||
|
history. A runtime that does not answer is `unknown` for every container, never `down`. What it counts is
|
||||||
|
kept in a file beside the node's state.
|
||||||
|
- **What a restart is.** The runtime's own count is read only to see it move. A count that moved is the
|
||||||
|
runtime restarting what exited, counted after the grace; a container made again (its identity changed with
|
||||||
|
no count moved) or started again by somebody (its start time moved with no count moved) is a new start,
|
||||||
|
from a fresh grace, its counted restarts kept. A container a maintenance window holds is `held`, and the
|
||||||
|
window ending is a start. A resource not running after its grace is `unhealthy` — `restarting` while the
|
||||||
|
runtime or the manager is restarting it, `down` otherwise — before any restart is counted.
|
||||||
|
- **What is judged.** A module's containers that stay up, its services stated running and its processes
|
||||||
|
that stay up. What the mesh declares in its own right (no module) and what an adopted machine holds as it
|
||||||
|
was found are not, in this phase.
|
||||||
|
- **The statement.** In every report the engine makes, judged right after the apply, so a resource the
|
||||||
|
apply started says `starting`; as an event on the machine's own subject (inside the grant every host
|
||||||
|
already has — no grant changed, and the writers table's row for the machine's report names the subject
|
||||||
|
beside the report's) on each change, again every minute while one is not healthy, and every five minutes
|
||||||
|
anyway, so a controller started again knows a healthy machine's state without waiting for its next apply.
|
||||||
|
Ordered by when the engine looked; an older statement is refused.
|
||||||
|
- **The controller** keeps the newest statement per machine in its store, with each module's run of
|
||||||
|
unhealthy statements, so the gate, `node show` and a restarted controller read the same word. The
|
||||||
|
condition joins the scopes as `module`. The gate reads a statement heard since the send; a resource
|
||||||
|
starting, unhealthy, held or unknown makes the judging not a pass — the build fails at the bound, as
|
||||||
|
ADR 0236 §3 says, rather than at once. A machine whose engine states nothing is judged as before.
|
||||||
|
- **The proof.** mesh-lab's replay `R-crashloop` raises a container whose program exits at start, has the
|
||||||
|
node-engine at its commit judge it through the runtime, and the controller at its commit judge the gate
|
||||||
|
from what the engine said. Proved: it fails on the trunk before this build, and passes on it.
|
||||||
|
|
||||||
|
Still to do for Phase A's "done when": every machine's report carrying a state for every long-running
|
||||||
|
resource, and a module stopped on purpose raised and cleared, are read on the live mesh once both are rolled
|
||||||
|
out — the node-engine first, then the controller.
|
||||||
|
|
||||||
## What is not decided here
|
## What is not decided here
|
||||||
|
|
||||||
- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).
|
- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).
|
||||||
|
|||||||
Reference in New Issue
Block a user