diff --git a/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md b/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md index a3fd357a..1ca7159c 100644 --- a/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md +++ b/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md @@ -1,6 +1,7 @@ --- -status: active +status: graduated initiated: 2026-10-06 +became: [02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md, 03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md] touches: - 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md - 03-DESIGN/01-to-be/18-building-a-module.md @@ -75,3 +76,10 @@ Evidence gathered on the live mesh, options weighed and a decision drafted, 2026 report, the condition raised by the controller on the second look, so the gate needs no new rule; a provider down said once at the provider; a proposed decision text with how each rule is checked, and the migration of the 68 modules. + +**Graduated 2026-10-07** on the operator's word (*"yes, turn it into a decision"*), as +[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) and +[to-be 48](../../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md). The record takes the proposed +decision, adding: a fifth state, `held`, for a resource under a maintenance step; a ceiling of five minutes on +grace plus failing looks, so a broken start is said inside the gate's bound; and the gate's *own health* reading +the stated health, so a resource still starting is not yet a pass. diff --git a/02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md b/02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md index 932a9a5a..5d8a0632 100644 --- a/02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md +++ b/02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md @@ -92,6 +92,13 @@ times, at least forty seconds apart and two minutes after the send, within ten m on one machine is judged the same way; only then does the plan go on. A policy of *together* is not gated: it is the module saying it must change everywhere at once. +> **The mechanism changed — 2026-10-07, by [ADR 0240](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).** +> What stands: the three judgings, their spacing and bound, and the points above. What moved: *a module's own +> health* is no longer only its tools served — every long-running resource of the module on that machine must +> also be stated healthy by the node-engine, a resource still starting is not yet a pass, and an unhealthy one +> is the condition `module...unhealthy` that *no condition raised since the send* already +> holds the module on. Designed in [to-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md). + **3. A build that fails its gate is put back there, once, and never sent again on its own.** The plan stops; the build is marked failed at its gate; the module's registered build goes back to the build the first machine ran before — the newest successful build of that commit still kept ([ADR 0189](0189-the-store-keeps-what-the-records-name.md) diff --git a/02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md b/02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md new file mode 100644 index 00000000..5b1ab8be --- /dev/null +++ b/02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md @@ -0,0 +1,263 @@ +--- +topic: what runs on it +status: accepted +date: 2026-10-07 +deciders: jochen +reconstructed: false +extends: 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md +--- + +# 240. A module says how it is healthy, and the node-engine judges it + +## Context + +The release gate ([ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md) +§2) judges a catalogue module on its first machine by what the mesh sees from outside: the build reported +applied, no witness put it back, no condition raised since the send about the machine or the module there, +and its tools served. None of that looks at what the module *runs*. To-be 45 names the gap in its Phase 4: +*container state in the node-engine's report, without which a container that crash-loops after its compose +applied is seen only through what it breaks.* + +Measured on the live mesh for [research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md) +([evidence](../01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md)), 2026-10-07: + +- **125 catalogue modules; 68 run something long-lived** — 49 a container that stays up, 19 a service + unit stated `running` and no such container. 50 run only their bundle in the node tools, 7 only files, + directories and packages. +- **No manifest can declare a health check, and nothing reads one.** The node-engine's report carries no + container or unit state; no probe of the self-check reads one. +- **19 of 73 long-running catalogue containers have an image that ships a check** (7 of 45 container + modules). The mesh never reads it. **Two of those were wrong in the mesh's hands**: the studio and the + flow editor read *unhealthy* while working, because their checks ask an address the program does not + bind in the mesh's configuration. Read without proof, they would have rolled back two good builds. +- **Every restart count is 0**, because a recreate loses it; the runtime's event history on the home + server is about a minute long, pushed out by its own check executions. +- **Of seven recent incidents**, liveness alone would have caught the agent server's crash loop (about a + hundred restarts, found by a person reading its log for something else); an HTTP readiness check the + web application that accepted TCP and answered nothing for eleven hours + ([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)); + only the module's own check the identity provider's refused administrator + ([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)). + A provisioner runtime that restarted until the overlay was up and then worked + ([issue 058](../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md)) + is the warning: a restart count without a start period reads churn that stops as a crash loop. +- **One sample is not a finding.** A single unanswered question raised an urgent alert nobody could read + ([issue 277](../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md)); + the self-check now raises such a finding on the second look (to-be 45 §4). + +The mesh now rolls a module out on its own and puts the previous build back when the first machine is not +healthy. That promise is only as good as "healthy" is, and for a module it is judged today from the outside. + +**Checked against GENESIS.** *Failure must be loud* — a crash loop behind every passing check is work that +reported success and did nothing. *Anything requiring a human to notice it will be noticed late* — eleven +hours, and a hundred restarts, each found by a person. *Evidence over assertion* — "applied" is an +assertion about a declaration; "it answers" is a measurement. *The mesh notices when something is wrong +before you do* ([effect](../00-META/effect.md)). *Long-lived user services rather than an orchestrator* +([context](../00-META/context.md)) is why this record judges and reports, and restarts nothing on health: +an orchestrator's liveness restart is the part it does not take. *The mesh is a guest on a personal node* +bounds the cost: the checks' floors below. Nothing here conflicts with GENESIS. + +## Considered Options + +**Who runs the checks.** + +1. *The container runtime's own check, read by the node-engine.* Rejected as the whole answer: containers + only — nothing for the 19 service-only modules or anything only a module's tool knows; every look is an + execution inside the container, which on the home server already pushes every lifecycle event out of + the runtime's history; an image check the module never stated is a check nobody owns, and two of 19 were + wrong. Kept as one *kind*, adopted by name and proved. +2. *The module's own tool answers "healthy", and that is the judgement.* Rejected as the judge: a bundle is + hosted by the node tools, so a tool that answers proves the bundle up, not the server; and a component + that judges itself is what [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) + rule 8 refuses for the core. Kept as one kind, for function no endpoint shows, and only beside a check + the module does not run itself. +3. *The node-engine runs everything itself, including commands, on its own schedule.* Rejected in part: a + command that only makes sense inside the container is better timed by the runtime's own retries and + start period than by an execution per look from outside. +4. **The node-engine owns every check and every verdict, and runs each kind where it is cheapest.** + Chosen: HTTP, TCP and unit checks it makes itself; a command it hands to the runtime as that + container's check and reads; a tool it asks through the node tools. One runner and one reader per + machine, for every hosting form, and it keeps what the runtime forgets. + +**Where the result goes.** + +1. *Only a field of the report.* Rejected alone: a report follows an apply, so a container that goes bad at + 03:00 waits for the next one. +2. *Only an event on each change.* Rejected alone: events are lost or replayed; an event is a sample, not + a state. +3. *A verb the controller calls per module when it judges.* Rejected: a pull per module per judging, and a + machine slow to answer reads as unhealthy. +4. **State in the report, its change on the bus, and a condition the controller raises on the second + look.** Chosen. It is how the core already says its own state; the gate, the self-check, the healers + ([ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)) and + the operator's conversation ([ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md)) all + already act on conditions. + +**What is judged.** + +1. *Liveness only.* Rejected: no declarations needed, but it misses issue 145 (the port was open and the + program running) and issue 179. +2. *Readiness only, where declared.* Rejected: nothing is judged until every module has declared, and + liveness alone would have caught the worst incident of the window. +3. **Liveness for every long-running resource at once, readiness where declared, the declaration required + over a migration.** Chosen: the gate means something from the first build. + +**What is done with an unhealthy module.** + +1. *Restart it, as an orchestrator's liveness probe does.* Rejected for this record: the evidence holds no + case where a restart would have fixed anything — the crash loop was restarting already — and a restart + hides the failure the gate is meant to see. Whether a healer restarts what stays unhealthy is left to a + record of its own under ADR 0231. +2. **Nothing restarts on health; the condition reaches a person or a healer.** Chosen. + +**A provider down.** With 12 consumers of the database provision and 36 of a route: + +1. *Ignore it.* Rejected: twelve conditions for one fault, and twelve gates failed for something none of + them did. +2. *Judge providers first, consumers only once their providers are healthy.* Rejected: a consumer broken on + its own is not said while its provider is down, which is when it is most needed. +3. **A consumer's check names the provision it exercises; while that provision's provider is unhealthy on + the record, the consumer's finding is held under the provider's condition and its gate waits.** Chosen. + It is how [issue 281](../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md) + already treats a machine-level fault: what is the machine's is never pinned on a module. + +**Where the declaration sits.** + +1. *One per module.* Rejected: a module of eleven containers (the mail module) could not say which is wrong. +2. **On each long-running resource**, beside the resource's other fields. Chosen. + +**The migration.** + +1. *Required at once.* Rejected: 68 modules to change before the next merge, and nothing judged until all + are. +2. *Optional for ever.* Rejected: 38 of 45 container modules would stay at liveness, and a field that is + optional for ever is a field half the catalogue never gets. +3. **Liveness at once; the declaration required by a date, counted down.** Chosen. + +## Decision + +**1. Liveness is judged for every long-running resource, with no declaration.** A container that stays up, +a process that stays up and a service stated `running` are *alive* when running and not restarted more than +once within the settle window after their grace period. The node-engine observes this itself on each tick +and keeps the restarts it counted across recreates and across its own restarts; it never relies on the +runtime's restart count or event history. A resource held still by an open maintenance step +([ADR 0189](0189-the-store-keeps-what-the-records-name.md)) is neither alive nor dead: it is said as held, +and judged again when the step ends. + +**2. A module declares how each long-running resource is ready**, in a field named `health` on that +resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint and the +status expected, a TCP connect to a declared endpoint, a command in the container, the unit's own +readiness, or one of the module's tools — with an interval (default 30 s, **not under 10 s**), a timeout +under the interval, a number of failing looks in a row before it is unhealthy (default 3, **not under 2**), +and a grace period after a start (default 60 s), in which failure does not count. The grace and the failing +looks together are at most five minutes, so a resource broken from its start is said within the gate's +bound. An endpoint is named by its `listens` name, never by a port or an address, so a check follows the +machine's ports as the endpoint does. A check of function that no endpoint shows is a module's own tool, and +only in addition to a check the module does not run itself. + +**3. The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes itself; +a command in the container it hands to the runtime as that container's check and reads the state; an +adopted image check it reads the same way; a tool it asks through the node tools. Nothing else on the machine +judges a module, and nothing else sets a container's check. + +**4. The state goes in the report, its change on the bus, and a condition is the controller's.** Each +report carries, per module and long-running resource, a state — healthy, unhealthy, starting, held, +unknown — since when, the failing streak and the restarts counted; each change is emitted as an event, and a +state that is not healthy is said again while it lasts. The controller keeps the last state per machine, +raises `module...unhealthy` when two consecutive statements say so, and clears it on the +first that does not. **The gate needs no new rule:** ADR 0236 §2's *a module's own health holds* now reads +the module's stated health — a judging is healthy only when every long-running resource of the module on that +machine is healthy, so a resource still starting is not yet a pass — and its *no condition raised since the +send* holds the module on this condition. Every start begins in `starting`, so a condition from before the +send clears at the new build's start and anything after it is the new build's. At the gate's bound the +build is put back, as ADR 0236 §3 says. + +**5. A provider down is said once, at the provider.** A check names the provision it exercises. While +that provision's provider for this consumer is unhealthy on the record, the consumer's finding is held under +the provider's condition — listed there as waiting on it, raised as nothing of its own — and the consumer's +gate waits rather than fails. What a consumer finds while its provider is healthy is its own. + +**6. Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition +reaches a person through the operator's conversation, or a healer; a healer that restarts on health is its +own decision. + +**7. A declaration is proved before it is trusted.** `module check` refuses a `health` field that names an +endpoint the module does not declare, an interval, timeout or count outside its bounds, or a tool the module +does not serve. A bed from mesh-lab, run by the catalogue's check on the build seat, starts every resource +whose declaration or image changed and requires its check to say healthy within its grace; an image check +adopted by name is proved the same way, because two of nineteen were wrong. This proves the check's +wiring — that it can see the program working — not the change, which the live mesh and the first machine's +gate still judge ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md) stands). + +**8. Every catalogue module that runs something long-lived declares one.** A catalogue-wide count of the +long-running resources without `health` may only go down. `module check` warns from this decision, and +refuses a long-running resource without `health` once the count reaches zero, or six weeks after liveness is +first judged live, whichever is first. A module running nothing long-lived — its bundle only, or files and +packages — declares none: the node tools serving its tools is its liveness, as the gate judges today. + +## Consequences + +- **Liveness alone, from the first build, would have caught the crash loop inside the gate's ten minutes.** + Readiness would have caught the silent web application in about a minute instead of eleven hours. The + identity provider's refused administrator needs the module's own tool, which it already has (ADR 0224 §5). +- **The node-engine grows a small scheduler, a field of its report and an event; the controller a condition + kind and its two-look rule; the gate a reading of state it already had a place for.** A machine of 45 + containers spends tens of milliseconds a look reading state, and about 3 s a minute on in-container + commands at the default interval. +- **A module whose first machine was unhealthy before the send no longer passes the gate for being + unchanged in its unhealthiness** — the start resets it to `starting`, and the new build is judged on its + own. +- **What a check names becomes load-bearing.** A check naming the wrong provision hides a consumer's own + fault under its provider; the bed and the first machine's gate are what catch that. +- **Harder:** a slow starter — a module that needs more than five minutes to be ready — cannot say so + within these bounds and fails its gate. That is accepted until one exists; it would be a change to the + bound, recorded. +- **A provider's health and its consumers' failures meet in two places**: this record's condition (the + provider's resource is unhealthy) and ADR 0224's standing (the provider keeps failing a consumer). They + are different facts, both kept; the hold of rule 5 reads only this record's. +- **Data is not health.** Whether a module's data is there, measured and backed up stays + [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)'s, + watched by the self-check's D13. + +## What this does not decide + +- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231). +- Health for scheduled work — whether the last scheduled run succeeded belongs to the record of the + scheduled step. +- Health of what a module's events do ([issue 276](../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md)) + — the event contract's, and the bus watchdog's. +- The core's own definitions (ADR 0236 §1, to-be 45 §8), which stand; a core component may later declare its + own through the same field. + +## How it is checked + +| Rule | Checked by | +|---|---| +| 1 Liveness without declaration | node-engine tests over a fake runtime and service manager: a container recreated keeps its counted restarts, and so does the engine restarted; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy; a resource under a maintenance step is held. A mesh-lab replay of the crash loop (a container whose program exits at start) fails its gate within the bound | +| 2 The field and its bounds | `module check` refuses each out-of-range part and each endpoint named by port or address, a test per refusal; a catalogue-wide test parses every `health` field | +| 3 The engine owns the verdict | a node-engine test that a declared command becomes the container's check and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change | +| 4 Report, event, condition, gate | controller tests: one unhealthy statement raises nothing and is listed unconfirmed; two raise; a healthy one clears; the gate fails a judging while a resource is starting or unhealthy and holds a module on this condition (the existing gate test, extended); a condition from before the send clears at the new build's start. A replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks | +| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; their gates wait, not fail; a consumer failing while its provider is healthy is raised on its own | +| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped | +| 7 Proved before trusted | the bed's run of every changed declaration in the catalogue's check; the replay of the studio's false *unhealthy* (a check that asks an address the program does not bind) fails the bed, not a machine | +| 8 Every long-running module declares | the catalogue-wide count of long-running resources without `health`, compared with the number kept in the catalogue: a merge may lower it and never raise it; after the date, `module check` refuses | +| live | every machine's report carries a state for every long-running resource; `conditions` raises `module...unhealthy` for a module stopped on purpose and clears it when it runs | + +## References + +- [Research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md) — the evidence, the + options and the recommendation this record takes. +- [To-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md), the design. +- [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md) + §2, whose module health this gives a content; + [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rules 5, 6 + and 8 and [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §2, §4 and §8 — the + condition, the second look, the witness that is never the component; + [ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md), + [ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md), + [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md), + [ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md), + [ADR 0189](0189-the-store-keeps-what-the-records-name.md), + [ADR 0149](0149-the-live-mesh-is-the-test-bed.md), + [ADR 0237](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md). +- Issues 058, 145, 179, 181, 276, 277, 281. diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index d7045644..7549f3f2 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -338,6 +338,7 @@ python3 00-META/checks/index.py fail if stale - **0232** — [A binding to a consumer's data moves only by a person](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md) - **0233** — [A module declares the data it holds, and the mesh protects and watches it from that declaration](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md) - **0235** — [The bus is backed up by its own snapshot of each stream, taken under the bus module's account](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md) +- **0240** — [A module says how it is healthy, and the node-engine judges it](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) ### How it is built diff --git a/03-DESIGN/01-to-be/32-what-a-module-declares.md b/03-DESIGN/01-to-be/32-what-a-module-declares.md index 40ddb42c..b424eaef 100644 --- a/03-DESIGN/01-to-be/32-what-a-module-declares.md +++ b/03-DESIGN/01-to-be/32-what-a-module-declares.md @@ -12,8 +12,9 @@ code: - mesh-tools src/main.ts - mesh-catalog modules/mesh-catalog - mesh-tools node-tools/internal/runtime (a module's state, ADR 0201) -updated: 2026-10-06 +updated: 2026-10-07 decisions: + - 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md - 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md - 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md - 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md @@ -284,6 +285,14 @@ module. The backup holder's composed `backup` and `data` lines are the bus-free holder reads, like every contribution (ADR 0212). The rules are in the record; `module check` holds them. +**A module says how each long-running resource is healthy.** *Added 2026-10-07, +[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).* +A container or process that stays up, and a service stated `running`, carries a field named `health`: one kind +of readiness check — the image's own adopted by name, HTTP or TCP to an endpoint named under `listens`, a +command in the container, the unit's own readiness, or one of the module's tools — with its interval, +timeout, failing looks, grace and the provision it exercises. Liveness needs no declaration. The node-engine +runs every check; `module check` holds the bounds. The design is [to-be 48](48-a-module-says-how-it-is-healthy.md). + ## 5. Seats A module declares a seat with its protocol, and the mesh enforces one holder at its scope diff --git a/03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md b/03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md index 7dada0c9..8e938154 100644 --- a/03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md +++ b/03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md @@ -4,6 +4,7 @@ status: in-progress code: [mesh-controller, mesh-host, mesh-tools, mesh-catalog, mesh-sdk, mesh-lab] updated: 2026-10-07 decisions: + - 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md - 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md - 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md - 02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md @@ -113,6 +114,11 @@ outlive the history. `bus..slow-consumer`. The last token is the kind's short word where it has one (`failing` for `provider-failing`), the kind elsewhere; a key is opaque to every reader, and the kind is the field. +*Added 2026-10-07, [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md):* +the scope `module`, for a module's own health on a machine — `module...unhealthy`, raised by +the controller on the second statement from the node-engine that a long-running resource is unhealthy, and +cleared on the first that does not. Designed in [to-be 48](48-a-module-says-how-it-is-healthy.md). + **The fields:** | Field | Holds | @@ -695,7 +701,9 @@ controller and node tools, restoring one not healthy in bound, the launcher's fo installer's grants. In mesh-catalog, the modules that keep `record` say why, and the forge's announcer says which files a merge deleted. On branches, not yet merged. **Not yet:** R3, R7, R8 on a lab mesh, and so the *done when*; container state in the node-engine's report, without which a container that -crash-loops after its compose applied is seen only through what it breaks; composing `not-reversible` for +crash-loops after its compose applied is seen only through what it breaks; *(designed 2026-10-07 by +[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) as +[to-be 48](48-a-module-says-how-it-is-healthy.md), whose phases track it)* composing `not-reversible` for a build whose migration cannot be undone; the next three core rollouts' verdicts. ### Phase 5 — Checks before merge, and the replays diff --git a/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md b/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md new file mode 100644 index 00000000..81b8e3f8 --- /dev/null +++ b/03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md @@ -0,0 +1,241 @@ +--- +layer: to-be +status: designed +code: [] +updated: 2026-10-07 +decisions: + - 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md + - 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md + - 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md +--- + +# 48 — A module says how it is healthy + +**Every long-running thing a module runs is judged alive by the node-engine on its machine, with no +declaration. Where the module declares how, it is also judged ready. The node-engine runs every check and +owns every verdict, states it in its report and on the bus, and the controller raises a condition on the +second look — which the release gate, the self-check, the healers and the operator's conversation already +act on.** ([ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).) + +This closes the gap to-be 45 names in its Phase 4: *container state in the node-engine's report*. The +core's own health definitions ([to-be 45](45-a-core-that-cannot-fail-silently.md) §8) stand beside it. + +## The parts + +``` + machine control node + ┌───────────────────────────────────────────┐ ┌───────────────────────────────────┐ + │ node-engine │ │ controller │ + │ ├─ liveness: every long-running resource │ report │ ├─ last state per machine │ + │ │ running? restarts counted (kept) │──────────► │ ├─ second statement → condition │ + │ ├─ readiness: the `health` field │ health │ │ module...unhealthy │ + │ │ http / tcp / unit ── itself │ events │ ├─ provider hold (waiting on) │ + │ │ exec / runtime ── runtime's check │──────────► │ ├─ the gate (ADR 0236 §2) │ + │ │ tool ── node tools │ │ └─ doctor, status, conditions │ + │ └─ state per resource, since, streak │ └───────────────┬───────────────────┘ + └───────────────────────────────────────────┘ │ condition events + operator's conversation (to-be 46), healers +``` + +One judge per machine, for every hosting form. Nothing else on the machine judges a module, and nothing +else sets a container's check. + +## 1. Liveness, for everything that stays up + +A **long-running resource** is a container that stays up, a process the mesh runs that stays up, or a +service unit the module states `running`. A container that runs once, a step, and anything on a schedule +are not long-running; their success is their step's and their schedule's. + +On every tick the node-engine reads each long-running resource's state from the runtime or the service +manager. A resource is **alive** when it is running and has not restarted more than once within the +**settle window** — ten minutes, the gate's bound — after its grace period. A resource that is not running +when it should be, or restarts twice in the settle window after grace, is **unhealthy**, with the reason +(`down`, `restarting`) and the restarts counted. + +**The node-engine counts restarts itself and keeps the count** in its own state on the machine, keyed by +module and resource, across recreates of the container and across its own restarts. The runtime's restart +count is lost on every recreate and its event history is about a minute long on a busy machine; neither is +read for a verdict. + +**The grace** is the resource's declared grace, or 60 s with none declared. A restart inside it is not +counted: churn that stops ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md)) +is not a crash loop. + +**A resource held still by an open maintenance step** ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)) +is `held`: neither alive nor dead, and judged again from a fresh grace when the step ends. + +## 2. Readiness, declared: the `health` field + +A long-running resource carries a field named `health`, beside its other fields. It says one **kind**, and +the timing: + +| Part | Says | Bounds | +|---|---|---| +| kind | how readiness is looked at — see below | one per resource | +| interval | how often to look | default 30 s; not under 10 s | +| timeout | how long one look may take | default 5 s; under the interval | +| failing looks | how many failing looks in a row make it unhealthy | default 3; not under 2 | +| grace | after each start, how long failure does not count | default 60 s | +| needs | the provision whose provider the check exercises, for §5 | none | + +The grace plus the failing looks at the interval are at most five minutes, so a resource broken from its +start is said within the gate's ten minutes with room for its judgings. + +**The kinds:** + +- **runtime** — the image's own check, adopted by name. A module that ships one *says* it does; an image + check a module does not adopt is not read, and is replaced by the module's own kind or none. +- **http** — a request to an endpoint the module declares under `listens`, a path, and the status + expected. +- **tcp** — a connect to an endpoint the module declares under `listens`. +- **exec** — a command run inside the container. +- **unit** — the unit's own readiness: active and not failed, and for a unit that notifies, notified. +- **tool** — one of the module's own tools, answering healthy or not with why. Only for function no endpoint + shows (the identity provider's administrator logs in; the database provider can create in a consumer's + database; the broker), and only beside a check of another kind on the same module: a module never judges + itself alone (ADR 0227 rule 8). + +**An endpoint is named, never a port or an address.** The check follows the machine's port for that endpoint +as the endpoint does; a port change moves the check with it. + +## 3. Who runs each kind + +| Kind | Run by | Read by the node-engine as | +|---|---|---| +| http, tcp | the node-engine, from the machine, to the endpoint's current port | the answer, within the timeout | +| unit | the node-engine, from the service manager | the unit's state | +| exec | the runtime: the node-engine sets the declared command as the container's check, with the declared timing | the container's health state | +| runtime | the runtime: the image's own command, with the declared timing | the container's health state | +| tool | the node tools, asked by the node-engine | the tool's answer, within the timeout | + +An HTTP or TCP check from the machine, not inside the container, tests the path a caller takes +([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)), +and costs no execution inside the container. A command is handed to the runtime so its own retries and start +period do the timing, without an execution per look from outside. + +## 4. The state, its statement, and the condition + +**Per module and long-running resource, the node-engine keeps a state:** `healthy`, `unhealthy` (with the +reason: down, restarting, or the check's failure in words), `starting` (inside grace), `held` (§1) or +`unknown` (nothing could be read). With it: since when, the failing streak, and the restarts counted. Every +start of a resource — a new build, a recreate, a restart — begins in `starting`. + +**Said three ways**, as the core already says its own state: + +- **in every report**, the current state of every resource; +- **as an event** on each change of state, on the node-engine's own subject; +- **again every minute** while a resource is not healthy, so a lost event is not a lost fault. + +**The controller keeps the last state per machine** and raises **`module...unhealthy`** +when two consecutive statements about the module say a resource is unhealthy. One statement is listed as +unconfirmed, as the self-check does a finding one look can be wrong about (to-be 45 §4). The first statement +that says no resource is unhealthy clears it. `module` joins the scopes of the condition store +(to-be 45 §2). + +- **Severity:** `warning`; `urgent` when consumers wait on it (§5), or when it has stood four hours. +- **Summary** in words, naming the module, the machine and the resource; the reason, the streak and the + restarts are evidence. Held to the content rule (ADR 0234 §6), as every condition. +- **Resolver:** `self` — it clears on observation. A healer that acts on it is a later record. + +**`status`** lists it with the other conditions; **`node show`** shows each module's resources with their +state and since when, so "is it working" has an answer without opening a terminal on the machine. + +## 5. The gate, unchanged in rule + +ADR 0236 §2 judges a module on its first machine three times, forty seconds apart, from two minutes after the +send and within ten minutes of it. Two of its points now read this design: + +- **A module's own health holds** when its tools are served (as before) **and every long-running resource of + the module on that machine is stated healthy.** `starting` is not yet a pass: a resource still in grace + makes the judging wait, not fail. +- **No condition raised since the send about the module** holds it on `module...unhealthy`. + Because every start begins in `starting`, a condition open before the send clears at the new build's start, + and one raised after it is the new build's. + +At the bound, a module not healthy is put back, as ADR 0236 §3 says. + +## 6. A provider down is said once + +A check that names, in `needs`, the provision it exercises is tied to that provision's provider for this +consumer — the provider the controller composed the consumer's grant against +([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is +about recording it beside the assignment; until then the grant is the record). While that provider's own +condition is open: + +``` + provider P unhealthy ─► module.P..unhealthy (urgent: consumers wait on it) + evidence: waiting on it — consumer A, consumer B, consumer C + consumer A failing ─► no condition of its own; its gate waits, not fails + P healthy again ─► held findings released: a consumer still failing is now its own +``` + +What a consumer finds while its provider is healthy, or by a check that names no provision, is its own. +A machine-level condition is never pinned on a module +([issue 281](../../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)). + +ADR 0224's `provider-failing` standing is a different fact — the provider failing *a consumer's +provisioning* — and stays as it is. + +## 7. Nothing restarts on health + +The runtime restarts what exits, as now. The node-engine never restarts, recreates or stops a resource for +being unhealthy. The condition reaches the operator through the conversation (to-be 46), or a healer once a +record gives one that act. + +## 8. A declaration is proved before it is trusted + +- **`module check`** refuses a `health` field naming an endpoint the module does not declare under + `listens`, a port or an address, an interval, timeout, count or grace outside its bounds, a tool the module + does not serve, or a `tool` kind with no other kind beside it on the module. +- **The bed.** mesh-lab provides a throwaway container runtime that the catalogue's check, run on the build + seat ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)), + uses to start every resource whose declaration or image changed, alone, and read its check. It must say + healthy within its grace. A check that names a provision in `needs` gets no provider on the bed, so there it + must only reach the program — an answer of any status, a connect — and is judged fully on the first + machine's gate. An adopted image check is proved the same way. +- **What the bed is not.** It proves the check's wiring — that it can see the program working — which is a + property of the image and the declaration, not of the mesh's state. Whether the change is good is judged on + the live mesh, by the first machine's gate ([ADR 0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md)). + +## 9. The migration + +**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only); 57 run +nothing long-lived and need no declaration. + +1. **Liveness first, with no declarations** (Phase A). From then every long-running resource is judged. The + catalogue's count of long-running resources without `health` is written down and starts. +2. **The seven modules whose images ship checks** adopt them by name; for the two that were wrong in the + mesh's configuration, the corrected check is what the bed proves. 19 containers are covered at once. +3. **The other container modules, and the containers without an image check in three of the seven,** declare + HTTP on their declared endpoint where they serve HTTP, TCP where they serve something else, a command where + neither shows readiness. 48 of 49 already declare the endpoint the check needs. +4. **The 19 service-only modules** declare the unit's own readiness, or TCP where the unit listens. +5. **Modules whose function no endpoint shows** add a tool check — the identity provider, the database + provider, the broker — as they are met, not in advance. +6. **The date:** when the count reaches zero, or six weeks after Phase A is live, whichever is first. From then + `module check` refuses a long-running resource without `health`. + +**Until then, and for ever for a module running nothing long-lived:** liveness of everything it runs that +stays up, and the gate's points as they stand. + +## Phases + +| Phase | Repository | Delivers | Done when | +|---|---|---|---| +| A — liveness and the statement | mesh-host (node-engine) | liveness for every long-running resource; restarts counted and kept; the state per resource in the report; the change event and its repetition; nothing restarted on health | the node-engine tests of ADR 0240 rules 1 and 6 pass; every machine's report carries a state for every long-running resource | +| | mesh-controller | the last state per machine; `module...unhealthy` on the second statement, cleared on the first that does not say it; the gate reading the stated health; `node show` and `status` | the controller tests of rule 4 pass; mesh-lab's replay of the crash loop fails its gate within the bound; a module stopped on purpose on the live mesh is raised and cleared | +| B — the field | mesh-controller | `health` parsed on every long-running resource; `module check`'s refusals; the catalogue-wide parse | a test per refusal; the whole catalogue passes `module check` | +| | mesh-host (node-engine) | the scheduler and the kinds: http, tcp, unit itself; exec and runtime as the container's check; tool through the node tools | the rule 3 tests pass; the replay of issue 145 raises the consumer within two looks | +| C — the provider hold | mesh-controller | `needs` read against the provider composed for the consumer; the consumer's finding held under the provider's condition; the consumer's gate waiting | the rule 5 test (one provider, three consumers, one condition) passes | +| D — the proof | mesh-lab, mesh-catalog | the bed; the catalogue's check starting every changed resource on it; adopted image checks proved | the replay of the studio's false *unhealthy* fails the bed, not a machine | +| E — the migration | mesh-catalog, mesh-controller | the declarations of §9 steps 2–5; the count kept in the catalogue and its test; `module check` refusing after the date | the count is zero, or the date has passed and `module check` refuses | + +Phases A and B may be built together; A is live first, because it judges without a single declaration. + +## What is not decided here + +- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231). +- Health of scheduled work, and of what a module's events do (issue 276). +- A slow starter that needs more than five minutes to be ready — a change to the bound, recorded, when one + exists. +- The core components declaring their own health through the same field.