Merge pull request 'ADR 0240 and to-be 48: a module says how it is healthy, and the node-engine judges it' (#159) from decision/0240-a-module-says-how-it-is-healthy into main

This commit was merged in pull request #159.
This commit is contained in:
2026-10-06 23:35:39 +00:00
7 changed files with 540 additions and 3 deletions
@@ -1,6 +1,7 @@
---
status: active
status: graduated
initiated: 2026-10-06
became: [02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md, 03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md]
touches:
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md
- 03-DESIGN/01-to-be/18-building-a-module.md
@@ -75,3 +76,10 @@ Evidence gathered on the live mesh, options weighed and a decision drafted, 2026
report, the condition raised by the controller on the second look, so the gate needs no new rule; a
provider down said once at the provider; a proposed decision text with how each rule is checked, and
the migration of the 68 modules.
**Graduated 2026-10-07** on the operator's word (*"yes, turn it into a decision"*), as
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) and
[to-be 48](../../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md). The record takes the proposed
decision, adding: a fifth state, `held`, for a resource under a maintenance step; a ceiling of five minutes on
grace plus failing looks, so a broken start is said inside the gate's bound; and the gate's *own health* reading
the stated health, so a resource still starting is not yet a pass.
@@ -92,6 +92,13 @@ times, at least forty seconds apart and two minutes after the send, within ten m
on one machine is judged the same way; only then does the plan go on. A policy of *together* is not
gated: it is the module saying it must change everywhere at once.
> **The mechanism changed — 2026-10-07, by [ADR 0240](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).**
> What stands: the three judgings, their spacing and bound, and the points above. What moved: *a module's own
> health* is no longer only its tools served — every long-running resource of the module on that machine must
> also be stated healthy by the node-engine, a resource still starting is not yet a pass, and an unhealthy one
> is the condition `module.<module>.<machine>.unhealthy` that *no condition raised since the send* already
> holds the module on. Designed in [to-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md).
**3. A build that fails its gate is put back there, once, and never sent again on its own.** The plan
stops; the build is marked failed at its gate; the module's registered build goes back to the build the
first machine ran before — the newest successful build of that commit still kept ([ADR 0189](0189-the-store-keeps-what-the-records-name.md)
@@ -0,0 +1,263 @@
---
topic: what runs on it
status: accepted
date: 2026-10-07
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
---
# 240. A module says how it is healthy, and the node-engine judges it
## Context
The release gate ([ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§2) judges a catalogue module on its first machine by what the mesh sees from outside: the build reported
applied, no witness put it back, no condition raised since the send about the machine or the module there,
and its tools served. None of that looks at what the module *runs*. To-be 45 names the gap in its Phase 4:
*container state in the node-engine's report, without which a container that crash-loops after its compose
applied is seen only through what it breaks.*
Measured on the live mesh for [research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md)
([evidence](../01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md)), 2026-10-07:
- **125 catalogue modules; 68 run something long-lived** — 49 a container that stays up, 19 a service
unit stated `running` and no such container. 50 run only their bundle in the node tools, 7 only files,
directories and packages.
- **No manifest can declare a health check, and nothing reads one.** The node-engine's report carries no
container or unit state; no probe of the self-check reads one.
- **19 of 73 long-running catalogue containers have an image that ships a check** (7 of 45 container
modules). The mesh never reads it. **Two of those were wrong in the mesh's hands**: the studio and the
flow editor read *unhealthy* while working, because their checks ask an address the program does not
bind in the mesh's configuration. Read without proof, they would have rolled back two good builds.
- **Every restart count is 0**, because a recreate loses it; the runtime's event history on the home
server is about a minute long, pushed out by its own check executions.
- **Of seven recent incidents**, liveness alone would have caught the agent server's crash loop (about a
hundred restarts, found by a person reading its log for something else); an HTTP readiness check the
web application that accepted TCP and answered nothing for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
only the module's own check the identity provider's refused administrator
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
A provisioner runtime that restarted until the overlay was up and then worked
([issue 058](../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md))
is the warning: a restart count without a start period reads churn that stops as a crash loop.
- **One sample is not a finding.** A single unanswered question raised an urgent alert nobody could read
([issue 277](../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md));
the self-check now raises such a finding on the second look (to-be 45 §4).
The mesh now rolls a module out on its own and puts the previous build back when the first machine is not
healthy. That promise is only as good as "healthy" is, and for a module it is judged today from the outside.
**Checked against GENESIS.** *Failure must be loud* — a crash loop behind every passing check is work that
reported success and did nothing. *Anything requiring a human to notice it will be noticed late* — eleven
hours, and a hundred restarts, each found by a person. *Evidence over assertion* — "applied" is an
assertion about a declaration; "it answers" is a measurement. *The mesh notices when something is wrong
before you do* ([effect](../00-META/effect.md)). *Long-lived user services rather than an orchestrator*
([context](../00-META/context.md)) is why this record judges and reports, and restarts nothing on health:
an orchestrator's liveness restart is the part it does not take. *The mesh is a guest on a personal node*
bounds the cost: the checks' floors below. Nothing here conflicts with GENESIS.
## Considered Options
**Who runs the checks.**
1. *The container runtime's own check, read by the node-engine.* Rejected as the whole answer: containers
only — nothing for the 19 service-only modules or anything only a module's tool knows; every look is an
execution inside the container, which on the home server already pushes every lifecycle event out of
the runtime's history; an image check the module never stated is a check nobody owns, and two of 19 were
wrong. Kept as one *kind*, adopted by name and proved.
2. *The module's own tool answers "healthy", and that is the judgement.* Rejected as the judge: a bundle is
hosted by the node tools, so a tool that answers proves the bundle up, not the server; and a component
that judges itself is what [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
rule 8 refuses for the core. Kept as one kind, for function no endpoint shows, and only beside a check
the module does not run itself.
3. *The node-engine runs everything itself, including commands, on its own schedule.* Rejected in part: a
command that only makes sense inside the container is better timed by the runtime's own retries and
start period than by an execution per look from outside.
4. **The node-engine owns every check and every verdict, and runs each kind where it is cheapest.**
Chosen: HTTP, TCP and unit checks it makes itself; a command it hands to the runtime as that
container's check and reads; a tool it asks through the node tools. One runner and one reader per
machine, for every hosting form, and it keeps what the runtime forgets.
**Where the result goes.**
1. *Only a field of the report.* Rejected alone: a report follows an apply, so a container that goes bad at
03:00 waits for the next one.
2. *Only an event on each change.* Rejected alone: events are lost or replayed; an event is a sample, not
a state.
3. *A verb the controller calls per module when it judges.* Rejected: a pull per module per judging, and a
machine slow to answer reads as unhealthy.
4. **State in the report, its change on the bus, and a condition the controller raises on the second
look.** Chosen. It is how the core already says its own state; the gate, the self-check, the healers
([ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)) and
the operator's conversation ([ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md)) all
already act on conditions.
**What is judged.**
1. *Liveness only.* Rejected: no declarations needed, but it misses issue 145 (the port was open and the
program running) and issue 179.
2. *Readiness only, where declared.* Rejected: nothing is judged until every module has declared, and
liveness alone would have caught the worst incident of the window.
3. **Liveness for every long-running resource at once, readiness where declared, the declaration required
over a migration.** Chosen: the gate means something from the first build.
**What is done with an unhealthy module.**
1. *Restart it, as an orchestrator's liveness probe does.* Rejected for this record: the evidence holds no
case where a restart would have fixed anything — the crash loop was restarting already — and a restart
hides the failure the gate is meant to see. Whether a healer restarts what stays unhealthy is left to a
record of its own under ADR 0231.
2. **Nothing restarts on health; the condition reaches a person or a healer.** Chosen.
**A provider down.** With 12 consumers of the database provision and 36 of a route:
1. *Ignore it.* Rejected: twelve conditions for one fault, and twelve gates failed for something none of
them did.
2. *Judge providers first, consumers only once their providers are healthy.* Rejected: a consumer broken on
its own is not said while its provider is down, which is when it is most needed.
3. **A consumer's check names the provision it exercises; while that provision's provider is unhealthy on
the record, the consumer's finding is held under the provider's condition and its gate waits.** Chosen.
It is how [issue 281](../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)
already treats a machine-level fault: what is the machine's is never pinned on a module.
**Where the declaration sits.**
1. *One per module.* Rejected: a module of eleven containers (the mail module) could not say which is wrong.
2. **On each long-running resource**, beside the resource's other fields. Chosen.
**The migration.**
1. *Required at once.* Rejected: 68 modules to change before the next merge, and nothing judged until all
are.
2. *Optional for ever.* Rejected: 38 of 45 container modules would stay at liveness, and a field that is
optional for ever is a field half the catalogue never gets.
3. **Liveness at once; the declaration required by a date, counted down.** Chosen.
## Decision
**1. Liveness is judged for every long-running resource, with no declaration.** A container that stays up,
a process that stays up and a service stated `running` are *alive* when running and not restarted more than
once within the settle window after their grace period. The node-engine observes this itself on each tick
and keeps the restarts it counted across recreates and across its own restarts; it never relies on the
runtime's restart count or event history. A resource held still by an open maintenance step
([ADR 0189](0189-the-store-keeps-what-the-records-name.md)) is neither alive nor dead: it is said as held,
and judged again when the step ends.
**2. A module declares how each long-running resource is ready**, in a field named `health` on that
resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint and the
status expected, a TCP connect to a declared endpoint, a command in the container, the unit's own
readiness, or one of the module's tools — with an interval (default 30 s, **not under 10 s**), a timeout
under the interval, a number of failing looks in a row before it is unhealthy (default 3, **not under 2**),
and a grace period after a start (default 60 s), in which failure does not count. The grace and the failing
looks together are at most five minutes, so a resource broken from its start is said within the gate's
bound. An endpoint is named by its `listens` name, never by a port or an address, so a check follows the
machine's ports as the endpoint does. A check of function that no endpoint shows is a module's own tool, and
only in addition to a check the module does not run itself.
**3. The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes itself;
a command in the container it hands to the runtime as that container's check and reads the state; an
adopted image check it reads the same way; a tool it asks through the node tools. Nothing else on the machine
judges a module, and nothing else sets a container's check.
**4. The state goes in the report, its change on the bus, and a condition is the controller's.** Each
report carries, per module and long-running resource, a state — healthy, unhealthy, starting, held,
unknown — since when, the failing streak and the restarts counted; each change is emitted as an event, and a
state that is not healthy is said again while it lasts. The controller keeps the last state per machine,
raises `module.<module>.<machine>.unhealthy` when two consecutive statements say so, and clears it on the
first that does not. **The gate needs no new rule:** ADR 0236 §2's *a module's own health holds* now reads
the module's stated health — a judging is healthy only when every long-running resource of the module on that
machine is healthy, so a resource still starting is not yet a pass — and its *no condition raised since the
send* holds the module on this condition. Every start begins in `starting`, so a condition from before the
send clears at the new build's start and anything after it is the new build's. At the gate's bound the
build is put back, as ADR 0236 §3 says.
**5. A provider down is said once, at the provider.** A check names the provision it exercises. While
that provision's provider for this consumer is unhealthy on the record, the consumer's finding is held under
the provider's condition — listed there as waiting on it, raised as nothing of its own — and the consumer's
gate waits rather than fails. What a consumer finds while its provider is healthy is its own.
**6. Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition
reaches a person through the operator's conversation, or a healer; a healer that restarts on health is its
own decision.
**7. A declaration is proved before it is trusted.** `module check` refuses a `health` field that names an
endpoint the module does not declare, an interval, timeout or count outside its bounds, or a tool the module
does not serve. A bed from mesh-lab, run by the catalogue's check on the build seat, starts every resource
whose declaration or image changed and requires its check to say healthy within its grace; an image check
adopted by name is proved the same way, because two of nineteen were wrong. This proves the check's
wiring — that it can see the program working — not the change, which the live mesh and the first machine's
gate still judge ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md) stands).
**8. Every catalogue module that runs something long-lived declares one.** A catalogue-wide count of the
long-running resources without `health` may only go down. `module check` warns from this decision, and
refuses a long-running resource without `health` once the count reaches zero, or six weeks after liveness is
first judged live, whichever is first. A module running nothing long-lived — its bundle only, or files and
packages — declares none: the node tools serving its tools is its liveness, as the gate judges today.
## Consequences
- **Liveness alone, from the first build, would have caught the crash loop inside the gate's ten minutes.**
Readiness would have caught the silent web application in about a minute instead of eleven hours. The
identity provider's refused administrator needs the module's own tool, which it already has (ADR 0224 §5).
- **The node-engine grows a small scheduler, a field of its report and an event; the controller a condition
kind and its two-look rule; the gate a reading of state it already had a place for.** A machine of 45
containers spends tens of milliseconds a look reading state, and about 3 s a minute on in-container
commands at the default interval.
- **A module whose first machine was unhealthy before the send no longer passes the gate for being
unchanged in its unhealthiness** — the start resets it to `starting`, and the new build is judged on its
own.
- **What a check names becomes load-bearing.** A check naming the wrong provision hides a consumer's own
fault under its provider; the bed and the first machine's gate are what catch that.
- **Harder:** a slow starter — a module that needs more than five minutes to be ready — cannot say so
within these bounds and fails its gate. That is accepted until one exists; it would be a change to the
bound, recorded.
- **A provider's health and its consumers' failures meet in two places**: this record's condition (the
provider's resource is unhealthy) and ADR 0224's standing (the provider keeps failing a consumer). They
are different facts, both kept; the hold of rule 5 reads only this record's.
- **Data is not health.** Whether a module's data is there, measured and backed up stays
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)'s,
watched by the self-check's D13.
## What this does not decide
- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231).
- Health for scheduled work — whether the last scheduled run succeeded belongs to the record of the
scheduled step.
- Health of what a module's events do ([issue 276](../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md))
— the event contract's, and the bus watchdog's.
- The core's own definitions (ADR 0236 §1, to-be 45 §8), which stand; a core component may later declare its
own through the same field.
## How it is checked
| Rule | Checked by |
|---|---|
| 1 Liveness without declaration | node-engine tests over a fake runtime and service manager: a container recreated keeps its counted restarts, and so does the engine restarted; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy; a resource under a maintenance step is held. A mesh-lab replay of the crash loop (a container whose program exits at start) fails its gate within the bound |
| 2 The field and its bounds | `module check` refuses each out-of-range part and each endpoint named by port or address, a test per refusal; a catalogue-wide test parses every `health` field |
| 3 The engine owns the verdict | a node-engine test that a declared command becomes the container's check and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change |
| 4 Report, event, condition, gate | controller tests: one unhealthy statement raises nothing and is listed unconfirmed; two raise; a healthy one clears; the gate fails a judging while a resource is starting or unhealthy and holds a module on this condition (the existing gate test, extended); a condition from before the send clears at the new build's start. A replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks |
| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; their gates wait, not fail; a consumer failing while its provider is healthy is raised on its own |
| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped |
| 7 Proved before trusted | the bed's run of every changed declaration in the catalogue's check; the replay of the studio's false *unhealthy* (a check that asks an address the program does not bind) fails the bed, not a machine |
| 8 Every long-running module declares | the catalogue-wide count of long-running resources without `health`, compared with the number kept in the catalogue: a merge may lower it and never raise it; after the date, `module check` refuses |
| live | every machine's report carries a state for every long-running resource; `conditions` raises `module.<module>.<machine>.unhealthy` for a module stopped on purpose and clears it when it runs |
## References
- [Research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md) — the evidence, the
options and the recommendation this record takes.
- [To-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md), the design.
- [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§2, whose module health this gives a content;
[ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rules 5, 6
and 8 and [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §2, §4 and §8 — the
condition, the second look, the witness that is never the component;
[ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md),
[ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md),
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md),
[ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md),
[ADR 0189](0189-the-store-keeps-what-the-records-name.md),
[ADR 0149](0149-the-live-mesh-is-the-test-bed.md),
[ADR 0237](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md).
- Issues 058, 145, 179, 181, 276, 277, 281.
+1
View File
@@ -338,6 +338,7 @@ python3 00-META/checks/index.py fail if stale
- **0232** — [A binding to a consumer's data moves only by a person](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md)
- **0233** — [A module declares the data it holds, and the mesh protects and watches it from that declaration](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
- **0235** — [The bus is backed up by its own snapshot of each stream, taken under the bus module's account](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)
- **0240** — [A module says how it is healthy, and the node-engine judges it](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
### How it is built
@@ -12,8 +12,9 @@ code:
- mesh-tools src/main.ts
- mesh-catalog modules/mesh-catalog
- mesh-tools node-tools/internal/runtime (a module's state, ADR 0201)
updated: 2026-10-06
updated: 2026-10-07
decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
@@ -284,6 +285,14 @@ module. The backup holder's composed `backup` and `data` lines are the bus-free
holder reads, like every contribution (ADR 0212). The rules are in the record; `module check` holds
them.
**A module says how each long-running resource is healthy.** *Added 2026-10-07,
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).*
A container or process that stays up, and a service stated `running`, carries a field named `health`: one kind
of readiness check — the image's own adopted by name, HTTP or TCP to an endpoint named under `listens`, a
command in the container, the unit's own readiness, or one of the module's tools — with its interval,
timeout, failing looks, grace and the provision it exercises. Liveness needs no declaration. The node-engine
runs every check; `module check` holds the bounds. The design is [to-be 48](48-a-module-says-how-it-is-healthy.md).
## 5. Seats
A module declares a seat with its protocol, and the mesh enforces one holder at its scope
@@ -4,6 +4,7 @@ status: in-progress
code: [mesh-controller, mesh-host, mesh-tools, mesh-catalog, mesh-sdk, mesh-lab]
updated: 2026-10-07
decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md
- 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
- 02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md
@@ -113,6 +114,11 @@ outlive the history.
`bus.<consumer>.slow-consumer`. The last token is the kind's short word where it has one (`failing` for
`provider-failing`), the kind elsewhere; a key is opaque to every reader, and the kind is the field.
*Added 2026-10-07, [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md):*
the scope `module`, for a module's own health on a machine — `module.<module>.<machine>.unhealthy`, raised by
the controller on the second statement from the node-engine that a long-running resource is unhealthy, and
cleared on the first that does not. Designed in [to-be 48](48-a-module-says-how-it-is-healthy.md).
**The fields:**
| Field | Holds |
@@ -695,7 +701,9 @@ controller and node tools, restoring one not healthy in bound, the launcher's fo
installer's grants. In mesh-catalog, the modules that keep `record` say why, and the forge's announcer
says which files a merge deleted. On branches, not yet merged. **Not yet:** R3, R7, R8 on a lab mesh, and
so the *done when*; container state in the node-engine's report, without which a container that
crash-loops after its compose applied is seen only through what it breaks; composing `not-reversible` for
crash-loops after its compose applied is seen only through what it breaks; *(designed 2026-10-07 by
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) as
[to-be 48](48-a-module-says-how-it-is-healthy.md), whose phases track it)* composing `not-reversible` for
a build whose migration cannot be undone; the next three core rollouts' verdicts.
### Phase 5 — Checks before merge, and the replays
@@ -0,0 +1,241 @@
---
layer: to-be
status: designed
code: []
updated: 2026-10-07
decisions:
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 48 — A module says how it is healthy
**Every long-running thing a module runs is judged alive by the node-engine on its machine, with no
declaration. Where the module declares how, it is also judged ready. The node-engine runs every check and
owns every verdict, states it in its report and on the bus, and the controller raises a condition on the
second look — which the release gate, the self-check, the healers and the operator's conversation already
act on.** ([ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).)
This closes the gap to-be 45 names in its Phase 4: *container state in the node-engine's report*. The
core's own health definitions ([to-be 45](45-a-core-that-cannot-fail-silently.md) §8) stand beside it.
## The parts
```
machine control node
┌───────────────────────────────────────────┐ ┌───────────────────────────────────┐
│ node-engine │ │ controller │
│ ├─ liveness: every long-running resource │ report │ ├─ last state per machine │
│ │ running? restarts counted (kept) │──────────► │ ├─ second statement → condition │
│ ├─ readiness: the `health` field │ health │ │ module.<m>.<machine>.unhealthy │
│ │ http / tcp / unit ── itself │ events │ ├─ provider hold (waiting on) │
│ │ exec / runtime ── runtime's check │──────────► │ ├─ the gate (ADR 0236 §2) │
│ │ tool ── node tools │ │ └─ doctor, status, conditions │
│ └─ state per resource, since, streak │ └───────────────┬───────────────────┘
└───────────────────────────────────────────┘ │ condition events
operator's conversation (to-be 46), healers
```
One judge per machine, for every hosting form. Nothing else on the machine judges a module, and nothing
else sets a container's check.
## 1. Liveness, for everything that stays up
A **long-running resource** is a container that stays up, a process the mesh runs that stays up, or a
service unit the module states `running`. A container that runs once, a step, and anything on a schedule
are not long-running; their success is their step's and their schedule's.
On every tick the node-engine reads each long-running resource's state from the runtime or the service
manager. A resource is **alive** when it is running and has not restarted more than once within the
**settle window** — ten minutes, the gate's bound — after its grace period. A resource that is not running
when it should be, or restarts twice in the settle window after grace, is **unhealthy**, with the reason
(`down`, `restarting`) and the restarts counted.
**The node-engine counts restarts itself and keeps the count** in its own state on the machine, keyed by
module and resource, across recreates of the container and across its own restarts. The runtime's restart
count is lost on every recreate and its event history is about a minute long on a busy machine; neither is
read for a verdict.
**The grace** is the resource's declared grace, or 60 s with none declared. A restart inside it is not
counted: churn that stops ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md))
is not a crash loop.
**A resource held still by an open maintenance step** ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
is `held`: neither alive nor dead, and judged again from a fresh grace when the step ends.
## 2. Readiness, declared: the `health` field
A long-running resource carries a field named `health`, beside its other fields. It says one **kind**, and
the timing:
| Part | Says | Bounds |
|---|---|---|
| kind | how readiness is looked at — see below | one per resource |
| interval | how often to look | default 30 s; not under 10 s |
| timeout | how long one look may take | default 5 s; under the interval |
| failing looks | how many failing looks in a row make it unhealthy | default 3; not under 2 |
| grace | after each start, how long failure does not count | default 60 s |
| needs | the provision whose provider the check exercises, for §5 | none |
The grace plus the failing looks at the interval are at most five minutes, so a resource broken from its
start is said within the gate's ten minutes with room for its judgings.
**The kinds:**
- **runtime** — the image's own check, adopted by name. A module that ships one *says* it does; an image
check a module does not adopt is not read, and is replaced by the module's own kind or none.
- **http** — a request to an endpoint the module declares under `listens`, a path, and the status
expected.
- **tcp** — a connect to an endpoint the module declares under `listens`.
- **exec** — a command run inside the container.
- **unit** — the unit's own readiness: active and not failed, and for a unit that notifies, notified.
- **tool** — one of the module's own tools, answering healthy or not with why. Only for function no endpoint
shows (the identity provider's administrator logs in; the database provider can create in a consumer's
database; the broker), and only beside a check of another kind on the same module: a module never judges
itself alone (ADR 0227 rule 8).
**An endpoint is named, never a port or an address.** The check follows the machine's port for that endpoint
as the endpoint does; a port change moves the check with it.
## 3. Who runs each kind
| Kind | Run by | Read by the node-engine as |
|---|---|---|
| http, tcp | the node-engine, from the machine, to the endpoint's current port | the answer, within the timeout |
| unit | the node-engine, from the service manager | the unit's state |
| exec | the runtime: the node-engine sets the declared command as the container's check, with the declared timing | the container's health state |
| runtime | the runtime: the image's own command, with the declared timing | the container's health state |
| tool | the node tools, asked by the node-engine | the tool's answer, within the timeout |
An HTTP or TCP check from the machine, not inside the container, tests the path a caller takes
([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)),
and costs no execution inside the container. A command is handed to the runtime so its own retries and start
period do the timing, without an execution per look from outside.
## 4. The state, its statement, and the condition
**Per module and long-running resource, the node-engine keeps a state:** `healthy`, `unhealthy` (with the
reason: down, restarting, or the check's failure in words), `starting` (inside grace), `held` (§1) or
`unknown` (nothing could be read). With it: since when, the failing streak, and the restarts counted. Every
start of a resource — a new build, a recreate, a restart — begins in `starting`.
**Said three ways**, as the core already says its own state:
- **in every report**, the current state of every resource;
- **as an event** on each change of state, on the node-engine's own subject;
- **again every minute** while a resource is not healthy, so a lost event is not a lost fault.
**The controller keeps the last state per machine** and raises **`module.<module>.<machine>.unhealthy`**
when two consecutive statements about the module say a resource is unhealthy. One statement is listed as
unconfirmed, as the self-check does a finding one look can be wrong about (to-be 45 §4). The first statement
that says no resource is unhealthy clears it. `module` joins the scopes of the condition store
(to-be 45 §2).
- **Severity:** `warning`; `urgent` when consumers wait on it (§5), or when it has stood four hours.
- **Summary** in words, naming the module, the machine and the resource; the reason, the streak and the
restarts are evidence. Held to the content rule (ADR 0234 §6), as every condition.
- **Resolver:** `self` — it clears on observation. A healer that acts on it is a later record.
**`status`** lists it with the other conditions; **`node show`** shows each module's resources with their
state and since when, so "is it working" has an answer without opening a terminal on the machine.
## 5. The gate, unchanged in rule
ADR 0236 §2 judges a module on its first machine three times, forty seconds apart, from two minutes after the
send and within ten minutes of it. Two of its points now read this design:
- **A module's own health holds** when its tools are served (as before) **and every long-running resource of
the module on that machine is stated healthy.** `starting` is not yet a pass: a resource still in grace
makes the judging wait, not fail.
- **No condition raised since the send about the module** holds it on `module.<module>.<machine>.unhealthy`.
Because every start begins in `starting`, a condition open before the send clears at the new build's start,
and one raised after it is the new build's.
At the bound, a module not healthy is put back, as ADR 0236 §3 says.
## 6. A provider down is said once
A check that names, in `needs`, the provision it exercises is tied to that provision's provider for this
consumer — the provider the controller composed the consumer's grant against
([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is
about recording it beside the assignment; until then the grant is the record). While that provider's own
condition is open:
```
provider P unhealthy ─► module.P.<machine>.unhealthy (urgent: consumers wait on it)
evidence: waiting on it — consumer A, consumer B, consumer C
consumer A failing ─► no condition of its own; its gate waits, not fails
P healthy again ─► held findings released: a consumer still failing is now its own
```
What a consumer finds while its provider is healthy, or by a check that names no provision, is its own.
A machine-level condition is never pinned on a module
([issue 281](../../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)).
ADR 0224's `provider-failing` standing is a different fact — the provider failing *a consumer's
provisioning* — and stays as it is.
## 7. Nothing restarts on health
The runtime restarts what exits, as now. The node-engine never restarts, recreates or stops a resource for
being unhealthy. The condition reaches the operator through the conversation (to-be 46), or a healer once a
record gives one that act.
## 8. A declaration is proved before it is trusted
- **`module check`** refuses a `health` field naming an endpoint the module does not declare under
`listens`, a port or an address, an interval, timeout, count or grace outside its bounds, a tool the module
does not serve, or a `tool` kind with no other kind beside it on the module.
- **The bed.** mesh-lab provides a throwaway container runtime that the catalogue's check, run on the build
seat ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)),
uses to start every resource whose declaration or image changed, alone, and read its check. It must say
healthy within its grace. A check that names a provision in `needs` gets no provider on the bed, so there it
must only reach the program — an answer of any status, a connect — and is judged fully on the first
machine's gate. An adopted image check is proved the same way.
- **What the bed is not.** It proves the check's wiring — that it can see the program working — which is a
property of the image and the declaration, not of the mesh's state. Whether the change is good is judged on
the live mesh, by the first machine's gate ([ADR 0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md)).
## 9. The migration
**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only); 57 run
nothing long-lived and need no declaration.
1. **Liveness first, with no declarations** (Phase A). From then every long-running resource is judged. The
catalogue's count of long-running resources without `health` is written down and starts.
2. **The seven modules whose images ship checks** adopt them by name; for the two that were wrong in the
mesh's configuration, the corrected check is what the bed proves. 19 containers are covered at once.
3. **The other container modules, and the containers without an image check in three of the seven,** declare
HTTP on their declared endpoint where they serve HTTP, TCP where they serve something else, a command where
neither shows readiness. 48 of 49 already declare the endpoint the check needs.
4. **The 19 service-only modules** declare the unit's own readiness, or TCP where the unit listens.
5. **Modules whose function no endpoint shows** add a tool check — the identity provider, the database
provider, the broker — as they are met, not in advance.
6. **The date:** when the count reaches zero, or six weeks after Phase A is live, whichever is first. From then
`module check` refuses a long-running resource without `health`.
**Until then, and for ever for a module running nothing long-lived:** liveness of everything it runs that
stays up, and the gate's points as they stand.
## Phases
| Phase | Repository | Delivers | Done when |
|---|---|---|---|
| A — liveness and the statement | mesh-host (node-engine) | liveness for every long-running resource; restarts counted and kept; the state per resource in the report; the change event and its repetition; nothing restarted on health | the node-engine tests of ADR 0240 rules 1 and 6 pass; every machine's report carries a state for every long-running resource |
| | mesh-controller | the last state per machine; `module.<module>.<machine>.unhealthy` on the second statement, cleared on the first that does not say it; the gate reading the stated health; `node show` and `status` | the controller tests of rule 4 pass; mesh-lab's replay of the crash loop fails its gate within the bound; a module stopped on purpose on the live mesh is raised and cleared |
| B — the field | mesh-controller | `health` parsed on every long-running resource; `module check`'s refusals; the catalogue-wide parse | a test per refusal; the whole catalogue passes `module check` |
| | mesh-host (node-engine) | the scheduler and the kinds: http, tcp, unit itself; exec and runtime as the container's check; tool through the node tools | the rule 3 tests pass; the replay of issue 145 raises the consumer within two looks |
| C — the provider hold | mesh-controller | `needs` read against the provider composed for the consumer; the consumer's finding held under the provider's condition; the consumer's gate waiting | the rule 5 test (one provider, three consumers, one condition) passes |
| D — the proof | mesh-lab, mesh-catalog | the bed; the catalogue's check starting every changed resource on it; adopted image checks proved | the replay of the studio's false *unhealthy* fails the bed, not a machine |
| E — the migration | mesh-catalog, mesh-controller | the declarations of §9 steps 2–5; the count kept in the catalogue and its test; `module check` refusing after the date | the count is zero, or the date has passed and `module check` refuses |
Phases A and B may be built together; A is live first, because it judges without a single declaration.
## What is not decided here
- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).
- Health of scheduled work, and of what a module's events do (issue 276).
- A slow starter that needs more than five minutes to be ready — a change to the bound, recorded, when one
exists.
- The core components declaring their own health through the same field.