ADR 0240 and to-be 48: a module says how it is healthy, and the node-engine judges it
The gate judged a module by what the mesh sees from outside, so a crash loop and an eleven-hour silent web app passed it. Graduates research 032 on the operator's word.
This commit is contained in:
@@ -1,6 +1,7 @@
|
||||
---
|
||||
status: active
|
||||
status: graduated
|
||||
initiated: 2026-10-06
|
||||
became: [02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md, 03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md]
|
||||
touches:
|
||||
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md
|
||||
- 03-DESIGN/01-to-be/18-building-a-module.md
|
||||
@@ -75,3 +76,10 @@ Evidence gathered on the live mesh, options weighed and a decision drafted, 2026
|
||||
report, the condition raised by the controller on the second look, so the gate needs no new rule; a
|
||||
provider down said once at the provider; a proposed decision text with how each rule is checked, and
|
||||
the migration of the 68 modules.
|
||||
|
||||
**Graduated 2026-10-07** on the operator's word (*"yes, turn it into a decision"*), as
|
||||
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) and
|
||||
[to-be 48](../../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md). The record takes the proposed
|
||||
decision, adding: a fifth state, `held`, for a resource under a maintenance step; a ceiling of five minutes on
|
||||
grace plus failing looks, so a broken start is said inside the gate's bound; and the gate's *own health* reading
|
||||
the stated health, so a resource still starting is not yet a pass.
|
||||
|
||||
+7
@@ -92,6 +92,13 @@ times, at least forty seconds apart and two minutes after the send, within ten m
|
||||
on one machine is judged the same way; only then does the plan go on. A policy of *together* is not
|
||||
gated: it is the module saying it must change everywhere at once.
|
||||
|
||||
> **The mechanism changed — 2026-10-07, by [ADR 0240](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).**
|
||||
> What stands: the three judgings, their spacing and bound, and the points above. What moved: *a module's own
|
||||
> health* is no longer only its tools served — every long-running resource of the module on that machine must
|
||||
> also be stated healthy by the node-engine, a resource still starting is not yet a pass, and an unhealthy one
|
||||
> is the condition `module.<module>.<machine>.unhealthy` that *no condition raised since the send* already
|
||||
> holds the module on. Designed in [to-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md).
|
||||
|
||||
**3. A build that fails its gate is put back there, once, and never sent again on its own.** The plan
|
||||
stops; the build is marked failed at its gate; the module's registered build goes back to the build the
|
||||
first machine ran before — the newest successful build of that commit still kept ([ADR 0189](0189-the-store-keeps-what-the-records-name.md)
|
||||
|
||||
@@ -0,0 +1,263 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
date: 2026-10-07
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
|
||||
---
|
||||
|
||||
# 240. A module says how it is healthy, and the node-engine judges it
|
||||
|
||||
## Context
|
||||
|
||||
The release gate ([ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
|
||||
§2) judges a catalogue module on its first machine by what the mesh sees from outside: the build reported
|
||||
applied, no witness put it back, no condition raised since the send about the machine or the module there,
|
||||
and its tools served. None of that looks at what the module *runs*. To-be 45 names the gap in its Phase 4:
|
||||
*container state in the node-engine's report, without which a container that crash-loops after its compose
|
||||
applied is seen only through what it breaks.*
|
||||
|
||||
Measured on the live mesh for [research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md)
|
||||
([evidence](../01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md)), 2026-10-07:
|
||||
|
||||
- **125 catalogue modules; 68 run something long-lived** — 49 a container that stays up, 19 a service
|
||||
unit stated `running` and no such container. 50 run only their bundle in the node tools, 7 only files,
|
||||
directories and packages.
|
||||
- **No manifest can declare a health check, and nothing reads one.** The node-engine's report carries no
|
||||
container or unit state; no probe of the self-check reads one.
|
||||
- **19 of 73 long-running catalogue containers have an image that ships a check** (7 of 45 container
|
||||
modules). The mesh never reads it. **Two of those were wrong in the mesh's hands**: the studio and the
|
||||
flow editor read *unhealthy* while working, because their checks ask an address the program does not
|
||||
bind in the mesh's configuration. Read without proof, they would have rolled back two good builds.
|
||||
- **Every restart count is 0**, because a recreate loses it; the runtime's event history on the home
|
||||
server is about a minute long, pushed out by its own check executions.
|
||||
- **Of seven recent incidents**, liveness alone would have caught the agent server's crash loop (about a
|
||||
hundred restarts, found by a person reading its log for something else); an HTTP readiness check the
|
||||
web application that accepted TCP and answered nothing for eleven hours
|
||||
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
|
||||
only the module's own check the identity provider's refused administrator
|
||||
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
|
||||
A provisioner runtime that restarted until the overlay was up and then worked
|
||||
([issue 058](../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md))
|
||||
is the warning: a restart count without a start period reads churn that stops as a crash loop.
|
||||
- **One sample is not a finding.** A single unanswered question raised an urgent alert nobody could read
|
||||
([issue 277](../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md));
|
||||
the self-check now raises such a finding on the second look (to-be 45 §4).
|
||||
|
||||
The mesh now rolls a module out on its own and puts the previous build back when the first machine is not
|
||||
healthy. That promise is only as good as "healthy" is, and for a module it is judged today from the outside.
|
||||
|
||||
**Checked against GENESIS.** *Failure must be loud* — a crash loop behind every passing check is work that
|
||||
reported success and did nothing. *Anything requiring a human to notice it will be noticed late* — eleven
|
||||
hours, and a hundred restarts, each found by a person. *Evidence over assertion* — "applied" is an
|
||||
assertion about a declaration; "it answers" is a measurement. *The mesh notices when something is wrong
|
||||
before you do* ([effect](../00-META/effect.md)). *Long-lived user services rather than an orchestrator*
|
||||
([context](../00-META/context.md)) is why this record judges and reports, and restarts nothing on health:
|
||||
an orchestrator's liveness restart is the part it does not take. *The mesh is a guest on a personal node*
|
||||
bounds the cost: the checks' floors below. Nothing here conflicts with GENESIS.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**Who runs the checks.**
|
||||
|
||||
1. *The container runtime's own check, read by the node-engine.* Rejected as the whole answer: containers
|
||||
only — nothing for the 19 service-only modules or anything only a module's tool knows; every look is an
|
||||
execution inside the container, which on the home server already pushes every lifecycle event out of
|
||||
the runtime's history; an image check the module never stated is a check nobody owns, and two of 19 were
|
||||
wrong. Kept as one *kind*, adopted by name and proved.
|
||||
2. *The module's own tool answers "healthy", and that is the judgement.* Rejected as the judge: a bundle is
|
||||
hosted by the node tools, so a tool that answers proves the bundle up, not the server; and a component
|
||||
that judges itself is what [ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
|
||||
rule 8 refuses for the core. Kept as one kind, for function no endpoint shows, and only beside a check
|
||||
the module does not run itself.
|
||||
3. *The node-engine runs everything itself, including commands, on its own schedule.* Rejected in part: a
|
||||
command that only makes sense inside the container is better timed by the runtime's own retries and
|
||||
start period than by an execution per look from outside.
|
||||
4. **The node-engine owns every check and every verdict, and runs each kind where it is cheapest.**
|
||||
Chosen: HTTP, TCP and unit checks it makes itself; a command it hands to the runtime as that
|
||||
container's check and reads; a tool it asks through the node tools. One runner and one reader per
|
||||
machine, for every hosting form, and it keeps what the runtime forgets.
|
||||
|
||||
**Where the result goes.**
|
||||
|
||||
1. *Only a field of the report.* Rejected alone: a report follows an apply, so a container that goes bad at
|
||||
03:00 waits for the next one.
|
||||
2. *Only an event on each change.* Rejected alone: events are lost or replayed; an event is a sample, not
|
||||
a state.
|
||||
3. *A verb the controller calls per module when it judges.* Rejected: a pull per module per judging, and a
|
||||
machine slow to answer reads as unhealthy.
|
||||
4. **State in the report, its change on the bus, and a condition the controller raises on the second
|
||||
look.** Chosen. It is how the core already says its own state; the gate, the self-check, the healers
|
||||
([ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)) and
|
||||
the operator's conversation ([ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md)) all
|
||||
already act on conditions.
|
||||
|
||||
**What is judged.**
|
||||
|
||||
1. *Liveness only.* Rejected: no declarations needed, but it misses issue 145 (the port was open and the
|
||||
program running) and issue 179.
|
||||
2. *Readiness only, where declared.* Rejected: nothing is judged until every module has declared, and
|
||||
liveness alone would have caught the worst incident of the window.
|
||||
3. **Liveness for every long-running resource at once, readiness where declared, the declaration required
|
||||
over a migration.** Chosen: the gate means something from the first build.
|
||||
|
||||
**What is done with an unhealthy module.**
|
||||
|
||||
1. *Restart it, as an orchestrator's liveness probe does.* Rejected for this record: the evidence holds no
|
||||
case where a restart would have fixed anything — the crash loop was restarting already — and a restart
|
||||
hides the failure the gate is meant to see. Whether a healer restarts what stays unhealthy is left to a
|
||||
record of its own under ADR 0231.
|
||||
2. **Nothing restarts on health; the condition reaches a person or a healer.** Chosen.
|
||||
|
||||
**A provider down.** With 12 consumers of the database provision and 36 of a route:
|
||||
|
||||
1. *Ignore it.* Rejected: twelve conditions for one fault, and twelve gates failed for something none of
|
||||
them did.
|
||||
2. *Judge providers first, consumers only once their providers are healthy.* Rejected: a consumer broken on
|
||||
its own is not said while its provider is down, which is when it is most needed.
|
||||
3. **A consumer's check names the provision it exercises; while that provision's provider is unhealthy on
|
||||
the record, the consumer's finding is held under the provider's condition and its gate waits.** Chosen.
|
||||
It is how [issue 281](../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)
|
||||
already treats a machine-level fault: what is the machine's is never pinned on a module.
|
||||
|
||||
**Where the declaration sits.**
|
||||
|
||||
1. *One per module.* Rejected: a module of eleven containers (the mail module) could not say which is wrong.
|
||||
2. **On each long-running resource**, beside the resource's other fields. Chosen.
|
||||
|
||||
**The migration.**
|
||||
|
||||
1. *Required at once.* Rejected: 68 modules to change before the next merge, and nothing judged until all
|
||||
are.
|
||||
2. *Optional for ever.* Rejected: 38 of 45 container modules would stay at liveness, and a field that is
|
||||
optional for ever is a field half the catalogue never gets.
|
||||
3. **Liveness at once; the declaration required by a date, counted down.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. Liveness is judged for every long-running resource, with no declaration.** A container that stays up,
|
||||
a process that stays up and a service stated `running` are *alive* when running and not restarted more than
|
||||
once within the settle window after their grace period. The node-engine observes this itself on each tick
|
||||
and keeps the restarts it counted across recreates and across its own restarts; it never relies on the
|
||||
runtime's restart count or event history. A resource held still by an open maintenance step
|
||||
([ADR 0189](0189-the-store-keeps-what-the-records-name.md)) is neither alive nor dead: it is said as held,
|
||||
and judged again when the step ends.
|
||||
|
||||
**2. A module declares how each long-running resource is ready**, in a field named `health` on that
|
||||
resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint and the
|
||||
status expected, a TCP connect to a declared endpoint, a command in the container, the unit's own
|
||||
readiness, or one of the module's tools — with an interval (default 30 s, **not under 10 s**), a timeout
|
||||
under the interval, a number of failing looks in a row before it is unhealthy (default 3, **not under 2**),
|
||||
and a grace period after a start (default 60 s), in which failure does not count. The grace and the failing
|
||||
looks together are at most five minutes, so a resource broken from its start is said within the gate's
|
||||
bound. An endpoint is named by its `listens` name, never by a port or an address, so a check follows the
|
||||
machine's ports as the endpoint does. A check of function that no endpoint shows is a module's own tool, and
|
||||
only in addition to a check the module does not run itself.
|
||||
|
||||
**3. The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes itself;
|
||||
a command in the container it hands to the runtime as that container's check and reads the state; an
|
||||
adopted image check it reads the same way; a tool it asks through the node tools. Nothing else on the machine
|
||||
judges a module, and nothing else sets a container's check.
|
||||
|
||||
**4. The state goes in the report, its change on the bus, and a condition is the controller's.** Each
|
||||
report carries, per module and long-running resource, a state — healthy, unhealthy, starting, held,
|
||||
unknown — since when, the failing streak and the restarts counted; each change is emitted as an event, and a
|
||||
state that is not healthy is said again while it lasts. The controller keeps the last state per machine,
|
||||
raises `module.<module>.<machine>.unhealthy` when two consecutive statements say so, and clears it on the
|
||||
first that does not. **The gate needs no new rule:** ADR 0236 §2's *a module's own health holds* now reads
|
||||
the module's stated health — a judging is healthy only when every long-running resource of the module on that
|
||||
machine is healthy, so a resource still starting is not yet a pass — and its *no condition raised since the
|
||||
send* holds the module on this condition. Every start begins in `starting`, so a condition from before the
|
||||
send clears at the new build's start and anything after it is the new build's. At the gate's bound the
|
||||
build is put back, as ADR 0236 §3 says.
|
||||
|
||||
**5. A provider down is said once, at the provider.** A check names the provision it exercises. While
|
||||
that provision's provider for this consumer is unhealthy on the record, the consumer's finding is held under
|
||||
the provider's condition — listed there as waiting on it, raised as nothing of its own — and the consumer's
|
||||
gate waits rather than fails. What a consumer finds while its provider is healthy is its own.
|
||||
|
||||
**6. Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition
|
||||
reaches a person through the operator's conversation, or a healer; a healer that restarts on health is its
|
||||
own decision.
|
||||
|
||||
**7. A declaration is proved before it is trusted.** `module check` refuses a `health` field that names an
|
||||
endpoint the module does not declare, an interval, timeout or count outside its bounds, or a tool the module
|
||||
does not serve. A bed from mesh-lab, run by the catalogue's check on the build seat, starts every resource
|
||||
whose declaration or image changed and requires its check to say healthy within its grace; an image check
|
||||
adopted by name is proved the same way, because two of nineteen were wrong. This proves the check's
|
||||
wiring — that it can see the program working — not the change, which the live mesh and the first machine's
|
||||
gate still judge ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md) stands).
|
||||
|
||||
**8. Every catalogue module that runs something long-lived declares one.** A catalogue-wide count of the
|
||||
long-running resources without `health` may only go down. `module check` warns from this decision, and
|
||||
refuses a long-running resource without `health` once the count reaches zero, or six weeks after liveness is
|
||||
first judged live, whichever is first. A module running nothing long-lived — its bundle only, or files and
|
||||
packages — declares none: the node tools serving its tools is its liveness, as the gate judges today.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Liveness alone, from the first build, would have caught the crash loop inside the gate's ten minutes.**
|
||||
Readiness would have caught the silent web application in about a minute instead of eleven hours. The
|
||||
identity provider's refused administrator needs the module's own tool, which it already has (ADR 0224 §5).
|
||||
- **The node-engine grows a small scheduler, a field of its report and an event; the controller a condition
|
||||
kind and its two-look rule; the gate a reading of state it already had a place for.** A machine of 45
|
||||
containers spends tens of milliseconds a look reading state, and about 3 s a minute on in-container
|
||||
commands at the default interval.
|
||||
- **A module whose first machine was unhealthy before the send no longer passes the gate for being
|
||||
unchanged in its unhealthiness** — the start resets it to `starting`, and the new build is judged on its
|
||||
own.
|
||||
- **What a check names becomes load-bearing.** A check naming the wrong provision hides a consumer's own
|
||||
fault under its provider; the bed and the first machine's gate are what catch that.
|
||||
- **Harder:** a slow starter — a module that needs more than five minutes to be ready — cannot say so
|
||||
within these bounds and fails its gate. That is accepted until one exists; it would be a change to the
|
||||
bound, recorded.
|
||||
- **A provider's health and its consumers' failures meet in two places**: this record's condition (the
|
||||
provider's resource is unhealthy) and ADR 0224's standing (the provider keeps failing a consumer). They
|
||||
are different facts, both kept; the hold of rule 5 reads only this record's.
|
||||
- **Data is not health.** Whether a module's data is there, measured and backed up stays
|
||||
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)'s,
|
||||
watched by the self-check's D13.
|
||||
|
||||
## What this does not decide
|
||||
|
||||
- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231).
|
||||
- Health for scheduled work — whether the last scheduled run succeeded belongs to the record of the
|
||||
scheduled step.
|
||||
- Health of what a module's events do ([issue 276](../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md))
|
||||
— the event contract's, and the bus watchdog's.
|
||||
- The core's own definitions (ADR 0236 §1, to-be 45 §8), which stand; a core component may later declare its
|
||||
own through the same field.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| 1 Liveness without declaration | node-engine tests over a fake runtime and service manager: a container recreated keeps its counted restarts, and so does the engine restarted; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy; a resource under a maintenance step is held. A mesh-lab replay of the crash loop (a container whose program exits at start) fails its gate within the bound |
|
||||
| 2 The field and its bounds | `module check` refuses each out-of-range part and each endpoint named by port or address, a test per refusal; a catalogue-wide test parses every `health` field |
|
||||
| 3 The engine owns the verdict | a node-engine test that a declared command becomes the container's check and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change |
|
||||
| 4 Report, event, condition, gate | controller tests: one unhealthy statement raises nothing and is listed unconfirmed; two raise; a healthy one clears; the gate fails a judging while a resource is starting or unhealthy and holds a module on this condition (the existing gate test, extended); a condition from before the send clears at the new build's start. A replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks |
|
||||
| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; their gates wait, not fail; a consumer failing while its provider is healthy is raised on its own |
|
||||
| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped |
|
||||
| 7 Proved before trusted | the bed's run of every changed declaration in the catalogue's check; the replay of the studio's false *unhealthy* (a check that asks an address the program does not bind) fails the bed, not a machine |
|
||||
| 8 Every long-running module declares | the catalogue-wide count of long-running resources without `health`, compared with the number kept in the catalogue: a merge may lower it and never raise it; after the date, `module check` refuses |
|
||||
| live | every machine's report carries a state for every long-running resource; `conditions` raises `module.<module>.<machine>.unhealthy` for a module stopped on purpose and clears it when it runs |
|
||||
|
||||
## References
|
||||
|
||||
- [Research 032](../01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md) — the evidence, the
|
||||
options and the recommendation this record takes.
|
||||
- [To-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md), the design.
|
||||
- [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
|
||||
§2, whose module health this gives a content;
|
||||
[ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md) rules 5, 6
|
||||
and 8 and [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §2, §4 and §8 — the
|
||||
condition, the second look, the witness that is never the component;
|
||||
[ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md),
|
||||
[ADR 0231](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md),
|
||||
[ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md),
|
||||
[ADR 0234](0234-the-mesh-holds-a-conversation-with-its-operator.md),
|
||||
[ADR 0189](0189-the-store-keeps-what-the-records-name.md),
|
||||
[ADR 0149](0149-the-live-mesh-is-the-test-bed.md),
|
||||
[ADR 0237](0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md).
|
||||
- Issues 058, 145, 179, 181, 276, 277, 281.
|
||||
@@ -338,6 +338,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0232** — [A binding to a consumer's data moves only by a person](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md)
|
||||
- **0233** — [A module declares the data it holds, and the mesh protects and watches it from that declaration](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
|
||||
- **0235** — [The bus is backed up by its own snapshot of each stream, taken under the bus module's account](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)
|
||||
- **0240** — [A module says how it is healthy, and the node-engine judges it](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
|
||||
|
||||
### How it is built
|
||||
|
||||
|
||||
@@ -12,8 +12,9 @@ code:
|
||||
- mesh-tools src/main.ts
|
||||
- mesh-catalog modules/mesh-catalog
|
||||
- mesh-tools node-tools/internal/runtime (a module's state, ADR 0201)
|
||||
updated: 2026-10-06
|
||||
updated: 2026-10-07
|
||||
decisions:
|
||||
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
|
||||
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
|
||||
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
|
||||
- 02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md
|
||||
@@ -284,6 +285,14 @@ module. The backup holder's composed `backup` and `data` lines are the bus-free
|
||||
holder reads, like every contribution (ADR 0212). The rules are in the record; `module check` holds
|
||||
them.
|
||||
|
||||
**A module says how each long-running resource is healthy.** *Added 2026-10-07,
|
||||
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).*
|
||||
A container or process that stays up, and a service stated `running`, carries a field named `health`: one kind
|
||||
of readiness check — the image's own adopted by name, HTTP or TCP to an endpoint named under `listens`, a
|
||||
command in the container, the unit's own readiness, or one of the module's tools — with its interval,
|
||||
timeout, failing looks, grace and the provision it exercises. Liveness needs no declaration. The node-engine
|
||||
runs every check; `module check` holds the bounds. The design is [to-be 48](48-a-module-says-how-it-is-healthy.md).
|
||||
|
||||
## 5. Seats
|
||||
|
||||
A module declares a seat with its protocol, and the mesh enforces one holder at its scope
|
||||
|
||||
@@ -4,6 +4,7 @@ status: in-progress
|
||||
code: [mesh-controller, mesh-host, mesh-tools, mesh-catalog, mesh-sdk, mesh-lab]
|
||||
updated: 2026-10-07
|
||||
decisions:
|
||||
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
|
||||
- 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md
|
||||
- 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
|
||||
- 02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md
|
||||
@@ -113,6 +114,11 @@ outlive the history.
|
||||
`bus.<consumer>.slow-consumer`. The last token is the kind's short word where it has one (`failing` for
|
||||
`provider-failing`), the kind elsewhere; a key is opaque to every reader, and the kind is the field.
|
||||
|
||||
*Added 2026-10-07, [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md):*
|
||||
the scope `module`, for a module's own health on a machine — `module.<module>.<machine>.unhealthy`, raised by
|
||||
the controller on the second statement from the node-engine that a long-running resource is unhealthy, and
|
||||
cleared on the first that does not. Designed in [to-be 48](48-a-module-says-how-it-is-healthy.md).
|
||||
|
||||
**The fields:**
|
||||
|
||||
| Field | Holds |
|
||||
@@ -695,7 +701,9 @@ controller and node tools, restoring one not healthy in bound, the launcher's fo
|
||||
installer's grants. In mesh-catalog, the modules that keep `record` say why, and the forge's announcer
|
||||
says which files a merge deleted. On branches, not yet merged. **Not yet:** R3, R7, R8 on a lab mesh, and
|
||||
so the *done when*; container state in the node-engine's report, without which a container that
|
||||
crash-loops after its compose applied is seen only through what it breaks; composing `not-reversible` for
|
||||
crash-loops after its compose applied is seen only through what it breaks; *(designed 2026-10-07 by
|
||||
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) as
|
||||
[to-be 48](48-a-module-says-how-it-is-healthy.md), whose phases track it)* composing `not-reversible` for
|
||||
a build whose migration cannot be undone; the next three core rollouts' verdicts.
|
||||
|
||||
### Phase 5 — Checks before merge, and the replays
|
||||
|
||||
@@ -0,0 +1,241 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: []
|
||||
updated: 2026-10-07
|
||||
decisions:
|
||||
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
|
||||
- 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
|
||||
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
|
||||
---
|
||||
|
||||
# 48 — A module says how it is healthy
|
||||
|
||||
**Every long-running thing a module runs is judged alive by the node-engine on its machine, with no
|
||||
declaration. Where the module declares how, it is also judged ready. The node-engine runs every check and
|
||||
owns every verdict, states it in its report and on the bus, and the controller raises a condition on the
|
||||
second look — which the release gate, the self-check, the healers and the operator's conversation already
|
||||
act on.** ([ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md).)
|
||||
|
||||
This closes the gap to-be 45 names in its Phase 4: *container state in the node-engine's report*. The
|
||||
core's own health definitions ([to-be 45](45-a-core-that-cannot-fail-silently.md) §8) stand beside it.
|
||||
|
||||
## The parts
|
||||
|
||||
```
|
||||
machine control node
|
||||
┌───────────────────────────────────────────┐ ┌───────────────────────────────────┐
|
||||
│ node-engine │ │ controller │
|
||||
│ ├─ liveness: every long-running resource │ report │ ├─ last state per machine │
|
||||
│ │ running? restarts counted (kept) │──────────► │ ├─ second statement → condition │
|
||||
│ ├─ readiness: the `health` field │ health │ │ module.<m>.<machine>.unhealthy │
|
||||
│ │ http / tcp / unit ── itself │ events │ ├─ provider hold (waiting on) │
|
||||
│ │ exec / runtime ── runtime's check │──────────► │ ├─ the gate (ADR 0236 §2) │
|
||||
│ │ tool ── node tools │ │ └─ doctor, status, conditions │
|
||||
│ └─ state per resource, since, streak │ └───────────────┬───────────────────┘
|
||||
└───────────────────────────────────────────┘ │ condition events
|
||||
operator's conversation (to-be 46), healers
|
||||
```
|
||||
|
||||
One judge per machine, for every hosting form. Nothing else on the machine judges a module, and nothing
|
||||
else sets a container's check.
|
||||
|
||||
## 1. Liveness, for everything that stays up
|
||||
|
||||
A **long-running resource** is a container that stays up, a process the mesh runs that stays up, or a
|
||||
service unit the module states `running`. A container that runs once, a step, and anything on a schedule
|
||||
are not long-running; their success is their step's and their schedule's.
|
||||
|
||||
On every tick the node-engine reads each long-running resource's state from the runtime or the service
|
||||
manager. A resource is **alive** when it is running and has not restarted more than once within the
|
||||
**settle window** — ten minutes, the gate's bound — after its grace period. A resource that is not running
|
||||
when it should be, or restarts twice in the settle window after grace, is **unhealthy**, with the reason
|
||||
(`down`, `restarting`) and the restarts counted.
|
||||
|
||||
**The node-engine counts restarts itself and keeps the count** in its own state on the machine, keyed by
|
||||
module and resource, across recreates of the container and across its own restarts. The runtime's restart
|
||||
count is lost on every recreate and its event history is about a minute long on a busy machine; neither is
|
||||
read for a verdict.
|
||||
|
||||
**The grace** is the resource's declared grace, or 60 s with none declared. A restart inside it is not
|
||||
counted: churn that stops ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md))
|
||||
is not a crash loop.
|
||||
|
||||
**A resource held still by an open maintenance step** ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
|
||||
is `held`: neither alive nor dead, and judged again from a fresh grace when the step ends.
|
||||
|
||||
## 2. Readiness, declared: the `health` field
|
||||
|
||||
A long-running resource carries a field named `health`, beside its other fields. It says one **kind**, and
|
||||
the timing:
|
||||
|
||||
| Part | Says | Bounds |
|
||||
|---|---|---|
|
||||
| kind | how readiness is looked at — see below | one per resource |
|
||||
| interval | how often to look | default 30 s; not under 10 s |
|
||||
| timeout | how long one look may take | default 5 s; under the interval |
|
||||
| failing looks | how many failing looks in a row make it unhealthy | default 3; not under 2 |
|
||||
| grace | after each start, how long failure does not count | default 60 s |
|
||||
| needs | the provision whose provider the check exercises, for §5 | none |
|
||||
|
||||
The grace plus the failing looks at the interval are at most five minutes, so a resource broken from its
|
||||
start is said within the gate's ten minutes with room for its judgings.
|
||||
|
||||
**The kinds:**
|
||||
|
||||
- **runtime** — the image's own check, adopted by name. A module that ships one *says* it does; an image
|
||||
check a module does not adopt is not read, and is replaced by the module's own kind or none.
|
||||
- **http** — a request to an endpoint the module declares under `listens`, a path, and the status
|
||||
expected.
|
||||
- **tcp** — a connect to an endpoint the module declares under `listens`.
|
||||
- **exec** — a command run inside the container.
|
||||
- **unit** — the unit's own readiness: active and not failed, and for a unit that notifies, notified.
|
||||
- **tool** — one of the module's own tools, answering healthy or not with why. Only for function no endpoint
|
||||
shows (the identity provider's administrator logs in; the database provider can create in a consumer's
|
||||
database; the broker), and only beside a check of another kind on the same module: a module never judges
|
||||
itself alone (ADR 0227 rule 8).
|
||||
|
||||
**An endpoint is named, never a port or an address.** The check follows the machine's port for that endpoint
|
||||
as the endpoint does; a port change moves the check with it.
|
||||
|
||||
## 3. Who runs each kind
|
||||
|
||||
| Kind | Run by | Read by the node-engine as |
|
||||
|---|---|---|
|
||||
| http, tcp | the node-engine, from the machine, to the endpoint's current port | the answer, within the timeout |
|
||||
| unit | the node-engine, from the service manager | the unit's state |
|
||||
| exec | the runtime: the node-engine sets the declared command as the container's check, with the declared timing | the container's health state |
|
||||
| runtime | the runtime: the image's own command, with the declared timing | the container's health state |
|
||||
| tool | the node tools, asked by the node-engine | the tool's answer, within the timeout |
|
||||
|
||||
An HTTP or TCP check from the machine, not inside the container, tests the path a caller takes
|
||||
([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)),
|
||||
and costs no execution inside the container. A command is handed to the runtime so its own retries and start
|
||||
period do the timing, without an execution per look from outside.
|
||||
|
||||
## 4. The state, its statement, and the condition
|
||||
|
||||
**Per module and long-running resource, the node-engine keeps a state:** `healthy`, `unhealthy` (with the
|
||||
reason: down, restarting, or the check's failure in words), `starting` (inside grace), `held` (§1) or
|
||||
`unknown` (nothing could be read). With it: since when, the failing streak, and the restarts counted. Every
|
||||
start of a resource — a new build, a recreate, a restart — begins in `starting`.
|
||||
|
||||
**Said three ways**, as the core already says its own state:
|
||||
|
||||
- **in every report**, the current state of every resource;
|
||||
- **as an event** on each change of state, on the node-engine's own subject;
|
||||
- **again every minute** while a resource is not healthy, so a lost event is not a lost fault.
|
||||
|
||||
**The controller keeps the last state per machine** and raises **`module.<module>.<machine>.unhealthy`**
|
||||
when two consecutive statements about the module say a resource is unhealthy. One statement is listed as
|
||||
unconfirmed, as the self-check does a finding one look can be wrong about (to-be 45 §4). The first statement
|
||||
that says no resource is unhealthy clears it. `module` joins the scopes of the condition store
|
||||
(to-be 45 §2).
|
||||
|
||||
- **Severity:** `warning`; `urgent` when consumers wait on it (§5), or when it has stood four hours.
|
||||
- **Summary** in words, naming the module, the machine and the resource; the reason, the streak and the
|
||||
restarts are evidence. Held to the content rule (ADR 0234 §6), as every condition.
|
||||
- **Resolver:** `self` — it clears on observation. A healer that acts on it is a later record.
|
||||
|
||||
**`status`** lists it with the other conditions; **`node show`** shows each module's resources with their
|
||||
state and since when, so "is it working" has an answer without opening a terminal on the machine.
|
||||
|
||||
## 5. The gate, unchanged in rule
|
||||
|
||||
ADR 0236 §2 judges a module on its first machine three times, forty seconds apart, from two minutes after the
|
||||
send and within ten minutes of it. Two of its points now read this design:
|
||||
|
||||
- **A module's own health holds** when its tools are served (as before) **and every long-running resource of
|
||||
the module on that machine is stated healthy.** `starting` is not yet a pass: a resource still in grace
|
||||
makes the judging wait, not fail.
|
||||
- **No condition raised since the send about the module** holds it on `module.<module>.<machine>.unhealthy`.
|
||||
Because every start begins in `starting`, a condition open before the send clears at the new build's start,
|
||||
and one raised after it is the new build's.
|
||||
|
||||
At the bound, a module not healthy is put back, as ADR 0236 §3 says.
|
||||
|
||||
## 6. A provider down is said once
|
||||
|
||||
A check that names, in `needs`, the provision it exercises is tied to that provision's provider for this
|
||||
consumer — the provider the controller composed the consumer's grant against
|
||||
([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is
|
||||
about recording it beside the assignment; until then the grant is the record). While that provider's own
|
||||
condition is open:
|
||||
|
||||
```
|
||||
provider P unhealthy ─► module.P.<machine>.unhealthy (urgent: consumers wait on it)
|
||||
evidence: waiting on it — consumer A, consumer B, consumer C
|
||||
consumer A failing ─► no condition of its own; its gate waits, not fails
|
||||
P healthy again ─► held findings released: a consumer still failing is now its own
|
||||
```
|
||||
|
||||
What a consumer finds while its provider is healthy, or by a check that names no provision, is its own.
|
||||
A machine-level condition is never pinned on a module
|
||||
([issue 281](../../04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md)).
|
||||
|
||||
ADR 0224's `provider-failing` standing is a different fact — the provider failing *a consumer's
|
||||
provisioning* — and stays as it is.
|
||||
|
||||
## 7. Nothing restarts on health
|
||||
|
||||
The runtime restarts what exits, as now. The node-engine never restarts, recreates or stops a resource for
|
||||
being unhealthy. The condition reaches the operator through the conversation (to-be 46), or a healer once a
|
||||
record gives one that act.
|
||||
|
||||
## 8. A declaration is proved before it is trusted
|
||||
|
||||
- **`module check`** refuses a `health` field naming an endpoint the module does not declare under
|
||||
`listens`, a port or an address, an interval, timeout, count or grace outside its bounds, a tool the module
|
||||
does not serve, or a `tool` kind with no other kind beside it on the module.
|
||||
- **The bed.** mesh-lab provides a throwaway container runtime that the catalogue's check, run on the build
|
||||
seat ([ADR 0237](../../02-DECISIONS/0237-a-change-is-judged-against-the-mesh-that-runs-before-it-merges-on-the-build-seat.md)),
|
||||
uses to start every resource whose declaration or image changed, alone, and read its check. It must say
|
||||
healthy within its grace. A check that names a provision in `needs` gets no provider on the bed, so there it
|
||||
must only reach the program — an answer of any status, a connect — and is judged fully on the first
|
||||
machine's gate. An adopted image check is proved the same way.
|
||||
- **What the bed is not.** It proves the check's wiring — that it can see the program working — which is a
|
||||
property of the image and the declaration, not of the mesh's state. Whether the change is good is judged on
|
||||
the live mesh, by the first machine's gate ([ADR 0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md)).
|
||||
|
||||
## 9. The migration
|
||||
|
||||
**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only); 57 run
|
||||
nothing long-lived and need no declaration.
|
||||
|
||||
1. **Liveness first, with no declarations** (Phase A). From then every long-running resource is judged. The
|
||||
catalogue's count of long-running resources without `health` is written down and starts.
|
||||
2. **The seven modules whose images ship checks** adopt them by name; for the two that were wrong in the
|
||||
mesh's configuration, the corrected check is what the bed proves. 19 containers are covered at once.
|
||||
3. **The other container modules, and the containers without an image check in three of the seven,** declare
|
||||
HTTP on their declared endpoint where they serve HTTP, TCP where they serve something else, a command where
|
||||
neither shows readiness. 48 of 49 already declare the endpoint the check needs.
|
||||
4. **The 19 service-only modules** declare the unit's own readiness, or TCP where the unit listens.
|
||||
5. **Modules whose function no endpoint shows** add a tool check — the identity provider, the database
|
||||
provider, the broker — as they are met, not in advance.
|
||||
6. **The date:** when the count reaches zero, or six weeks after Phase A is live, whichever is first. From then
|
||||
`module check` refuses a long-running resource without `health`.
|
||||
|
||||
**Until then, and for ever for a module running nothing long-lived:** liveness of everything it runs that
|
||||
stays up, and the gate's points as they stand.
|
||||
|
||||
## Phases
|
||||
|
||||
| Phase | Repository | Delivers | Done when |
|
||||
|---|---|---|---|
|
||||
| A — liveness and the statement | mesh-host (node-engine) | liveness for every long-running resource; restarts counted and kept; the state per resource in the report; the change event and its repetition; nothing restarted on health | the node-engine tests of ADR 0240 rules 1 and 6 pass; every machine's report carries a state for every long-running resource |
|
||||
| | mesh-controller | the last state per machine; `module.<module>.<machine>.unhealthy` on the second statement, cleared on the first that does not say it; the gate reading the stated health; `node show` and `status` | the controller tests of rule 4 pass; mesh-lab's replay of the crash loop fails its gate within the bound; a module stopped on purpose on the live mesh is raised and cleared |
|
||||
| B — the field | mesh-controller | `health` parsed on every long-running resource; `module check`'s refusals; the catalogue-wide parse | a test per refusal; the whole catalogue passes `module check` |
|
||||
| | mesh-host (node-engine) | the scheduler and the kinds: http, tcp, unit itself; exec and runtime as the container's check; tool through the node tools | the rule 3 tests pass; the replay of issue 145 raises the consumer within two looks |
|
||||
| C — the provider hold | mesh-controller | `needs` read against the provider composed for the consumer; the consumer's finding held under the provider's condition; the consumer's gate waiting | the rule 5 test (one provider, three consumers, one condition) passes |
|
||||
| D — the proof | mesh-lab, mesh-catalog | the bed; the catalogue's check starting every changed resource on it; adopted image checks proved | the replay of the studio's false *unhealthy* fails the bed, not a machine |
|
||||
| E — the migration | mesh-catalog, mesh-controller | the declarations of §9 steps 2–5; the count kept in the catalogue and its test; `module check` refusing after the date | the count is zero, or the date has passed and `module check` refuses |
|
||||
|
||||
Phases A and B may be built together; A is live first, because it judges without a single declaration.
|
||||
|
||||
## What is not decided here
|
||||
|
||||
- Whether a healer restarts what stays unhealthy (its own record, under ADR 0231).
|
||||
- Health of scheduled work, and of what a module's events do (issue 276).
|
||||
- A slow starter that needs more than five minutes to be ready — a change to the bound, recorded, when one
|
||||
exists.
|
||||
- The core components declaring their own health through the same field.
|
||||
Reference in New Issue
Block a user