Research 032: a module says how it is healthy

The release gate judges a module from outside (applied, no new condition,
tools answer); a container that restarts in a loop or a service that fails
its own work passes it. Investigate a health declaration in every manifest.
This commit is contained in:
jochen
2026-10-06 20:50:25 +02:00
parent 68e432ec0b
commit 5e54632626
@@ -0,0 +1,59 @@
---
status: active
initiated: 2026-10-06
touches:
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md
- 03-DESIGN/01-to-be/18-building-a-module.md
- the module manifest
- the node-engine
- the release gate
---
# 032 — A module says how it is healthy
## What is investigated
Whether, and how, every module should declare **how its own health is checked**, as a field of its
manifest, so the mesh can judge a module the way it already judges its core parts.
Today the release gate ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)) judges a
module on its first machine by what the mesh can see from outside: the declaration applied, no new
condition about the module or its machine since the send, and its tools answering. It does not see
whether the module's own service works. A container that applied cleanly and then restarts in a loop,
a web application whose port is open but whose pages fail, a database that accepts connections and
refuses queries — each passes the gate and is caught only through what it breaks later, if anything
notices at all. To-be 45 names this gap ("no container state in the gate").
The questions:
1. **What a health declaration says.** The kinds of check a module may declare (a container's own
health state, an HTTP request and the answer expected, a TCP connect, a command run inside the
service, a query, a tool of the module's own that answers "healthy"), its interval, its timeout,
how many failures in a row count, and a start period during which failure does not count.
2. **Who runs the checks, and where the result goes.** The node-engine on the machine, the node tools,
or the module itself; how the result reaches the controller (the report, an event, a seat verb); and
what the release gate, the self-check and the healers do with it.
3. **What the mesh already has to build on.** Container runtimes' own healthchecks, which many images
ship and which the mesh neither reads nor sets today; systemd's own state for services; the tools a
module already serves.
4. **What "healthy" covers.** Liveness (it runs), readiness (it serves), and whether a module's health
may depend on what it requires (a provider down makes its consumers unhealthy — said once, at the
provider, not once per consumer).
5. **The cost and the noise.** How often checks may run across all modules on a small machine, and how
a check avoids the single-sample flaw the self-check already met (issue 277).
6. **Migration.** How every catalogue module gets a declaration, what a module without one is judged
by, and whether `module check` should require one.
## Why
The mesh now rolls a module out on its own, one machine first, and puts the previous build back when
the first machine is not healthy. That promise is only as good as "healthy" is, and for a module it is
today judged from the outside. Making each module say how its health is checked turns "applied" into
"working", for the gate, for the self-check and for whoever asks.
## What it touches
The module manifest and `module check`; the node-engine, which would run or read the checks; the
release gate and the doctor probes of to-be 45; the healers, which may restart what stays unhealthy;
the operator's conversation (ADR 0234), which carries what stays unhealthy; and every catalogue module,
each of which would gain a declaration.