Research 032: measure module health on the live mesh and propose a decision
mesh/merge-gate pass: the change touches no module of the mesh's graph
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered

Liveness is judged nowhere and readiness cannot be declared; the evidence,
options and a graduation-ready recommendation say how a module states it.
This commit is contained in:
jochen
2026-10-07 01:18:00 +02:00
parent 4405a3048c
commit ef6ab814b0
4 changed files with 475 additions and 0 deletions
@@ -57,3 +57,21 @@ The module manifest and `module check`; the node-engine, which would run or read
release gate and the doctor probes of to-be 45; the healers, which may restart what stays unhealthy;
the operator's conversation (ADR 0234), which carries what stays unhealthy; and every catalogue module,
each of which would gain a declaration.
## Where it stands
Evidence gathered on the live mesh, options weighed and a decision drafted, 2026-10-07:
- [01 — The evidence](01-evidence.md): 125 catalogue modules, 68 of them running something long-lived
(49 a container, 19 a service); no manifest can declare a check and nothing reads one; 19 of 73
long-running catalogue containers carry an image check, two of which were wrong in the mesh's hands;
every restart count is 0 because a recreate loses it, and the runtime's event history is about a minute
long. Of seven recent incidents, liveness alone would have caught the crash loop, readiness the
eleven-hour silent web app, and only a module's own tool the refused identity-provider admin.
- [02 — Options](02-options.md): who runs the checks, where the result goes, liveness and readiness,
dependency-aware health, the field's shape and the migration.
- [03 — Recommendation](03-recommendation.md): liveness judged for every long-running resource at once;
readiness declared per resource in a field named `health` and run by the node-engine; the state in the
report, the condition raised by the controller on the second look, so the gate needs no new rule; a
provider down said once at the provider; a proposed decision text with how each rule is checked, and
the migration of the 68 modules.
@@ -0,0 +1,196 @@
# 01 — The evidence
Measured on 2026-10-07 on the live mesh of four machines (a home server, a control node, a workstation
and a laptop), read only: the three catalogue repositories at their trunk, the controller's and the
node-engine's source at their trunk, each machine's container runtime and service manager, and the
controller's own answers. Counts, not anecdotes. Machines are named by role.
## 1. What the catalogue runs
The three catalogue repositories hold **125 module manifests**: 113 in the main catalogue, 12 in the
media catalogue. The photos repository holds application code and no manifest; its two instances are
modules in the main catalogue.
Each manifest sorted by the **longest-lived thing it runs** (a container that stays up first, then a
service the manifest says is running, then the mesh's own process that stays up, then a bundle only):
| What a module runs | Modules |
|---|---|
| at least one container that stays up | **49** |
| no such container, but a service unit stated `running` | **19** |
| only its bundle — tools and handlers hosted by the node tools | **50** |
| only files, directories and packages | **7** |
Underneath:
| Resource | Count | Of which |
|---|---|---|
| container | **85** in 50 modules | 77 stay up, 6 run once (a step), 2 on a schedule; **71 distinct images** |
| service (an existing unit put in a state) | **31** in 21 modules | 26 stated `running`, 5 stateless (the machine's lifecycle) |
| process (the mesh's own code in a unit it writes) | **20** in 18 modules | 17 run once, 2 on a schedule, **1** stays up |
| bundle artifact | **101** in 99 modules | |
And what a check could be built from:
- **50** modules declare `listens` (an endpoint, a port, a protocol) — **48 of the 49** container
modules. A TCP or HTTP check has its target named already.
- **53** modules declare `tools`. **One** tool in the whole catalogue is a health tool by name
(`dbus_health`); **18** modules have a tool named `…_status`.
- **67** modules `require` something; the most-required provisions are `route` (36 consumers),
`x11-display` (15) and `postgres-database` (12). A database down is, today, potentially twelve
consumers failing at once.
- **No manifest declares a health check.** The container resource has no field for one — its fields are
image, environment, files, ports, volumes, arguments, names, networks, capabilities, logging, the
restart triggers and the run-once and schedule modes — and the node-engine passes the runtime no
health option. Whatever check runs is the image's own, run by the runtime by default.
## 2. What the images already ship, and what the runtime says today
Every container on every machine inspected — its state, its health, its restart count, and its
**image's** healthcheck from the image configuration:
| | |
|---|---|
| containers on the four machines | **115** (94 running) |
| of those, declared by the catalogue and running | **79 instances** of **73** of the 77 long-running containers (4 are assigned nowhere) |
| long-running catalogue containers whose **image ships a HEALTHCHECK** | **19 of 73 (26 %)** |
| …in how many modules | **7 of 45** modules with a running container (mail, the hosted database suite, a spreadsheet app, the chat client, a flow editor, the media server, the certificate authority) |
| …modules whose every container has one | **4** |
| running container modules with **no health state at all** | **38 of 45** |
| catalogue containers reporting `healthy` / `unhealthy` / `starting` | **19 / 0 / 0** |
| containers reporting `unhealthy` | 5, all exited, all on the workstation, all predating the mesh (none declared) |
**Two of the nineteen image checks were wrong under the mesh's own configuration** until a catalogue fix:
- the hosted database suite's studio — the framework binds the address the runtime puts in `HOSTNAME`, so
it answered only on its network address while its image's check asked `localhost`: *unhealthy for ever
while working* (fixed 2026-10-05);
- the flow editor — the image's check reads its settings from a path the module mounted elsewhere: *ran
fine but reported unhealthy forever* (fixed 2026-09-30).
So 2 of 19 shipped checks (≈ 10 %) gave a false *unhealthy* in the mesh's hands. Read blindly, they would
have failed two good builds at the gate.
The intervals the images chose range from **2 s to 60 s**: a database every 2 s, three web apps every
5 s, an API gateway every 10 s, the rest 30 s or 60 s. Start periods range from none to 360 s (the
antivirus). Retries from 3 to 20.
### The restart count the runtime keeps is lost
**All 94 running containers report a restart count of 0** — including the agent server that, per
[issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md), had restarted about
a hundred times in a crash loop days before. The counter belongs to a container, and the node-engine
recreates a container whenever its declaration or a file it reads changes; the fix recreated it and the
history went with it. A restart count is only evidence if something outside the container keeps it.
### The runtime's event history is about a minute long
The runtime keeps its most recent ~250 events in memory. On the home server, a query for the last 30
minutes and for the last minute both answered ~250 events, **every one an `exec_*` event** — the image
health checks running (84 creates, 85 dies, 84 starts). A 24-hour query for container lifecycle events,
through the mesh's own `docker_events` tool, answered **zero**. The images' checks crowd every lifecycle
event out of the history within a minute. A container that died an hour ago leaves no trace there.
## 3. What the service manager says today
- **Failed system units:** 0 on three machines; 2 on the workstation, both mounts that the mesh does not
declare. One failed user unit on each of two machines, neither the mesh's.
- The mesh's own units (`mesh-ca-trust`, `mesh-filter`, the power units, the controller, the node tools,
the node-engine): all `active`, **`NRestarts` 0**.
- systemd already gives, per unit and for free: `ActiveState`/`SubState`, `is-failed`, `NRestarts`,
the time it entered its state, and — through a unit's own `ExecStartPost`/`WatchdogSec` where software
supports it — readiness. None of it is read by the mesh for a module's unit.
## 4. What the node-engine reports today
The node-engine's report (its `Report` message) carries: what applied and what failed per resource,
refusals, the declaration's digest and order, held and stray containers, reachable sockets on an adopted
machine, packet filters and the firewall found, windows (a scheduled step holding containers still), the
machine's profile, the engine's own build, outward links, rollbacks its witnesses decided, and whether it
reads `witness`. Separately, an `Alive` heartbeat with its interval.
**It carries no container state, no unit state and no health.** The controller's `node` answer for the
home server shows the consequence: its assigned modules, its strays, its filters and its capabilities —
and nothing about whether any of its 45 catalogue containers is running.
The node-engine has one judging mechanism already: the **witness** (to-be 45 §8) — a process's `witness`
field, `lease` or `ping` or `none`, judges a new build of the controller or the node tools and restores
the previous one when it is not healthy in bound. Only the two core processes use it.
A **run-once step** is already a health gate in practice: a step exiting non-zero fails the apply and
everything placed after it. **23** steps exist (6 containers, 17 processes); one of them, in the hosted
database suite, is a loop that waits up to five minutes for its analytics service's `/health` before the
services that need it start.
## 5. What the release gate judges a module by
The gate on a plan's first machine ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md);
`judgeHealth` in the controller) passes a module there when, on three judgings at least 40 s apart over at
least two minutes, within ten minutes of the send:
1. no witness on the machine put a core build back since the send;
2. the machine's last report is current and **applied** (failed or refused is broken);
3. **no condition** raised since the send names the module on that machine, or the machine as a whole
(issue 281 sorted the latter out);
4. for the controller, the node-engine and the node tools: their own health definitions (lease held and
ready; the engine's build reported; the node tools answering the bus);
5. for any other module **that declares tools**: the machine's node tools serve them.
That is the whole of it for a catalogue module. For the 49 container modules nothing in 1–5 looks at a
container: point 5 asks the node tools, which host the module's bundle, not its container. **A container
that applied and then crash-loops passes all five.** To-be 45's Phase 4 note says so: *"container state in
the node-engine's report, without which a container that crash-loops after its compose applied is seen
only through what it breaks"*.
None of the self-check's probes (D1–D13, the data and delivery probes, H-controller, H-engine, H-tools,
H-bus, DG) reads a module's container or unit state either.
## 6. The incidents, and what a declaration would have done
Every incident of the last ten days where a module applied and did not work, or looked as if it did not:
| Incident | What the gate and the self-check saw | Caught by, after | Would a declaration have caught it? |
|---|---|---|---|
| The agent server crash-looped: its database lacked the vector extension the provider never created (catalogue fix 2026-10-05) | applied; tools served (its bundle runs in the node tools, not the container); no condition | a person reading its log for another reason ([issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md)), **about a hundred restarts**; the change readying it for the home server merged 2026-09-30, the provider's fix 2026-10-05 | **Yes, by liveness alone** — *running and not restarting* — inside the gate's ten minutes; an HTTP check on its declared `web` endpoint the same |
| A web app accepted TCP and answered no HTTP request, its database unreachable after a firewall change ([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)) | *"all doing what they were told"* | a person, **after eleven hours** | **Yes, by an HTTP readiness check** within two looks (a minute at 30 s). **Not** by a TCP check — the port was open — and not by liveness |
| The identity provider's admin refused the secret the mesh minted; every consumer's client failed ([issue 179](../../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)) | the server up and serving; its image ships no check | a person, the module having been broken since it moved to the mesh; **twice** (2026-10-01, again 2026-10-05) | **Only by a module's own check** — *the admin logs in* — which the module now runs itself and announces (ADR 0224). No container check sees it |
| The studio and the flow editor read unhealthy while working (§2) | nothing: the mesh does not read health | a person reading `docker ps` | The opposite case: **a check read without proving it would have rolled back two good builds.** A declaration must be owned by the module and proved before it is trusted |
| Two media managers refused a new recycle-bin folder; their settings step exited non-zero ([issue 279](../../04-ISSUES/279-a-folder-cannot-be-owned-by-the-account-a-module-runs-as/00-report.md)) | the apply failed → the gate failed, at once | the gate | Already caught: a step is a gate. A health declaration adds nothing here |
| The media server's event handler threw on an empty answer and was offered each event five times ([issue 276](../../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md)) | — | the bus watchdog (S9), `max-deliveries` | **No, and it should not**: the server was healthy; handling an event is the event contract's, not health's |
| A provisioner runtime restarted several times at start until the overlay was up, then worked ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md)) | — | the lab | A warning: **a restart-counting check without a start period reads churn that stops as a crash loop** |
Of the seven: **two** a declaration catches that nothing catches today (the crash loop, by liveness; the
silent web app, by readiness), **one** only a module's own tool can catch (the admin), **one** shows what
an unproved check does wrong, and **three** are not health at all or are already caught.
## 7. Cost and noise
- **Reading every container's state on a machine**, health and restart count included: one `ps` of
38–49 containers took **17–21 ms**; one inspect of all of them **30–37 ms** (control node and home
server, five and three runs).
- **What the images' own checks cost today**, from the runtime's record of each check's start and end
(95 checks over 19 containers): **median 33 ms, slowest 138 ms**. On the home server they run
**77 checks a minute** — 8 containers, three of them every 2–5 s — for about **5 s of exec time a
minute**. On the control node, 11 containers at 30 s: 22 checks a minute, under 1 s.
- **The smallest machines in this mesh** have 12 cores (the control node) and 31 GB (the laptop); the busiest runs 45 catalogue containers.
One check per long-running resource at 30 s is **90 checks a minute** on the busiest machine — about
3 s of exec time a minute for in-container checks, a few hundred milliseconds for HTTP or TCP checks
the engine makes itself. The cost is not CPU; it is that every in-container check is an `exec` in the
runtime's event history, which on the home server is already one minute long.
- **Single-sample noise.** [Issue 277](../../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md):
one DNS question unanswered on a machine starting twenty containers was raised urgent; thirty asked by
hand a moment later were all answered. Its rule — *a finding one look can be wrong about is raised on
the second look in a row* — is the runtime's own `retries` under another name; the images use 3 to 20.
The gate's own three passes 40 s apart are a third form of the same idea.
## 8. What this says
1. **Most of the catalogue has no health anywhere.** 38 of 45 running container modules, all 19
service-only modules and every bundle module have nothing that says *working* beyond *applied*.
2. **What exists is not read and not owned.** 19 image checks run, 2 were wrong in the mesh's hands,
and neither the report nor the gate nor the self-check reads any of them.
3. **The runtime's memory is too short to rely on.** Restart counts vanish with each recreate; the
event history is a minute long. Liveness must be observed and kept by the node-engine.
4. **Liveness alone would have caught the worst incident; readiness the longest; only a module's own
tool the identity provider's.** All three kinds are needed, and none is enough alone.
5. **The cost is small and the noise is known.** Tens of milliseconds a look; the two-look rule exists.
@@ -0,0 +1,141 @@
# 02 — Options
Each question of the [overview](00-overview.md), the options for it, and what the
[evidence](01-evidence.md) says about each. The recommendation is [03](03-recommendation.md).
## A. Who runs the checks
### A1 — The container runtime's own HEALTHCHECK, read by the node-engine
The image's check, or one the module sets, run by the runtime inside the container; the node-engine reads
the `health` state it keeps.
- **For:** 19 images already ship one; it runs inside the container's namespace, so a check of
`localhost` sees what the program sees; the runtime handles interval, timeout, retries and start period.
- **Against:** containers only — nothing for the 19 service-only modules, the one long-running process,
or anything a module's own tool knows. Every check is an `exec`, and on the home server the 77 a minute
already push every lifecycle event out of the runtime's history (01 §2). Two of 19 shipped checks were
wrong in the mesh's configuration; an image check the module never stated is a check nobody owns. The
runtime does nothing when a container turns unhealthy (it restarts only on exit), so the state is only
useful to whoever reads it.
### A2 — The node-engine runs every check itself
HTTP and TCP from the machine, `exec` through the runtime, a unit's state through the service manager, a
tool through the node tools.
- **For:** one runner and one reader per machine, for every hosting form; it keeps what the runtime
forgets (restarts across recreates, 01 §2); a check from outside the container tests the path a caller
takes, which issue 145 says matters; HTTP and TCP cost no `exec`.
- **Against:** the engine grows a scheduler; a check that only makes sense inside the container (a CLI,
a pid file) still needs an `exec`, which A1 does better.
### A3 — The module's own tool answers "healthy"
The bundle exposes a health tool; something calls it.
- **For:** the only kind that catches a failure of function — the identity provider's refused admin
(issue 179) — and the module knows what "working" means for it.
- **Against:** a bundle is hosted by the node tools, not by the service; a tool that answers proves the
bundle is up, not the server. One tool in the catalogue is a health tool today. And a module that judges
itself is the thing ADR 0227 rule 8 refuses for the core: *the component being replaced is never the
judge*. For readiness of function it is the only option; for liveness it must not be the only one.
### A4 — A mix, by kind (the shape the evidence points at)
The engine owns every check and every verdict. It runs HTTP, TCP and unit checks itself; it delegates an
in-container command to the runtime by setting the container's HEALTHCHECK from the declaration (so the
runtime's retries and start period do the timing) and reads the state; it asks a module's health tool
through the node tools. Liveness — running and not restarting — it observes for every long-running
resource with no declaration at all.
## B. Where the result goes
| Option | For | Against |
|---|---|---|
| **B1 — A field of the report**: per module, per long-running resource, a *state* (healthy, unhealthy, starting, unknown), since when, the failing streak and the restarts the engine counted | the report is how a machine states facts about itself; the gate already reads the last report; one place to read | a report is sent after an apply, not on a change — a container that goes bad at 03:00 waits for the next report |
| **B2 — An event on each transition** (`module.<m>.<machine>.health` changed) | immediate; the controller can keep the state | events are lost or replayed; an event alone is a sample, not a state |
| **B3 — A condition, raised by the controller** | the gate already fails a module on a new condition naming it on that machine (01 §5, point 3); the operator's conversation (ADR 0234) and the healers (ADR 0231) already act on conditions; the two-look rule and the content rule already apply | a condition is the controller's word, raised from something — it needs B1 or B2 under it |
| **B4 — A seat verb the controller calls** (`health <module>`) | always current | a pull per module per judging; a machine that is slow to answer reads as unhealthy |
B1 + B2 + B3 together is how the core already works: the engine states its state in the report and
emits the change; the controller keeps the last state per machine, raises the condition on the second
look, and clears it on the look that no longer sees it. That makes **the gate need no new rule**: an
unhealthy module raises a condition naming it, and point 3 already holds the gate on it.
## C. Liveness and readiness
- **Liveness** — *it runs*: a container running and not restarted within a window; a unit `active` and
not `failed`, `NRestarts` not climbing; a process up. Observable for every long-running resource with
no declaration. It alone catches the worst incident of the window (the crash loop). It needs a **start
period** and a **settle rule**, or it reads issue 058's churn-that-stops as a crash loop.
- **Readiness** — *it serves*: the declared check passes. Only the module can say what serving means.
- **Function** — *it does its job*: the identity provider's admin logs in. A module's own tool, and only
for what no endpoint shows.
Options: judge liveness only (cheap, no declarations, misses issue 145 and 179); judge readiness only
(misses nothing that readiness sees, but every module must declare before anything is judged); **judge
liveness everywhere at once and readiness where declared**, making the declaration required over a
migration. The last keeps the gate meaningful from the first day.
What the mesh **does** with each is a separate choice. Restarting a container that is unhealthy is what
an orchestrator's liveness probe does; the runtime here does not, and a restart hides the failure the
gate is meant to see. Options: the healers restart what stays unhealthy (ADR 0231's shape: act on what
observation raised, say whether it worked), or nothing restarts on health and the condition reaches a
person. The evidence has no case where a restart would have fixed anything — the crash loop was
restarting already.
## D. Dependency-aware health
A module requiring a provision fails when its provider fails. With 12 consumers of the database
provision and 36 of a route, a provider down would be a dozen conditions said once each, and a dozen
gates failed for something none of them did.
| Option | What it costs |
|---|---|
| D1 — Ignore it | twelve conditions for one fault; the operator learns which one matters by reading all of them; the gate blames the consumers |
| D2 — A consumer's check names the provision it exercises; when that provider is unhealthy on the record, the consumer's finding is **held under the provider's** — said as *waiting on* the provider, not raised on its own — and its gate waits rather than fails | one condition, at the provider; needs the controller to know which provider answers which consumer — it does: it composes the grants ([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is about recording it) |
| D3 — Order the judging: providers first, consumers only when their providers are healthy | simple, but a consumer broken on its own is not said while a provider is down, which is the moment it is most needed |
D2 is how issue 281 already treats a machine-level condition: what is about the machine is the
machine's, and is never pinned on the module the gate is kept on.
## E. The manifest field
Research may sketch a shape; a design doc may only name the field. Where the declaration lives:
- **E1 — On the resource** (`health` on a container, a process, a service): the check belongs to the
thing it checks; a module with three containers states three. Matches how `restart-on` and `witness`
already sit on the resource.
- **E2 — One per module** (a top-level `health`): one verdict, but a module of eleven containers (the
mail module) cannot say which one is wrong.
The sketch for E1 — a field named `health`, holding:
| Part | Meaning | Default |
|---|---|---|
| kind | `runtime` (the image's own check, adopted as is), `http` (an endpoint from `listens`, a path, the status expected), `tcp` (an endpoint from `listens`), `exec` (a command in the container), `unit` (the unit's own readiness), `tool` (one of the module's tools answering) | — |
| every | interval | 30 s; not under 10 s |
| timeout | how long one look may take | 5 s; under `every` |
| after | looks failing in a row before it is unhealthy | 3; **not under 2** (issue 277) |
| grace | after a start, how long failure does not count | 60 s |
| needs | the provision whose provider it exercises, for D2 | none |
An HTTP or TCP check names an endpoint by its `listens` name, never a port or an address, so a check
follows the machine's ports the way the endpoint does. A `runtime` kind is an explicit adoption of the
image's check, so a module that ships one *says* it does, and `module check` can prove it on a lab.
## F. Migration
- **F1 — Required at once** — every long-running resource declares, or `module check` refuses. 68
modules to touch before the next merge; nothing is judged until all are done.
- **F2 — Liveness at once, declaration required by a date** — every long-running resource is judged by
liveness from the first build; `module check` warns, then refuses after a stated date; a catalogue-wide
test counts the modules still without one, and the count only goes down.
- **F3 — Optional for ever** — modules without one are judged by liveness and tools served. The 38 of 45
container modules with nothing today would stay at liveness.
What a module without a declaration is judged by in F2 and F3: liveness of every long-running resource
plus today's five points (01 §5). A bundle-only module (50) and a files-only module (7) need no
declaration: they run nothing long-lived of their own, and the node tools serving their tools is their
liveness.
@@ -0,0 +1,120 @@
# 03 — Recommendation
From the [evidence](01-evidence.md) and the [options](02-options.md): the node-engine judges every
long-running thing a module runs — **liveness for all of them, at once, with no declaration**, and
**readiness where the module declares how** — states it in its report, and the controller raises it as a
condition on the second look, so the gate, the self-check and the operator's conversation act on it with
no new rule of their own. Every catalogue module that runs something long-lived declares its check
within a stated migration, and `module check` then requires it.
## Proposed decision text
Ready for graduation through playbook 02. The number is assigned then.
> **A module says how it is healthy, and the node-engine judges it**
>
> **Context.** The release gate (ADR 0236) judges a catalogue module on its first machine by what the
> mesh sees from outside: the declaration applied, no new condition about it or its machine, its tools
> served. None of that looks at what the module runs. Of 125 catalogue modules, 49 run a container that
> stays up and 19 a service unit; no manifest can declare a health check, the node-engine reports no
> container or unit state, and no probe reads one. 19 of 73 long-running catalogue containers have an
> image check the mesh never reads, two of which were wrong in the mesh's configuration. In ten days a
> container crash-looped about a hundred times and a web app answered nothing for eleven hours, both
> while every check passed.
>
> **Decision.**
>
> 1. **Liveness is judged for every long-running resource, with no declaration.** A container that stays
> up, a process that stays up and a service stated `running` are *alive* when running and not
> restarted more than once within the settle window after their grace period. The node-engine
> observes this itself on each tick and keeps the restarts it counted across recreates; it never
> relies on the runtime's restart count or event history. A resource held still by an open window (ADR 0189) is
> neither alive nor dead: it is said as held, and judged again when the window closes.
> 2. **A module declares how each long-running resource is ready**, in a field named `health` on that
> resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint
> and the status expected, a TCP connect to a declared endpoint, a command in the container, the
> unit's own readiness, or one of the module's tools — with an interval (default 30 s, not under
> 10 s), a timeout under the interval, a number of failing looks in a row (default 3, **not under 2**),
> and a grace period after a start. An endpoint is named by its `listens` name, never by a port or an
> address. A check of function that no endpoint shows is a module's own tool, and only in addition to
> a check the module does not run itself.
> 3. **The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes
> itself; a command in the container it hands to the runtime as that container's healthcheck and reads
> the state; a tool it asks through the node tools. Nothing else on the machine judges a module.
> 4. **The state goes in the report, its change on the bus, and a condition is the controller's.** Each
> report carries, per module and long-running resource, a state — healthy, unhealthy, starting,
> unknown — since when, the failing streak and the counted restarts; each transition is emitted as an
> event. The controller keeps the last state per machine and raises `module.<module>.<machine>`
> unhealthy as a condition when two consecutive states say so, and clears it on the first that does
> not. The release gate, unchanged, holds a module on a condition naming it; at its bound the build is
> put back.
> 5. **A provider down is said once, at the provider.** A check names the provision it exercises. While
> that provision's provider is unhealthy on the record, the consumer's finding is held under the
> provider's condition — said as waiting on it — and the consumer's gate waits rather than fails.
> What a consumer finds while its provider is healthy is its own.
> 6. **Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition
> reaches a person or a healer (ADR 0231); a healer that restarts on health is its own decision.
> 7. **A declaration is proved before it is trusted.** `module check` refuses a `health` field that names
> an endpoint the module does not declare, an interval or count below the floor, or a tool the module
> does not serve. A lab bed applies every changed declaration and requires it healthy within its grace;
> an image check adopted by name is proved the same way, because two of nineteen were wrong.
> 8. **Every catalogue module that runs something long-lived declares one.** `module check` warns from
> the decision and refuses a long-running resource without `health` after the migration's date. A
> module running nothing long-lived — its bundle only, or files and packages — declares none: the node
> tools serving its tools is its liveness, as the gate judges today.
>
> **Consequences.** Liveness alone, from the first build, would have caught the crash loop inside the
> gate's ten minutes. Readiness would have caught the silent web app in a minute instead of eleven hours.
> The identity provider's refused admin needs the module's own tool, which it now has. The node-engine
> grows a small scheduler and a report field; the controller a condition kind; the gate nothing. A
> machine of 45 containers spends tens of milliseconds a look reading state, and about 3 s a minute on
> in-container commands at the default interval.
## How each rule is checked
| Rule | Checked by |
|---|---|
| 1 Liveness without declaration | node-engine unit tests over a fake runtime and service manager: a container recreated keeps its counted restarts; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy. A lab bed that replays the crash loop (a container whose program exits at start) fails the gate within its bound |
| 2 The field and its floors | `module check` refuses each out-of-range part, with a test per refusal; a catalogue-wide test parses every `health` field |
| 3 The engine owns the verdict | a test that a declared command becomes the container's healthcheck and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change |
| 4 Report, event, condition | controller tests: one unhealthy state raises nothing and is listed unconfirmed; two raise; a healthy state clears; the gate holds a module on that condition (an existing gate test, extended with this condition's kind). A lab replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks |
| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; the consumers' gates wait, not fail |
| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped |
| 7 Proved before trusted | the lab bed's run of every changed declaration on every catalogue merge; the replay of the studio's false *unhealthy* (a check that asks `localhost` while the program binds elsewhere) fails the bed, not a machine |
| 8 Every long-running module declares | a catalogue-wide test that counts the modules with a long-running resource and no `health`: it may only go down, and is zero by the date; after it, `module check` refuses |
## The migration
**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only);
57 run nothing long-lived and need no declaration.
1. **Build first, without declarations.** The node-engine's liveness, the report field, the events, the
controller's condition and its two-look rule, and the dependency hold. From this step every
long-running resource is judged by liveness. The catalogue-wide counter starts at 68.
2. **The seven modules whose images ship checks** adopt them by name — and, for the two that were
wrong, the fixed configuration is what the lab proves. 19 of their containers are covered at once.
3. **The other 42 container modules, and the containers without an image check in three of the seven,** declare an HTTP check on their declared endpoint where they serve
HTTP, TCP where they serve something else, a command where neither shows readiness. 48 of 49
already declare the endpoint the check needs.
4. **The 19 service-only modules** declare the unit's own readiness, or a TCP check where the unit
listens; most are machine software (a resolver, a time daemon, a session manager), where `active` and
not `failed` is already the honest answer.
5. **Modules whose function no endpoint shows** add a tool check: the identity provider (its admin logs
in), the database provider (it can create in a consumer's database), the broker. Identified as they
are met, not in advance.
6. **The date:** when the counter reaches zero, or six weeks after step 1, whichever is first. From then
`module check` refuses a long-running resource without `health`.
**What a module without a declaration is judged by, until then and for ever if it runs nothing
long-lived:** liveness of everything it runs that stays up, and the gate's five points as they stand —
applied, no witness put it back, no new condition naming it or its machine, the core's own definitions,
its tools served.
## What this does not decide
- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231).
- Health for scheduled work — whether the last scheduled run succeeded is data the windows of ADR 0189
already carry, and belongs to that record.
- Health of what a module's events do (issue 276) — the event contract's, and the bus watchdog's.
- The core's own definitions (to-be 45 §8), which stand; a core component may later declare its own
through the same field.