From ef6ab814b0539a09fea3ad14f7f2c26f5a82a910 Mon Sep 17 00:00:00 2001 From: jochen Date: Wed, 7 Oct 2026 01:18:00 +0200 Subject: [PATCH] Research 032: measure module health on the live mesh and propose a decision Liveness is judged nowhere and readiness cannot be declared; the evidence, options and a graduation-ready recommendation say how a module states it. --- .../00-overview.md | 18 ++ .../01-evidence.md | 196 ++++++++++++++++++ .../02-options.md | 141 +++++++++++++ .../03-recommendation.md | 120 +++++++++++ 4 files changed, 475 insertions(+) create mode 100644 01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md create mode 100644 01-RESEARCH/032-a-module-says-how-it-is-healthy/02-options.md create mode 100644 01-RESEARCH/032-a-module-says-how-it-is-healthy/03-recommendation.md diff --git a/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md b/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md index de7ffe88..a3fd357a 100644 --- a/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md +++ b/01-RESEARCH/032-a-module-says-how-it-is-healthy/00-overview.md @@ -57,3 +57,21 @@ The module manifest and `module check`; the node-engine, which would run or read release gate and the doctor probes of to-be 45; the healers, which may restart what stays unhealthy; the operator's conversation (ADR 0234), which carries what stays unhealthy; and every catalogue module, each of which would gain a declaration. + +## Where it stands + +Evidence gathered on the live mesh, options weighed and a decision drafted, 2026-10-07: + +- [01 — The evidence](01-evidence.md): 125 catalogue modules, 68 of them running something long-lived + (49 a container, 19 a service); no manifest can declare a check and nothing reads one; 19 of 73 + long-running catalogue containers carry an image check, two of which were wrong in the mesh's hands; + every restart count is 0 because a recreate loses it, and the runtime's event history is about a minute + long. Of seven recent incidents, liveness alone would have caught the crash loop, readiness the + eleven-hour silent web app, and only a module's own tool the refused identity-provider admin. +- [02 — Options](02-options.md): who runs the checks, where the result goes, liveness and readiness, + dependency-aware health, the field's shape and the migration. +- [03 — Recommendation](03-recommendation.md): liveness judged for every long-running resource at once; + readiness declared per resource in a field named `health` and run by the node-engine; the state in the + report, the condition raised by the controller on the second look, so the gate needs no new rule; a + provider down said once at the provider; a proposed decision text with how each rule is checked, and + the migration of the 68 modules. diff --git a/01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md b/01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md new file mode 100644 index 00000000..6460f806 --- /dev/null +++ b/01-RESEARCH/032-a-module-says-how-it-is-healthy/01-evidence.md @@ -0,0 +1,196 @@ +# 01 — The evidence + +Measured on 2026-10-07 on the live mesh of four machines (a home server, a control node, a workstation +and a laptop), read only: the three catalogue repositories at their trunk, the controller's and the +node-engine's source at their trunk, each machine's container runtime and service manager, and the +controller's own answers. Counts, not anecdotes. Machines are named by role. + +## 1. What the catalogue runs + +The three catalogue repositories hold **125 module manifests**: 113 in the main catalogue, 12 in the +media catalogue. The photos repository holds application code and no manifest; its two instances are +modules in the main catalogue. + +Each manifest sorted by the **longest-lived thing it runs** (a container that stays up first, then a +service the manifest says is running, then the mesh's own process that stays up, then a bundle only): + +| What a module runs | Modules | +|---|---| +| at least one container that stays up | **49** | +| no such container, but a service unit stated `running` | **19** | +| only its bundle — tools and handlers hosted by the node tools | **50** | +| only files, directories and packages | **7** | + +Underneath: + +| Resource | Count | Of which | +|---|---|---| +| container | **85** in 50 modules | 77 stay up, 6 run once (a step), 2 on a schedule; **71 distinct images** | +| service (an existing unit put in a state) | **31** in 21 modules | 26 stated `running`, 5 stateless (the machine's lifecycle) | +| process (the mesh's own code in a unit it writes) | **20** in 18 modules | 17 run once, 2 on a schedule, **1** stays up | +| bundle artifact | **101** in 99 modules | | + +And what a check could be built from: + +- **50** modules declare `listens` (an endpoint, a port, a protocol) — **48 of the 49** container + modules. A TCP or HTTP check has its target named already. +- **53** modules declare `tools`. **One** tool in the whole catalogue is a health tool by name + (`dbus_health`); **18** modules have a tool named `…_status`. +- **67** modules `require` something; the most-required provisions are `route` (36 consumers), + `x11-display` (15) and `postgres-database` (12). A database down is, today, potentially twelve + consumers failing at once. +- **No manifest declares a health check.** The container resource has no field for one — its fields are + image, environment, files, ports, volumes, arguments, names, networks, capabilities, logging, the + restart triggers and the run-once and schedule modes — and the node-engine passes the runtime no + health option. Whatever check runs is the image's own, run by the runtime by default. + +## 2. What the images already ship, and what the runtime says today + +Every container on every machine inspected — its state, its health, its restart count, and its +**image's** healthcheck from the image configuration: + +| | | +|---|---| +| containers on the four machines | **115** (94 running) | +| of those, declared by the catalogue and running | **79 instances** of **73** of the 77 long-running containers (4 are assigned nowhere) | +| long-running catalogue containers whose **image ships a HEALTHCHECK** | **19 of 73 (26 %)** | +| …in how many modules | **7 of 45** modules with a running container (mail, the hosted database suite, a spreadsheet app, the chat client, a flow editor, the media server, the certificate authority) | +| …modules whose every container has one | **4** | +| running container modules with **no health state at all** | **38 of 45** | +| catalogue containers reporting `healthy` / `unhealthy` / `starting` | **19 / 0 / 0** | +| containers reporting `unhealthy` | 5, all exited, all on the workstation, all predating the mesh (none declared) | + +**Two of the nineteen image checks were wrong under the mesh's own configuration** until a catalogue fix: + +- the hosted database suite's studio — the framework binds the address the runtime puts in `HOSTNAME`, so + it answered only on its network address while its image's check asked `localhost`: *unhealthy for ever + while working* (fixed 2026-10-05); +- the flow editor — the image's check reads its settings from a path the module mounted elsewhere: *ran + fine but reported unhealthy forever* (fixed 2026-09-30). + +So 2 of 19 shipped checks (≈ 10 %) gave a false *unhealthy* in the mesh's hands. Read blindly, they would +have failed two good builds at the gate. + +The intervals the images chose range from **2 s to 60 s**: a database every 2 s, three web apps every +5 s, an API gateway every 10 s, the rest 30 s or 60 s. Start periods range from none to 360 s (the +antivirus). Retries from 3 to 20. + +### The restart count the runtime keeps is lost + +**All 94 running containers report a restart count of 0** — including the agent server that, per +[issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md), had restarted about +a hundred times in a crash loop days before. The counter belongs to a container, and the node-engine +recreates a container whenever its declaration or a file it reads changes; the fix recreated it and the +history went with it. A restart count is only evidence if something outside the container keeps it. + +### The runtime's event history is about a minute long + +The runtime keeps its most recent ~250 events in memory. On the home server, a query for the last 30 +minutes and for the last minute both answered ~250 events, **every one an `exec_*` event** — the image +health checks running (84 creates, 85 dies, 84 starts). A 24-hour query for container lifecycle events, +through the mesh's own `docker_events` tool, answered **zero**. The images' checks crowd every lifecycle +event out of the history within a minute. A container that died an hour ago leaves no trace there. + +## 3. What the service manager says today + +- **Failed system units:** 0 on three machines; 2 on the workstation, both mounts that the mesh does not + declare. One failed user unit on each of two machines, neither the mesh's. +- The mesh's own units (`mesh-ca-trust`, `mesh-filter`, the power units, the controller, the node tools, + the node-engine): all `active`, **`NRestarts` 0**. +- systemd already gives, per unit and for free: `ActiveState`/`SubState`, `is-failed`, `NRestarts`, + the time it entered its state, and — through a unit's own `ExecStartPost`/`WatchdogSec` where software + supports it — readiness. None of it is read by the mesh for a module's unit. + +## 4. What the node-engine reports today + +The node-engine's report (its `Report` message) carries: what applied and what failed per resource, +refusals, the declaration's digest and order, held and stray containers, reachable sockets on an adopted +machine, packet filters and the firewall found, windows (a scheduled step holding containers still), the +machine's profile, the engine's own build, outward links, rollbacks its witnesses decided, and whether it +reads `witness`. Separately, an `Alive` heartbeat with its interval. + +**It carries no container state, no unit state and no health.** The controller's `node` answer for the +home server shows the consequence: its assigned modules, its strays, its filters and its capabilities — +and nothing about whether any of its 45 catalogue containers is running. + +The node-engine has one judging mechanism already: the **witness** (to-be 45 §8) — a process's `witness` +field, `lease` or `ping` or `none`, judges a new build of the controller or the node tools and restores +the previous one when it is not healthy in bound. Only the two core processes use it. + +A **run-once step** is already a health gate in practice: a step exiting non-zero fails the apply and +everything placed after it. **23** steps exist (6 containers, 17 processes); one of them, in the hosted +database suite, is a loop that waits up to five minutes for its analytics service's `/health` before the +services that need it start. + +## 5. What the release gate judges a module by + +The gate on a plan's first machine ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md); +`judgeHealth` in the controller) passes a module there when, on three judgings at least 40 s apart over at +least two minutes, within ten minutes of the send: + +1. no witness on the machine put a core build back since the send; +2. the machine's last report is current and **applied** (failed or refused is broken); +3. **no condition** raised since the send names the module on that machine, or the machine as a whole + (issue 281 sorted the latter out); +4. for the controller, the node-engine and the node tools: their own health definitions (lease held and + ready; the engine's build reported; the node tools answering the bus); +5. for any other module **that declares tools**: the machine's node tools serve them. + +That is the whole of it for a catalogue module. For the 49 container modules nothing in 1–5 looks at a +container: point 5 asks the node tools, which host the module's bundle, not its container. **A container +that applied and then crash-loops passes all five.** To-be 45's Phase 4 note says so: *"container state in +the node-engine's report, without which a container that crash-loops after its compose applied is seen +only through what it breaks"*. + +None of the self-check's probes (D1–D13, the data and delivery probes, H-controller, H-engine, H-tools, +H-bus, DG) reads a module's container or unit state either. + +## 6. The incidents, and what a declaration would have done + +Every incident of the last ten days where a module applied and did not work, or looked as if it did not: + +| Incident | What the gate and the self-check saw | Caught by, after | Would a declaration have caught it? | +|---|---|---|---| +| The agent server crash-looped: its database lacked the vector extension the provider never created (catalogue fix 2026-10-05) | applied; tools served (its bundle runs in the node tools, not the container); no condition | a person reading its log for another reason ([issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md)), **about a hundred restarts**; the change readying it for the home server merged 2026-09-30, the provider's fix 2026-10-05 | **Yes, by liveness alone** — *running and not restarting* — inside the gate's ten minutes; an HTTP check on its declared `web` endpoint the same | +| A web app accepted TCP and answered no HTTP request, its database unreachable after a firewall change ([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)) | *"all doing what they were told"* | a person, **after eleven hours** | **Yes, by an HTTP readiness check** within two looks (a minute at 30 s). **Not** by a TCP check — the port was open — and not by liveness | +| The identity provider's admin refused the secret the mesh minted; every consumer's client failed ([issue 179](../../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)) | the server up and serving; its image ships no check | a person, the module having been broken since it moved to the mesh; **twice** (2026-10-01, again 2026-10-05) | **Only by a module's own check** — *the admin logs in* — which the module now runs itself and announces (ADR 0224). No container check sees it | +| The studio and the flow editor read unhealthy while working (§2) | nothing: the mesh does not read health | a person reading `docker ps` | The opposite case: **a check read without proving it would have rolled back two good builds.** A declaration must be owned by the module and proved before it is trusted | +| Two media managers refused a new recycle-bin folder; their settings step exited non-zero ([issue 279](../../04-ISSUES/279-a-folder-cannot-be-owned-by-the-account-a-module-runs-as/00-report.md)) | the apply failed → the gate failed, at once | the gate | Already caught: a step is a gate. A health declaration adds nothing here | +| The media server's event handler threw on an empty answer and was offered each event five times ([issue 276](../../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md)) | — | the bus watchdog (S9), `max-deliveries` | **No, and it should not**: the server was healthy; handling an event is the event contract's, not health's | +| A provisioner runtime restarted several times at start until the overlay was up, then worked ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md)) | — | the lab | A warning: **a restart-counting check without a start period reads churn that stops as a crash loop** | + +Of the seven: **two** a declaration catches that nothing catches today (the crash loop, by liveness; the +silent web app, by readiness), **one** only a module's own tool can catch (the admin), **one** shows what +an unproved check does wrong, and **three** are not health at all or are already caught. + +## 7. Cost and noise + +- **Reading every container's state on a machine**, health and restart count included: one `ps` of + 38–49 containers took **17–21 ms**; one inspect of all of them **30–37 ms** (control node and home + server, five and three runs). +- **What the images' own checks cost today**, from the runtime's record of each check's start and end + (95 checks over 19 containers): **median 33 ms, slowest 138 ms**. On the home server they run + **77 checks a minute** — 8 containers, three of them every 2–5 s — for about **5 s of exec time a + minute**. On the control node, 11 containers at 30 s: 22 checks a minute, under 1 s. +- **The smallest machines in this mesh** have 12 cores (the control node) and 31 GB (the laptop); the busiest runs 45 catalogue containers. + One check per long-running resource at 30 s is **90 checks a minute** on the busiest machine — about + 3 s of exec time a minute for in-container checks, a few hundred milliseconds for HTTP or TCP checks + the engine makes itself. The cost is not CPU; it is that every in-container check is an `exec` in the + runtime's event history, which on the home server is already one minute long. +- **Single-sample noise.** [Issue 277](../../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md): + one DNS question unanswered on a machine starting twenty containers was raised urgent; thirty asked by + hand a moment later were all answered. Its rule — *a finding one look can be wrong about is raised on + the second look in a row* — is the runtime's own `retries` under another name; the images use 3 to 20. + The gate's own three passes 40 s apart are a third form of the same idea. + +## 8. What this says + +1. **Most of the catalogue has no health anywhere.** 38 of 45 running container modules, all 19 + service-only modules and every bundle module have nothing that says *working* beyond *applied*. +2. **What exists is not read and not owned.** 19 image checks run, 2 were wrong in the mesh's hands, + and neither the report nor the gate nor the self-check reads any of them. +3. **The runtime's memory is too short to rely on.** Restart counts vanish with each recreate; the + event history is a minute long. Liveness must be observed and kept by the node-engine. +4. **Liveness alone would have caught the worst incident; readiness the longest; only a module's own + tool the identity provider's.** All three kinds are needed, and none is enough alone. +5. **The cost is small and the noise is known.** Tens of milliseconds a look; the two-look rule exists. diff --git a/01-RESEARCH/032-a-module-says-how-it-is-healthy/02-options.md b/01-RESEARCH/032-a-module-says-how-it-is-healthy/02-options.md new file mode 100644 index 00000000..3d0e6aa4 --- /dev/null +++ b/01-RESEARCH/032-a-module-says-how-it-is-healthy/02-options.md @@ -0,0 +1,141 @@ +# 02 — Options + +Each question of the [overview](00-overview.md), the options for it, and what the +[evidence](01-evidence.md) says about each. The recommendation is [03](03-recommendation.md). + +## A. Who runs the checks + +### A1 — The container runtime's own HEALTHCHECK, read by the node-engine + +The image's check, or one the module sets, run by the runtime inside the container; the node-engine reads +the `health` state it keeps. + +- **For:** 19 images already ship one; it runs inside the container's namespace, so a check of + `localhost` sees what the program sees; the runtime handles interval, timeout, retries and start period. +- **Against:** containers only — nothing for the 19 service-only modules, the one long-running process, + or anything a module's own tool knows. Every check is an `exec`, and on the home server the 77 a minute + already push every lifecycle event out of the runtime's history (01 §2). Two of 19 shipped checks were + wrong in the mesh's configuration; an image check the module never stated is a check nobody owns. The + runtime does nothing when a container turns unhealthy (it restarts only on exit), so the state is only + useful to whoever reads it. + +### A2 — The node-engine runs every check itself + +HTTP and TCP from the machine, `exec` through the runtime, a unit's state through the service manager, a +tool through the node tools. + +- **For:** one runner and one reader per machine, for every hosting form; it keeps what the runtime + forgets (restarts across recreates, 01 §2); a check from outside the container tests the path a caller + takes, which issue 145 says matters; HTTP and TCP cost no `exec`. +- **Against:** the engine grows a scheduler; a check that only makes sense inside the container (a CLI, + a pid file) still needs an `exec`, which A1 does better. + +### A3 — The module's own tool answers "healthy" + +The bundle exposes a health tool; something calls it. + +- **For:** the only kind that catches a failure of function — the identity provider's refused admin + (issue 179) — and the module knows what "working" means for it. +- **Against:** a bundle is hosted by the node tools, not by the service; a tool that answers proves the + bundle is up, not the server. One tool in the catalogue is a health tool today. And a module that judges + itself is the thing ADR 0227 rule 8 refuses for the core: *the component being replaced is never the + judge*. For readiness of function it is the only option; for liveness it must not be the only one. + +### A4 — A mix, by kind (the shape the evidence points at) + +The engine owns every check and every verdict. It runs HTTP, TCP and unit checks itself; it delegates an +in-container command to the runtime by setting the container's HEALTHCHECK from the declaration (so the +runtime's retries and start period do the timing) and reads the state; it asks a module's health tool +through the node tools. Liveness — running and not restarting — it observes for every long-running +resource with no declaration at all. + +## B. Where the result goes + +| Option | For | Against | +|---|---|---| +| **B1 — A field of the report**: per module, per long-running resource, a *state* (healthy, unhealthy, starting, unknown), since when, the failing streak and the restarts the engine counted | the report is how a machine states facts about itself; the gate already reads the last report; one place to read | a report is sent after an apply, not on a change — a container that goes bad at 03:00 waits for the next report | +| **B2 — An event on each transition** (`module...health` changed) | immediate; the controller can keep the state | events are lost or replayed; an event alone is a sample, not a state | +| **B3 — A condition, raised by the controller** | the gate already fails a module on a new condition naming it on that machine (01 §5, point 3); the operator's conversation (ADR 0234) and the healers (ADR 0231) already act on conditions; the two-look rule and the content rule already apply | a condition is the controller's word, raised from something — it needs B1 or B2 under it | +| **B4 — A seat verb the controller calls** (`health `) | always current | a pull per module per judging; a machine that is slow to answer reads as unhealthy | + +B1 + B2 + B3 together is how the core already works: the engine states its state in the report and +emits the change; the controller keeps the last state per machine, raises the condition on the second +look, and clears it on the look that no longer sees it. That makes **the gate need no new rule**: an +unhealthy module raises a condition naming it, and point 3 already holds the gate on it. + +## C. Liveness and readiness + +- **Liveness** — *it runs*: a container running and not restarted within a window; a unit `active` and + not `failed`, `NRestarts` not climbing; a process up. Observable for every long-running resource with + no declaration. It alone catches the worst incident of the window (the crash loop). It needs a **start + period** and a **settle rule**, or it reads issue 058's churn-that-stops as a crash loop. +- **Readiness** — *it serves*: the declared check passes. Only the module can say what serving means. +- **Function** — *it does its job*: the identity provider's admin logs in. A module's own tool, and only + for what no endpoint shows. + +Options: judge liveness only (cheap, no declarations, misses issue 145 and 179); judge readiness only +(misses nothing that readiness sees, but every module must declare before anything is judged); **judge +liveness everywhere at once and readiness where declared**, making the declaration required over a +migration. The last keeps the gate meaningful from the first day. + +What the mesh **does** with each is a separate choice. Restarting a container that is unhealthy is what +an orchestrator's liveness probe does; the runtime here does not, and a restart hides the failure the +gate is meant to see. Options: the healers restart what stays unhealthy (ADR 0231's shape: act on what +observation raised, say whether it worked), or nothing restarts on health and the condition reaches a +person. The evidence has no case where a restart would have fixed anything — the crash loop was +restarting already. + +## D. Dependency-aware health + +A module requiring a provision fails when its provider fails. With 12 consumers of the database +provision and 36 of a route, a provider down would be a dozen conditions said once each, and a dozen +gates failed for something none of them did. + +| Option | What it costs | +|---|---| +| D1 — Ignore it | twelve conditions for one fault; the operator learns which one matters by reading all of them; the gate blames the consumers | +| D2 — A consumer's check names the provision it exercises; when that provider is unhealthy on the record, the consumer's finding is **held under the provider's** — said as *waiting on* the provider, not raised on its own — and its gate waits rather than fails | one condition, at the provider; needs the controller to know which provider answers which consumer — it does: it composes the grants ([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is about recording it) | +| D3 — Order the judging: providers first, consumers only when their providers are healthy | simple, but a consumer broken on its own is not said while a provider is down, which is the moment it is most needed | + +D2 is how issue 281 already treats a machine-level condition: what is about the machine is the +machine's, and is never pinned on the module the gate is kept on. + +## E. The manifest field + +Research may sketch a shape; a design doc may only name the field. Where the declaration lives: + +- **E1 — On the resource** (`health` on a container, a process, a service): the check belongs to the + thing it checks; a module with three containers states three. Matches how `restart-on` and `witness` + already sit on the resource. +- **E2 — One per module** (a top-level `health`): one verdict, but a module of eleven containers (the + mail module) cannot say which one is wrong. + +The sketch for E1 — a field named `health`, holding: + +| Part | Meaning | Default | +|---|---|---| +| kind | `runtime` (the image's own check, adopted as is), `http` (an endpoint from `listens`, a path, the status expected), `tcp` (an endpoint from `listens`), `exec` (a command in the container), `unit` (the unit's own readiness), `tool` (one of the module's tools answering) | — | +| every | interval | 30 s; not under 10 s | +| timeout | how long one look may take | 5 s; under `every` | +| after | looks failing in a row before it is unhealthy | 3; **not under 2** (issue 277) | +| grace | after a start, how long failure does not count | 60 s | +| needs | the provision whose provider it exercises, for D2 | none | + +An HTTP or TCP check names an endpoint by its `listens` name, never a port or an address, so a check +follows the machine's ports the way the endpoint does. A `runtime` kind is an explicit adoption of the +image's check, so a module that ships one *says* it does, and `module check` can prove it on a lab. + +## F. Migration + +- **F1 — Required at once** — every long-running resource declares, or `module check` refuses. 68 + modules to touch before the next merge; nothing is judged until all are done. +- **F2 — Liveness at once, declaration required by a date** — every long-running resource is judged by + liveness from the first build; `module check` warns, then refuses after a stated date; a catalogue-wide + test counts the modules still without one, and the count only goes down. +- **F3 — Optional for ever** — modules without one are judged by liveness and tools served. The 38 of 45 + container modules with nothing today would stay at liveness. + +What a module without a declaration is judged by in F2 and F3: liveness of every long-running resource +plus today's five points (01 §5). A bundle-only module (50) and a files-only module (7) need no +declaration: they run nothing long-lived of their own, and the node tools serving their tools is their +liveness. diff --git a/01-RESEARCH/032-a-module-says-how-it-is-healthy/03-recommendation.md b/01-RESEARCH/032-a-module-says-how-it-is-healthy/03-recommendation.md new file mode 100644 index 00000000..fae63724 --- /dev/null +++ b/01-RESEARCH/032-a-module-says-how-it-is-healthy/03-recommendation.md @@ -0,0 +1,120 @@ +# 03 — Recommendation + +From the [evidence](01-evidence.md) and the [options](02-options.md): the node-engine judges every +long-running thing a module runs — **liveness for all of them, at once, with no declaration**, and +**readiness where the module declares how** — states it in its report, and the controller raises it as a +condition on the second look, so the gate, the self-check and the operator's conversation act on it with +no new rule of their own. Every catalogue module that runs something long-lived declares its check +within a stated migration, and `module check` then requires it. + +## Proposed decision text + +Ready for graduation through playbook 02. The number is assigned then. + +> **A module says how it is healthy, and the node-engine judges it** +> +> **Context.** The release gate (ADR 0236) judges a catalogue module on its first machine by what the +> mesh sees from outside: the declaration applied, no new condition about it or its machine, its tools +> served. None of that looks at what the module runs. Of 125 catalogue modules, 49 run a container that +> stays up and 19 a service unit; no manifest can declare a health check, the node-engine reports no +> container or unit state, and no probe reads one. 19 of 73 long-running catalogue containers have an +> image check the mesh never reads, two of which were wrong in the mesh's configuration. In ten days a +> container crash-looped about a hundred times and a web app answered nothing for eleven hours, both +> while every check passed. +> +> **Decision.** +> +> 1. **Liveness is judged for every long-running resource, with no declaration.** A container that stays +> up, a process that stays up and a service stated `running` are *alive* when running and not +> restarted more than once within the settle window after their grace period. The node-engine +> observes this itself on each tick and keeps the restarts it counted across recreates; it never +> relies on the runtime's restart count or event history. A resource held still by an open window (ADR 0189) is +> neither alive nor dead: it is said as held, and judged again when the window closes. +> 2. **A module declares how each long-running resource is ready**, in a field named `health` on that +> resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint +> and the status expected, a TCP connect to a declared endpoint, a command in the container, the +> unit's own readiness, or one of the module's tools — with an interval (default 30 s, not under +> 10 s), a timeout under the interval, a number of failing looks in a row (default 3, **not under 2**), +> and a grace period after a start. An endpoint is named by its `listens` name, never by a port or an +> address. A check of function that no endpoint shows is a module's own tool, and only in addition to +> a check the module does not run itself. +> 3. **The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes +> itself; a command in the container it hands to the runtime as that container's healthcheck and reads +> the state; a tool it asks through the node tools. Nothing else on the machine judges a module. +> 4. **The state goes in the report, its change on the bus, and a condition is the controller's.** Each +> report carries, per module and long-running resource, a state — healthy, unhealthy, starting, +> unknown — since when, the failing streak and the counted restarts; each transition is emitted as an +> event. The controller keeps the last state per machine and raises `module..` +> unhealthy as a condition when two consecutive states say so, and clears it on the first that does +> not. The release gate, unchanged, holds a module on a condition naming it; at its bound the build is +> put back. +> 5. **A provider down is said once, at the provider.** A check names the provision it exercises. While +> that provision's provider is unhealthy on the record, the consumer's finding is held under the +> provider's condition — said as waiting on it — and the consumer's gate waits rather than fails. +> What a consumer finds while its provider is healthy is its own. +> 6. **Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition +> reaches a person or a healer (ADR 0231); a healer that restarts on health is its own decision. +> 7. **A declaration is proved before it is trusted.** `module check` refuses a `health` field that names +> an endpoint the module does not declare, an interval or count below the floor, or a tool the module +> does not serve. A lab bed applies every changed declaration and requires it healthy within its grace; +> an image check adopted by name is proved the same way, because two of nineteen were wrong. +> 8. **Every catalogue module that runs something long-lived declares one.** `module check` warns from +> the decision and refuses a long-running resource without `health` after the migration's date. A +> module running nothing long-lived — its bundle only, or files and packages — declares none: the node +> tools serving its tools is its liveness, as the gate judges today. +> +> **Consequences.** Liveness alone, from the first build, would have caught the crash loop inside the +> gate's ten minutes. Readiness would have caught the silent web app in a minute instead of eleven hours. +> The identity provider's refused admin needs the module's own tool, which it now has. The node-engine +> grows a small scheduler and a report field; the controller a condition kind; the gate nothing. A +> machine of 45 containers spends tens of milliseconds a look reading state, and about 3 s a minute on +> in-container commands at the default interval. + +## How each rule is checked + +| Rule | Checked by | +|---|---| +| 1 Liveness without declaration | node-engine unit tests over a fake runtime and service manager: a container recreated keeps its counted restarts; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy. A lab bed that replays the crash loop (a container whose program exits at start) fails the gate within its bound | +| 2 The field and its floors | `module check` refuses each out-of-range part, with a test per refusal; a catalogue-wide test parses every `health` field | +| 3 The engine owns the verdict | a test that a declared command becomes the container's healthcheck and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change | +| 4 Report, event, condition | controller tests: one unhealthy state raises nothing and is listed unconfirmed; two raise; a healthy state clears; the gate holds a module on that condition (an existing gate test, extended with this condition's kind). A lab replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks | +| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; the consumers' gates wait, not fail | +| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped | +| 7 Proved before trusted | the lab bed's run of every changed declaration on every catalogue merge; the replay of the studio's false *unhealthy* (a check that asks `localhost` while the program binds elsewhere) fails the bed, not a machine | +| 8 Every long-running module declares | a catalogue-wide test that counts the modules with a long-running resource and no `health`: it may only go down, and is zero by the date; after it, `module check` refuses | + +## The migration + +**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only); +57 run nothing long-lived and need no declaration. + +1. **Build first, without declarations.** The node-engine's liveness, the report field, the events, the + controller's condition and its two-look rule, and the dependency hold. From this step every + long-running resource is judged by liveness. The catalogue-wide counter starts at 68. +2. **The seven modules whose images ship checks** adopt them by name — and, for the two that were + wrong, the fixed configuration is what the lab proves. 19 of their containers are covered at once. +3. **The other 42 container modules, and the containers without an image check in three of the seven,** declare an HTTP check on their declared endpoint where they serve + HTTP, TCP where they serve something else, a command where neither shows readiness. 48 of 49 + already declare the endpoint the check needs. +4. **The 19 service-only modules** declare the unit's own readiness, or a TCP check where the unit + listens; most are machine software (a resolver, a time daemon, a session manager), where `active` and + not `failed` is already the honest answer. +5. **Modules whose function no endpoint shows** add a tool check: the identity provider (its admin logs + in), the database provider (it can create in a consumer's database), the broker. Identified as they + are met, not in advance. +6. **The date:** when the counter reaches zero, or six weeks after step 1, whichever is first. From then + `module check` refuses a long-running resource without `health`. + +**What a module without a declaration is judged by, until then and for ever if it runs nothing +long-lived:** liveness of everything it runs that stays up, and the gate's five points as they stand — +applied, no witness put it back, no new condition naming it or its machine, the core's own definitions, +its tools served. + +## What this does not decide + +- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231). +- Health for scheduled work — whether the last scheduled run succeeded is data the windows of ADR 0189 + already carry, and belongs to that record. +- Health of what a module's events do (issue 276) — the event contract's, and the bus watchdog's. +- The core's own definitions (to-be 45 §8), which stand; a core component may later declare its own + through the same field.