Research 032: measure module health on the live mesh and propose a decision
Liveness is judged nowhere and readiness cannot be declared; the evidence, options and a graduation-ready recommendation say how a module states it.
This commit is contained in:
@@ -57,3 +57,21 @@ The module manifest and `module check`; the node-engine, which would run or read
|
||||
release gate and the doctor probes of to-be 45; the healers, which may restart what stays unhealthy;
|
||||
the operator's conversation (ADR 0234), which carries what stays unhealthy; and every catalogue module,
|
||||
each of which would gain a declaration.
|
||||
|
||||
## Where it stands
|
||||
|
||||
Evidence gathered on the live mesh, options weighed and a decision drafted, 2026-10-07:
|
||||
|
||||
- [01 — The evidence](01-evidence.md): 125 catalogue modules, 68 of them running something long-lived
|
||||
(49 a container, 19 a service); no manifest can declare a check and nothing reads one; 19 of 73
|
||||
long-running catalogue containers carry an image check, two of which were wrong in the mesh's hands;
|
||||
every restart count is 0 because a recreate loses it, and the runtime's event history is about a minute
|
||||
long. Of seven recent incidents, liveness alone would have caught the crash loop, readiness the
|
||||
eleven-hour silent web app, and only a module's own tool the refused identity-provider admin.
|
||||
- [02 — Options](02-options.md): who runs the checks, where the result goes, liveness and readiness,
|
||||
dependency-aware health, the field's shape and the migration.
|
||||
- [03 — Recommendation](03-recommendation.md): liveness judged for every long-running resource at once;
|
||||
readiness declared per resource in a field named `health` and run by the node-engine; the state in the
|
||||
report, the condition raised by the controller on the second look, so the gate needs no new rule; a
|
||||
provider down said once at the provider; a proposed decision text with how each rule is checked, and
|
||||
the migration of the 68 modules.
|
||||
|
||||
@@ -0,0 +1,196 @@
|
||||
# 01 — The evidence
|
||||
|
||||
Measured on 2026-10-07 on the live mesh of four machines (a home server, a control node, a workstation
|
||||
and a laptop), read only: the three catalogue repositories at their trunk, the controller's and the
|
||||
node-engine's source at their trunk, each machine's container runtime and service manager, and the
|
||||
controller's own answers. Counts, not anecdotes. Machines are named by role.
|
||||
|
||||
## 1. What the catalogue runs
|
||||
|
||||
The three catalogue repositories hold **125 module manifests**: 113 in the main catalogue, 12 in the
|
||||
media catalogue. The photos repository holds application code and no manifest; its two instances are
|
||||
modules in the main catalogue.
|
||||
|
||||
Each manifest sorted by the **longest-lived thing it runs** (a container that stays up first, then a
|
||||
service the manifest says is running, then the mesh's own process that stays up, then a bundle only):
|
||||
|
||||
| What a module runs | Modules |
|
||||
|---|---|
|
||||
| at least one container that stays up | **49** |
|
||||
| no such container, but a service unit stated `running` | **19** |
|
||||
| only its bundle — tools and handlers hosted by the node tools | **50** |
|
||||
| only files, directories and packages | **7** |
|
||||
|
||||
Underneath:
|
||||
|
||||
| Resource | Count | Of which |
|
||||
|---|---|---|
|
||||
| container | **85** in 50 modules | 77 stay up, 6 run once (a step), 2 on a schedule; **71 distinct images** |
|
||||
| service (an existing unit put in a state) | **31** in 21 modules | 26 stated `running`, 5 stateless (the machine's lifecycle) |
|
||||
| process (the mesh's own code in a unit it writes) | **20** in 18 modules | 17 run once, 2 on a schedule, **1** stays up |
|
||||
| bundle artifact | **101** in 99 modules | |
|
||||
|
||||
And what a check could be built from:
|
||||
|
||||
- **50** modules declare `listens` (an endpoint, a port, a protocol) — **48 of the 49** container
|
||||
modules. A TCP or HTTP check has its target named already.
|
||||
- **53** modules declare `tools`. **One** tool in the whole catalogue is a health tool by name
|
||||
(`dbus_health`); **18** modules have a tool named `…_status`.
|
||||
- **67** modules `require` something; the most-required provisions are `route` (36 consumers),
|
||||
`x11-display` (15) and `postgres-database` (12). A database down is, today, potentially twelve
|
||||
consumers failing at once.
|
||||
- **No manifest declares a health check.** The container resource has no field for one — its fields are
|
||||
image, environment, files, ports, volumes, arguments, names, networks, capabilities, logging, the
|
||||
restart triggers and the run-once and schedule modes — and the node-engine passes the runtime no
|
||||
health option. Whatever check runs is the image's own, run by the runtime by default.
|
||||
|
||||
## 2. What the images already ship, and what the runtime says today
|
||||
|
||||
Every container on every machine inspected — its state, its health, its restart count, and its
|
||||
**image's** healthcheck from the image configuration:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| containers on the four machines | **115** (94 running) |
|
||||
| of those, declared by the catalogue and running | **79 instances** of **73** of the 77 long-running containers (4 are assigned nowhere) |
|
||||
| long-running catalogue containers whose **image ships a HEALTHCHECK** | **19 of 73 (26 %)** |
|
||||
| …in how many modules | **7 of 45** modules with a running container (mail, the hosted database suite, a spreadsheet app, the chat client, a flow editor, the media server, the certificate authority) |
|
||||
| …modules whose every container has one | **4** |
|
||||
| running container modules with **no health state at all** | **38 of 45** |
|
||||
| catalogue containers reporting `healthy` / `unhealthy` / `starting` | **19 / 0 / 0** |
|
||||
| containers reporting `unhealthy` | 5, all exited, all on the workstation, all predating the mesh (none declared) |
|
||||
|
||||
**Two of the nineteen image checks were wrong under the mesh's own configuration** until a catalogue fix:
|
||||
|
||||
- the hosted database suite's studio — the framework binds the address the runtime puts in `HOSTNAME`, so
|
||||
it answered only on its network address while its image's check asked `localhost`: *unhealthy for ever
|
||||
while working* (fixed 2026-10-05);
|
||||
- the flow editor — the image's check reads its settings from a path the module mounted elsewhere: *ran
|
||||
fine but reported unhealthy forever* (fixed 2026-09-30).
|
||||
|
||||
So 2 of 19 shipped checks (≈ 10 %) gave a false *unhealthy* in the mesh's hands. Read blindly, they would
|
||||
have failed two good builds at the gate.
|
||||
|
||||
The intervals the images chose range from **2 s to 60 s**: a database every 2 s, three web apps every
|
||||
5 s, an API gateway every 10 s, the rest 30 s or 60 s. Start periods range from none to 360 s (the
|
||||
antivirus). Retries from 3 to 20.
|
||||
|
||||
### The restart count the runtime keeps is lost
|
||||
|
||||
**All 94 running containers report a restart count of 0** — including the agent server that, per
|
||||
[issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md), had restarted about
|
||||
a hundred times in a crash loop days before. The counter belongs to a container, and the node-engine
|
||||
recreates a container whenever its declaration or a file it reads changes; the fix recreated it and the
|
||||
history went with it. A restart count is only evidence if something outside the container keeps it.
|
||||
|
||||
### The runtime's event history is about a minute long
|
||||
|
||||
The runtime keeps its most recent ~250 events in memory. On the home server, a query for the last 30
|
||||
minutes and for the last minute both answered ~250 events, **every one an `exec_*` event** — the image
|
||||
health checks running (84 creates, 85 dies, 84 starts). A 24-hour query for container lifecycle events,
|
||||
through the mesh's own `docker_events` tool, answered **zero**. The images' checks crowd every lifecycle
|
||||
event out of the history within a minute. A container that died an hour ago leaves no trace there.
|
||||
|
||||
## 3. What the service manager says today
|
||||
|
||||
- **Failed system units:** 0 on three machines; 2 on the workstation, both mounts that the mesh does not
|
||||
declare. One failed user unit on each of two machines, neither the mesh's.
|
||||
- The mesh's own units (`mesh-ca-trust`, `mesh-filter`, the power units, the controller, the node tools,
|
||||
the node-engine): all `active`, **`NRestarts` 0**.
|
||||
- systemd already gives, per unit and for free: `ActiveState`/`SubState`, `is-failed`, `NRestarts`,
|
||||
the time it entered its state, and — through a unit's own `ExecStartPost`/`WatchdogSec` where software
|
||||
supports it — readiness. None of it is read by the mesh for a module's unit.
|
||||
|
||||
## 4. What the node-engine reports today
|
||||
|
||||
The node-engine's report (its `Report` message) carries: what applied and what failed per resource,
|
||||
refusals, the declaration's digest and order, held and stray containers, reachable sockets on an adopted
|
||||
machine, packet filters and the firewall found, windows (a scheduled step holding containers still), the
|
||||
machine's profile, the engine's own build, outward links, rollbacks its witnesses decided, and whether it
|
||||
reads `witness`. Separately, an `Alive` heartbeat with its interval.
|
||||
|
||||
**It carries no container state, no unit state and no health.** The controller's `node` answer for the
|
||||
home server shows the consequence: its assigned modules, its strays, its filters and its capabilities —
|
||||
and nothing about whether any of its 45 catalogue containers is running.
|
||||
|
||||
The node-engine has one judging mechanism already: the **witness** (to-be 45 §8) — a process's `witness`
|
||||
field, `lease` or `ping` or `none`, judges a new build of the controller or the node tools and restores
|
||||
the previous one when it is not healthy in bound. Only the two core processes use it.
|
||||
|
||||
A **run-once step** is already a health gate in practice: a step exiting non-zero fails the apply and
|
||||
everything placed after it. **23** steps exist (6 containers, 17 processes); one of them, in the hosted
|
||||
database suite, is a loop that waits up to five minutes for its analytics service's `/health` before the
|
||||
services that need it start.
|
||||
|
||||
## 5. What the release gate judges a module by
|
||||
|
||||
The gate on a plan's first machine ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md);
|
||||
`judgeHealth` in the controller) passes a module there when, on three judgings at least 40 s apart over at
|
||||
least two minutes, within ten minutes of the send:
|
||||
|
||||
1. no witness on the machine put a core build back since the send;
|
||||
2. the machine's last report is current and **applied** (failed or refused is broken);
|
||||
3. **no condition** raised since the send names the module on that machine, or the machine as a whole
|
||||
(issue 281 sorted the latter out);
|
||||
4. for the controller, the node-engine and the node tools: their own health definitions (lease held and
|
||||
ready; the engine's build reported; the node tools answering the bus);
|
||||
5. for any other module **that declares tools**: the machine's node tools serve them.
|
||||
|
||||
That is the whole of it for a catalogue module. For the 49 container modules nothing in 1–5 looks at a
|
||||
container: point 5 asks the node tools, which host the module's bundle, not its container. **A container
|
||||
that applied and then crash-loops passes all five.** To-be 45's Phase 4 note says so: *"container state in
|
||||
the node-engine's report, without which a container that crash-loops after its compose applied is seen
|
||||
only through what it breaks"*.
|
||||
|
||||
None of the self-check's probes (D1–D13, the data and delivery probes, H-controller, H-engine, H-tools,
|
||||
H-bus, DG) reads a module's container or unit state either.
|
||||
|
||||
## 6. The incidents, and what a declaration would have done
|
||||
|
||||
Every incident of the last ten days where a module applied and did not work, or looked as if it did not:
|
||||
|
||||
| Incident | What the gate and the self-check saw | Caught by, after | Would a declaration have caught it? |
|
||||
|---|---|---|---|
|
||||
| The agent server crash-looped: its database lacked the vector extension the provider never created (catalogue fix 2026-10-05) | applied; tools served (its bundle runs in the node tools, not the container); no condition | a person reading its log for another reason ([issue 268](../../04-ISSUES/268-letta-printed-its-passwords-into-its-log/00-report.md)), **about a hundred restarts**; the change readying it for the home server merged 2026-09-30, the provider's fix 2026-10-05 | **Yes, by liveness alone** — *running and not restarting* — inside the gate's ten minutes; an HTTP check on its declared `web` endpoint the same |
|
||||
| A web app accepted TCP and answered no HTTP request, its database unreachable after a firewall change ([issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)) | *"all doing what they were told"* | a person, **after eleven hours** | **Yes, by an HTTP readiness check** within two looks (a minute at 30 s). **Not** by a TCP check — the port was open — and not by liveness |
|
||||
| The identity provider's admin refused the secret the mesh minted; every consumer's client failed ([issue 179](../../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)) | the server up and serving; its image ships no check | a person, the module having been broken since it moved to the mesh; **twice** (2026-10-01, again 2026-10-05) | **Only by a module's own check** — *the admin logs in* — which the module now runs itself and announces (ADR 0224). No container check sees it |
|
||||
| The studio and the flow editor read unhealthy while working (§2) | nothing: the mesh does not read health | a person reading `docker ps` | The opposite case: **a check read without proving it would have rolled back two good builds.** A declaration must be owned by the module and proved before it is trusted |
|
||||
| Two media managers refused a new recycle-bin folder; their settings step exited non-zero ([issue 279](../../04-ISSUES/279-a-folder-cannot-be-owned-by-the-account-a-module-runs-as/00-report.md)) | the apply failed → the gate failed, at once | the gate | Already caught: a step is a gate. A health declaration adds nothing here |
|
||||
| The media server's event handler threw on an empty answer and was offered each event five times ([issue 276](../../04-ISSUES/276-a-handler-that-did-its-work-was-offered-it-five-times/00-report.md)) | — | the bus watchdog (S9), `max-deliveries` | **No, and it should not**: the server was healthy; handling an event is the event contract's, not health's |
|
||||
| A provisioner runtime restarted several times at start until the overlay was up, then worked ([issue 058](../../04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md)) | — | the lab | A warning: **a restart-counting check without a start period reads churn that stops as a crash loop** |
|
||||
|
||||
Of the seven: **two** a declaration catches that nothing catches today (the crash loop, by liveness; the
|
||||
silent web app, by readiness), **one** only a module's own tool can catch (the admin), **one** shows what
|
||||
an unproved check does wrong, and **three** are not health at all or are already caught.
|
||||
|
||||
## 7. Cost and noise
|
||||
|
||||
- **Reading every container's state on a machine**, health and restart count included: one `ps` of
|
||||
38–49 containers took **17–21 ms**; one inspect of all of them **30–37 ms** (control node and home
|
||||
server, five and three runs).
|
||||
- **What the images' own checks cost today**, from the runtime's record of each check's start and end
|
||||
(95 checks over 19 containers): **median 33 ms, slowest 138 ms**. On the home server they run
|
||||
**77 checks a minute** — 8 containers, three of them every 2–5 s — for about **5 s of exec time a
|
||||
minute**. On the control node, 11 containers at 30 s: 22 checks a minute, under 1 s.
|
||||
- **The smallest machines in this mesh** have 12 cores (the control node) and 31 GB (the laptop); the busiest runs 45 catalogue containers.
|
||||
One check per long-running resource at 30 s is **90 checks a minute** on the busiest machine — about
|
||||
3 s of exec time a minute for in-container checks, a few hundred milliseconds for HTTP or TCP checks
|
||||
the engine makes itself. The cost is not CPU; it is that every in-container check is an `exec` in the
|
||||
runtime's event history, which on the home server is already one minute long.
|
||||
- **Single-sample noise.** [Issue 277](../../04-ISSUES/277-one-unanswered-question-was-an-urgent-alert-nobody-could-read/00-report.md):
|
||||
one DNS question unanswered on a machine starting twenty containers was raised urgent; thirty asked by
|
||||
hand a moment later were all answered. Its rule — *a finding one look can be wrong about is raised on
|
||||
the second look in a row* — is the runtime's own `retries` under another name; the images use 3 to 20.
|
||||
The gate's own three passes 40 s apart are a third form of the same idea.
|
||||
|
||||
## 8. What this says
|
||||
|
||||
1. **Most of the catalogue has no health anywhere.** 38 of 45 running container modules, all 19
|
||||
service-only modules and every bundle module have nothing that says *working* beyond *applied*.
|
||||
2. **What exists is not read and not owned.** 19 image checks run, 2 were wrong in the mesh's hands,
|
||||
and neither the report nor the gate nor the self-check reads any of them.
|
||||
3. **The runtime's memory is too short to rely on.** Restart counts vanish with each recreate; the
|
||||
event history is a minute long. Liveness must be observed and kept by the node-engine.
|
||||
4. **Liveness alone would have caught the worst incident; readiness the longest; only a module's own
|
||||
tool the identity provider's.** All three kinds are needed, and none is enough alone.
|
||||
5. **The cost is small and the noise is known.** Tens of milliseconds a look; the two-look rule exists.
|
||||
@@ -0,0 +1,141 @@
|
||||
# 02 — Options
|
||||
|
||||
Each question of the [overview](00-overview.md), the options for it, and what the
|
||||
[evidence](01-evidence.md) says about each. The recommendation is [03](03-recommendation.md).
|
||||
|
||||
## A. Who runs the checks
|
||||
|
||||
### A1 — The container runtime's own HEALTHCHECK, read by the node-engine
|
||||
|
||||
The image's check, or one the module sets, run by the runtime inside the container; the node-engine reads
|
||||
the `health` state it keeps.
|
||||
|
||||
- **For:** 19 images already ship one; it runs inside the container's namespace, so a check of
|
||||
`localhost` sees what the program sees; the runtime handles interval, timeout, retries and start period.
|
||||
- **Against:** containers only — nothing for the 19 service-only modules, the one long-running process,
|
||||
or anything a module's own tool knows. Every check is an `exec`, and on the home server the 77 a minute
|
||||
already push every lifecycle event out of the runtime's history (01 §2). Two of 19 shipped checks were
|
||||
wrong in the mesh's configuration; an image check the module never stated is a check nobody owns. The
|
||||
runtime does nothing when a container turns unhealthy (it restarts only on exit), so the state is only
|
||||
useful to whoever reads it.
|
||||
|
||||
### A2 — The node-engine runs every check itself
|
||||
|
||||
HTTP and TCP from the machine, `exec` through the runtime, a unit's state through the service manager, a
|
||||
tool through the node tools.
|
||||
|
||||
- **For:** one runner and one reader per machine, for every hosting form; it keeps what the runtime
|
||||
forgets (restarts across recreates, 01 §2); a check from outside the container tests the path a caller
|
||||
takes, which issue 145 says matters; HTTP and TCP cost no `exec`.
|
||||
- **Against:** the engine grows a scheduler; a check that only makes sense inside the container (a CLI,
|
||||
a pid file) still needs an `exec`, which A1 does better.
|
||||
|
||||
### A3 — The module's own tool answers "healthy"
|
||||
|
||||
The bundle exposes a health tool; something calls it.
|
||||
|
||||
- **For:** the only kind that catches a failure of function — the identity provider's refused admin
|
||||
(issue 179) — and the module knows what "working" means for it.
|
||||
- **Against:** a bundle is hosted by the node tools, not by the service; a tool that answers proves the
|
||||
bundle is up, not the server. One tool in the catalogue is a health tool today. And a module that judges
|
||||
itself is the thing ADR 0227 rule 8 refuses for the core: *the component being replaced is never the
|
||||
judge*. For readiness of function it is the only option; for liveness it must not be the only one.
|
||||
|
||||
### A4 — A mix, by kind (the shape the evidence points at)
|
||||
|
||||
The engine owns every check and every verdict. It runs HTTP, TCP and unit checks itself; it delegates an
|
||||
in-container command to the runtime by setting the container's HEALTHCHECK from the declaration (so the
|
||||
runtime's retries and start period do the timing) and reads the state; it asks a module's health tool
|
||||
through the node tools. Liveness — running and not restarting — it observes for every long-running
|
||||
resource with no declaration at all.
|
||||
|
||||
## B. Where the result goes
|
||||
|
||||
| Option | For | Against |
|
||||
|---|---|---|
|
||||
| **B1 — A field of the report**: per module, per long-running resource, a *state* (healthy, unhealthy, starting, unknown), since when, the failing streak and the restarts the engine counted | the report is how a machine states facts about itself; the gate already reads the last report; one place to read | a report is sent after an apply, not on a change — a container that goes bad at 03:00 waits for the next report |
|
||||
| **B2 — An event on each transition** (`module.<m>.<machine>.health` changed) | immediate; the controller can keep the state | events are lost or replayed; an event alone is a sample, not a state |
|
||||
| **B3 — A condition, raised by the controller** | the gate already fails a module on a new condition naming it on that machine (01 §5, point 3); the operator's conversation (ADR 0234) and the healers (ADR 0231) already act on conditions; the two-look rule and the content rule already apply | a condition is the controller's word, raised from something — it needs B1 or B2 under it |
|
||||
| **B4 — A seat verb the controller calls** (`health <module>`) | always current | a pull per module per judging; a machine that is slow to answer reads as unhealthy |
|
||||
|
||||
B1 + B2 + B3 together is how the core already works: the engine states its state in the report and
|
||||
emits the change; the controller keeps the last state per machine, raises the condition on the second
|
||||
look, and clears it on the look that no longer sees it. That makes **the gate need no new rule**: an
|
||||
unhealthy module raises a condition naming it, and point 3 already holds the gate on it.
|
||||
|
||||
## C. Liveness and readiness
|
||||
|
||||
- **Liveness** — *it runs*: a container running and not restarted within a window; a unit `active` and
|
||||
not `failed`, `NRestarts` not climbing; a process up. Observable for every long-running resource with
|
||||
no declaration. It alone catches the worst incident of the window (the crash loop). It needs a **start
|
||||
period** and a **settle rule**, or it reads issue 058's churn-that-stops as a crash loop.
|
||||
- **Readiness** — *it serves*: the declared check passes. Only the module can say what serving means.
|
||||
- **Function** — *it does its job*: the identity provider's admin logs in. A module's own tool, and only
|
||||
for what no endpoint shows.
|
||||
|
||||
Options: judge liveness only (cheap, no declarations, misses issue 145 and 179); judge readiness only
|
||||
(misses nothing that readiness sees, but every module must declare before anything is judged); **judge
|
||||
liveness everywhere at once and readiness where declared**, making the declaration required over a
|
||||
migration. The last keeps the gate meaningful from the first day.
|
||||
|
||||
What the mesh **does** with each is a separate choice. Restarting a container that is unhealthy is what
|
||||
an orchestrator's liveness probe does; the runtime here does not, and a restart hides the failure the
|
||||
gate is meant to see. Options: the healers restart what stays unhealthy (ADR 0231's shape: act on what
|
||||
observation raised, say whether it worked), or nothing restarts on health and the condition reaches a
|
||||
person. The evidence has no case where a restart would have fixed anything — the crash loop was
|
||||
restarting already.
|
||||
|
||||
## D. Dependency-aware health
|
||||
|
||||
A module requiring a provision fails when its provider fails. With 12 consumers of the database
|
||||
provision and 36 of a route, a provider down would be a dozen conditions said once each, and a dozen
|
||||
gates failed for something none of them did.
|
||||
|
||||
| Option | What it costs |
|
||||
|---|---|
|
||||
| D1 — Ignore it | twelve conditions for one fault; the operator learns which one matters by reading all of them; the gate blames the consumers |
|
||||
| D2 — A consumer's check names the provision it exercises; when that provider is unhealthy on the record, the consumer's finding is **held under the provider's** — said as *waiting on* the provider, not raised on its own — and its gate waits rather than fails | one condition, at the provider; needs the controller to know which provider answers which consumer — it does: it composes the grants ([issue 181](../../04-ISSUES/181-an-assignment-does-not-record-which-provider-answers-it/00-report.md) is about recording it) |
|
||||
| D3 — Order the judging: providers first, consumers only when their providers are healthy | simple, but a consumer broken on its own is not said while a provider is down, which is the moment it is most needed |
|
||||
|
||||
D2 is how issue 281 already treats a machine-level condition: what is about the machine is the
|
||||
machine's, and is never pinned on the module the gate is kept on.
|
||||
|
||||
## E. The manifest field
|
||||
|
||||
Research may sketch a shape; a design doc may only name the field. Where the declaration lives:
|
||||
|
||||
- **E1 — On the resource** (`health` on a container, a process, a service): the check belongs to the
|
||||
thing it checks; a module with three containers states three. Matches how `restart-on` and `witness`
|
||||
already sit on the resource.
|
||||
- **E2 — One per module** (a top-level `health`): one verdict, but a module of eleven containers (the
|
||||
mail module) cannot say which one is wrong.
|
||||
|
||||
The sketch for E1 — a field named `health`, holding:
|
||||
|
||||
| Part | Meaning | Default |
|
||||
|---|---|---|
|
||||
| kind | `runtime` (the image's own check, adopted as is), `http` (an endpoint from `listens`, a path, the status expected), `tcp` (an endpoint from `listens`), `exec` (a command in the container), `unit` (the unit's own readiness), `tool` (one of the module's tools answering) | — |
|
||||
| every | interval | 30 s; not under 10 s |
|
||||
| timeout | how long one look may take | 5 s; under `every` |
|
||||
| after | looks failing in a row before it is unhealthy | 3; **not under 2** (issue 277) |
|
||||
| grace | after a start, how long failure does not count | 60 s |
|
||||
| needs | the provision whose provider it exercises, for D2 | none |
|
||||
|
||||
An HTTP or TCP check names an endpoint by its `listens` name, never a port or an address, so a check
|
||||
follows the machine's ports the way the endpoint does. A `runtime` kind is an explicit adoption of the
|
||||
image's check, so a module that ships one *says* it does, and `module check` can prove it on a lab.
|
||||
|
||||
## F. Migration
|
||||
|
||||
- **F1 — Required at once** — every long-running resource declares, or `module check` refuses. 68
|
||||
modules to touch before the next merge; nothing is judged until all are done.
|
||||
- **F2 — Liveness at once, declaration required by a date** — every long-running resource is judged by
|
||||
liveness from the first build; `module check` warns, then refuses after a stated date; a catalogue-wide
|
||||
test counts the modules still without one, and the count only goes down.
|
||||
- **F3 — Optional for ever** — modules without one are judged by liveness and tools served. The 38 of 45
|
||||
container modules with nothing today would stay at liveness.
|
||||
|
||||
What a module without a declaration is judged by in F2 and F3: liveness of every long-running resource
|
||||
plus today's five points (01 §5). A bundle-only module (50) and a files-only module (7) need no
|
||||
declaration: they run nothing long-lived of their own, and the node tools serving their tools is their
|
||||
liveness.
|
||||
@@ -0,0 +1,120 @@
|
||||
# 03 — Recommendation
|
||||
|
||||
From the [evidence](01-evidence.md) and the [options](02-options.md): the node-engine judges every
|
||||
long-running thing a module runs — **liveness for all of them, at once, with no declaration**, and
|
||||
**readiness where the module declares how** — states it in its report, and the controller raises it as a
|
||||
condition on the second look, so the gate, the self-check and the operator's conversation act on it with
|
||||
no new rule of their own. Every catalogue module that runs something long-lived declares its check
|
||||
within a stated migration, and `module check` then requires it.
|
||||
|
||||
## Proposed decision text
|
||||
|
||||
Ready for graduation through playbook 02. The number is assigned then.
|
||||
|
||||
> **A module says how it is healthy, and the node-engine judges it**
|
||||
>
|
||||
> **Context.** The release gate (ADR 0236) judges a catalogue module on its first machine by what the
|
||||
> mesh sees from outside: the declaration applied, no new condition about it or its machine, its tools
|
||||
> served. None of that looks at what the module runs. Of 125 catalogue modules, 49 run a container that
|
||||
> stays up and 19 a service unit; no manifest can declare a health check, the node-engine reports no
|
||||
> container or unit state, and no probe reads one. 19 of 73 long-running catalogue containers have an
|
||||
> image check the mesh never reads, two of which were wrong in the mesh's configuration. In ten days a
|
||||
> container crash-looped about a hundred times and a web app answered nothing for eleven hours, both
|
||||
> while every check passed.
|
||||
>
|
||||
> **Decision.**
|
||||
>
|
||||
> 1. **Liveness is judged for every long-running resource, with no declaration.** A container that stays
|
||||
> up, a process that stays up and a service stated `running` are *alive* when running and not
|
||||
> restarted more than once within the settle window after their grace period. The node-engine
|
||||
> observes this itself on each tick and keeps the restarts it counted across recreates; it never
|
||||
> relies on the runtime's restart count or event history. A resource held still by an open window (ADR 0189) is
|
||||
> neither alive nor dead: it is said as held, and judged again when the window closes.
|
||||
> 2. **A module declares how each long-running resource is ready**, in a field named `health` on that
|
||||
> resource: one kind — the image's own check adopted by name, an HTTP request to a declared endpoint
|
||||
> and the status expected, a TCP connect to a declared endpoint, a command in the container, the
|
||||
> unit's own readiness, or one of the module's tools — with an interval (default 30 s, not under
|
||||
> 10 s), a timeout under the interval, a number of failing looks in a row (default 3, **not under 2**),
|
||||
> and a grace period after a start. An endpoint is named by its `listens` name, never by a port or an
|
||||
> address. A check of function that no endpoint shows is a module's own tool, and only in addition to
|
||||
> a check the module does not run itself.
|
||||
> 3. **The node-engine runs every check and owns every verdict.** HTTP, TCP and unit checks it makes
|
||||
> itself; a command in the container it hands to the runtime as that container's healthcheck and reads
|
||||
> the state; a tool it asks through the node tools. Nothing else on the machine judges a module.
|
||||
> 4. **The state goes in the report, its change on the bus, and a condition is the controller's.** Each
|
||||
> report carries, per module and long-running resource, a state — healthy, unhealthy, starting,
|
||||
> unknown — since when, the failing streak and the counted restarts; each transition is emitted as an
|
||||
> event. The controller keeps the last state per machine and raises `module.<module>.<machine>`
|
||||
> unhealthy as a condition when two consecutive states say so, and clears it on the first that does
|
||||
> not. The release gate, unchanged, holds a module on a condition naming it; at its bound the build is
|
||||
> put back.
|
||||
> 5. **A provider down is said once, at the provider.** A check names the provision it exercises. While
|
||||
> that provision's provider is unhealthy on the record, the consumer's finding is held under the
|
||||
> provider's condition — said as waiting on it — and the consumer's gate waits rather than fails.
|
||||
> What a consumer finds while its provider is healthy is its own.
|
||||
> 6. **Nothing is restarted for being unhealthy.** The runtime restarts what exits, as now. A condition
|
||||
> reaches a person or a healer (ADR 0231); a healer that restarts on health is its own decision.
|
||||
> 7. **A declaration is proved before it is trusted.** `module check` refuses a `health` field that names
|
||||
> an endpoint the module does not declare, an interval or count below the floor, or a tool the module
|
||||
> does not serve. A lab bed applies every changed declaration and requires it healthy within its grace;
|
||||
> an image check adopted by name is proved the same way, because two of nineteen were wrong.
|
||||
> 8. **Every catalogue module that runs something long-lived declares one.** `module check` warns from
|
||||
> the decision and refuses a long-running resource without `health` after the migration's date. A
|
||||
> module running nothing long-lived — its bundle only, or files and packages — declares none: the node
|
||||
> tools serving its tools is its liveness, as the gate judges today.
|
||||
>
|
||||
> **Consequences.** Liveness alone, from the first build, would have caught the crash loop inside the
|
||||
> gate's ten minutes. Readiness would have caught the silent web app in a minute instead of eleven hours.
|
||||
> The identity provider's refused admin needs the module's own tool, which it now has. The node-engine
|
||||
> grows a small scheduler and a report field; the controller a condition kind; the gate nothing. A
|
||||
> machine of 45 containers spends tens of milliseconds a look reading state, and about 3 s a minute on
|
||||
> in-container commands at the default interval.
|
||||
|
||||
## How each rule is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| 1 Liveness without declaration | node-engine unit tests over a fake runtime and service manager: a container recreated keeps its counted restarts; a restart inside grace is not counted; two restarts within the settle window after grace make it unhealthy. A lab bed that replays the crash loop (a container whose program exits at start) fails the gate within its bound |
|
||||
| 2 The field and its floors | `module check` refuses each out-of-range part, with a test per refusal; a catalogue-wide test parses every `health` field |
|
||||
| 3 The engine owns the verdict | a test that a declared command becomes the container's healthcheck and nothing else sets one; a test that an HTTP check dials the endpoint's current port after a port change |
|
||||
| 4 Report, event, condition | controller tests: one unhealthy state raises nothing and is listed unconfirmed; two raise; a healthy state clears; the gate holds a module on that condition (an existing gate test, extended with this condition's kind). A lab replay of issue 145 — a database made unreachable — raises the consumer's condition within two looks |
|
||||
| 5 Said once at the provider | a controller test with one unhealthy database provider and three consumers failing: one condition, at the provider, the consumers listed as waiting; the consumers' gates wait, not fail |
|
||||
| 6 No restart on health | a node-engine test that an unhealthy container is not restarted, recreated or stopped |
|
||||
| 7 Proved before trusted | the lab bed's run of every changed declaration on every catalogue merge; the replay of the studio's false *unhealthy* (a check that asks `localhost` while the program binds elsewhere) fails the bed, not a machine |
|
||||
| 8 Every long-running module declares | a catalogue-wide test that counts the modules with a long-running resource and no `health`: it may only go down, and is zero by the date; after it, `module check` refuses |
|
||||
|
||||
## The migration
|
||||
|
||||
**Today:** 68 modules run something long-lived (49 with a container, 19 with a running service only);
|
||||
57 run nothing long-lived and need no declaration.
|
||||
|
||||
1. **Build first, without declarations.** The node-engine's liveness, the report field, the events, the
|
||||
controller's condition and its two-look rule, and the dependency hold. From this step every
|
||||
long-running resource is judged by liveness. The catalogue-wide counter starts at 68.
|
||||
2. **The seven modules whose images ship checks** adopt them by name — and, for the two that were
|
||||
wrong, the fixed configuration is what the lab proves. 19 of their containers are covered at once.
|
||||
3. **The other 42 container modules, and the containers without an image check in three of the seven,** declare an HTTP check on their declared endpoint where they serve
|
||||
HTTP, TCP where they serve something else, a command where neither shows readiness. 48 of 49
|
||||
already declare the endpoint the check needs.
|
||||
4. **The 19 service-only modules** declare the unit's own readiness, or a TCP check where the unit
|
||||
listens; most are machine software (a resolver, a time daemon, a session manager), where `active` and
|
||||
not `failed` is already the honest answer.
|
||||
5. **Modules whose function no endpoint shows** add a tool check: the identity provider (its admin logs
|
||||
in), the database provider (it can create in a consumer's database), the broker. Identified as they
|
||||
are met, not in advance.
|
||||
6. **The date:** when the counter reaches zero, or six weeks after step 1, whichever is first. From then
|
||||
`module check` refuses a long-running resource without `health`.
|
||||
|
||||
**What a module without a declaration is judged by, until then and for ever if it runs nothing
|
||||
long-lived:** liveness of everything it runs that stays up, and the gate's five points as they stand —
|
||||
applied, no witness put it back, no new condition naming it or its machine, the core's own definitions,
|
||||
its tools served.
|
||||
|
||||
## What this does not decide
|
||||
|
||||
- Whether a healer restarts what stays unhealthy (rule 6 leaves it to its own record, under ADR 0231).
|
||||
- Health for scheduled work — whether the last scheduled run succeeded is data the windows of ADR 0189
|
||||
already carry, and belongs to that record.
|
||||
- Health of what a module's events do (issue 276) — the event contract's, and the bus watchdog's.
|
||||
- The core's own definitions (to-be 45 §8), which stand; a core component may later declare its own
|
||||
through the same field.
|
||||
Reference in New Issue
Block a user