Files
hq/02-DECISIONS/0184-a-service-the-mesh-asked-to-run-is-still-running-a-moment-later.md
T

4.2 KiB

topic, status, date, deciders, reconstructed, extends
topic status date deciders reconstructed extends
the mesh accepted 2026-10-02 jochen false 02-DECISIONS/0005-the-node-host.md

184. A service the mesh asked to run is still running a moment later

Context

The host already refuses to take a service manager's word for it. Three places in one function read a unit back after acting on it, each with a comment saying why: a service manager accepting a command says the transaction was accepted, not that the unit is running — one that starts and immediately dies satisfies it. The intent was right and the implementation did not reach it.

On 2026-10-02 the mesh composed a fail2ban jail whose pattern the daemon refused. The host wrote the files, restarted the service, read the unit back and reported restarted. The unit was active at that instant and failed 221 milliseconds later, which the unit's own record states. Both public machines then kept no bans at all — every jail, not the one at fault — and nothing in the mesh said so. The fault was found by calling a tool that needed the daemon, not by the mesh noticing.

The read-back races the failure. A service manager returns when it has started the process; a daemon that reads its configuration, refuses it and exits does so a fraction of a second afterwards. One look sees activating or active whatever the process is about to do, and the host reports success for a machine that is already wrong — the one shape of failure this host exists to refuse (ADR 0005).

A command the module declares — test the configuration before restarting — was considered and rejected. The link carries no actions (ADR 0005), and a verification command is a command: a declaration that carried one would be remote execution over the bus, arriving as root on every machine, which is a far larger door than the fault it closes. The host does not need one. It already knows what it asked for.

Decision

1. A unit the host has just asked to run is read twice, with a pause between the reads long enough for a daemon that refuses its configuration to have exited. Not running at the second look is a failure of that resource, named with the unit and the state it is in — the same failure the single read was always meant to catch.

2. It is never a wait for a unit to come up. A unit still starting reads as running at both looks and is accepted, exactly as before. What the second look catches is a unit that was running and is not any more. A service asked to be stopped is not waited on at all.

3. The host tests nothing and runs nothing of a module's. The second look is the host checking the state it was told to establish, which is its whole job; the declaration gains no vocabulary, and no command reaches a machine that did not already come from a built artifact.

Consequences

  • Every apply that starts, restarts or reloads a service spends a moment confirming it. The cost is bounded by the number of services that changed in that apply, which is usually none.
  • A module whose configuration the mesh composes — the packet filter, the intrusion prevention, the resolver — now fails its apply when the composition is bad, instead of reporting success onto a dead daemon. status names the machine, which is how the operator finds out.
  • It does not prevent the bad composition. ADR 0179's manifest check is what refuses the one that caused this, at merge time; this record is what makes the next one visible within a minute rather than invisible until something asks the daemon a question.

How this is checked

Rule Checked by
A unit that is running at the first look and dead at the second fails the apply, naming the unit and its state a host test over a service manager that answers as systemd does
A unit still starting is accepted at both looks a host test
A service asked to be stopped is not waited on a host test

References