Issue 210: the host re-creates the node's runtime on every reconcile, restarting it every ten minutes #317

Merged
mesh-admin merged 2 commits from issues/210-the-host-re-creates-the-runtime-every-cycle into main 2026-10-03 11:20:27 +00:00
@@ -0,0 +1,51 @@
---
status: located
opened: 2026-10-03
located-in:
- mesh-host
fixed-by:
amended-design:
---
# 210 — The host re-creates the node's runtime on every reconcile, restarting it every ten minutes
## What was observed
2026-10-03, on the laptop, while proving [design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md)
WP4. The host's log says the same thing at every reconcile, ten minutes apart, since the runtime
module first arrived at 00:04 — 82 times in the day's first thirteen hours:
```
created node-tools.runtime (node-tools): 239 file(s), running as node-tools.service
```
and systemd confirms it: `node-tools.service` is stopped and started at 12:39, 12:49, 12:59, 13:08.
Nothing else in those reconciles changed; every other resource is `kept`. The declaration is the
same one each time — no push happened between the cycles.
## Why it matters beyond this instance
The runtime is every module's tools on the node ([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)).
A restart every ten minutes drops every call in flight at that moment, re-subscribes every
membership, and re-imports every bundle; a bundle that is slow to load leaves the node's tools
silent for that long, ten minutes out of every ten. And the log reports it as success, so nothing
in `status` shows a node whose tools blink. The host's rule is that a resource reports `unchanged`
when the machine is as declared; a `process` here never does.
## Diagnosis
Owner mesh-host, `internal/apply/process.go`. The decision is: the digest of the fetched bundle
plus the unit text is `want`; when the host's record of what it last wrote equals `want` and no
`restart-on` resource changed, the process is left alone. Two things were ruled out on the laptop:
the credential the runtime restarts on, which has not been written since the evening before (the
host would also say `updated`, not `created`, for a restart it owed to another resource); and the
unit text, rendered with sorted keys and so stable. What fails is the record. The host's state file
holds an entry for `node-tools.runtime` — applied at the last cycle, with **no `wrote` digest at
all** — while the sibling entry for the runtime's credential carries its digest. The apply sets the
digest on its outcome on every path that installs the daemon, and the loop that records outcomes
copies it into the record for every kind; between the two, a `process` outcome arrives with its
digest empty. **Fix direction:** find where a `process` outcome loses its digest on the way to the
record, and a test that applies the same `process` declaration twice against a recorded store and
asserts the second outcome is `unchanged` with no restart — the test the shape never had. Observed
on the laptop's journal and state; the other three machines' host logs are not readable by the
operator account over SSH, and the behaviour is the host's, not the machine's.