49 lines
2.5 KiB
Markdown
49 lines
2.5 KiB
Markdown
---
|
|
status: located
|
|
opened: 2026-10-03
|
|
located-in:
|
|
- mesh-host
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 210 — The host re-creates the node's runtime on every reconcile, restarting it every ten minutes
|
|
|
|
## What was observed
|
|
|
|
2026-10-03, on the laptop, while proving [design 38](../../03-DESIGN/01-to-be/38-building-the-operators-machine.md)
|
|
WP4. The host's log says the same thing at every reconcile, ten minutes apart, since the runtime
|
|
module first arrived at 00:04 — 82 times in the day's first thirteen hours:
|
|
|
|
```
|
|
created node-tools.runtime (node-tools): 239 file(s), running as node-tools.service
|
|
```
|
|
|
|
and systemd confirms it: `node-tools.service` is stopped and started at 12:39, 12:49, 12:59, 13:08.
|
|
Nothing else in those reconciles changed; every other resource is `kept`. The declaration is the
|
|
same one each time — no push happened between the cycles.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
The runtime is every module's tools on the node ([ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md)).
|
|
A restart every ten minutes drops every call in flight at that moment, re-subscribes every
|
|
membership, and re-imports every bundle; a bundle that is slow to load leaves the node's tools
|
|
silent for that long, ten minutes out of every ten. And the log reports it as success, so nothing
|
|
in `status` shows a node whose tools blink. The host's rule is that a resource reports `unchanged`
|
|
when the machine is as declared; a `process` here never does.
|
|
|
|
## Diagnosis
|
|
|
|
Owner mesh-host, `internal/apply/process.go`. The decision is: the digest of the fetched bundle
|
|
plus the unit text is `want`; when the host's record of what it last wrote equals `want` and no
|
|
`restart-on` resource changed, the process is left alone. The unit text is rendered with sorted
|
|
keys, so `want` is stable. The log says `created` — the word the code uses only when the previous
|
|
record is empty (with a record it would say `updated`) — so what fails is the record: for this
|
|
resource the host does not find what it wrote last cycle. Two candidates, to be settled in the
|
|
host's tests: the applied-store does not persist or does not key a `process` outcome the way it is
|
|
read back, or the outcome's `wrote` is dropped before it is saved. **Fix direction:** a test that
|
|
applies the same `process` declaration twice against a recorded store and asserts the second
|
|
outcome is `unchanged` with no restart. Observed only on the laptop's journal; the other three
|
|
machines' host logs are not readable by the operator account over SSH, and the behaviour is the
|
|
host's, not the machine's.
|