docker inspect <name> resolves across every object kind, not just
containers. A module regularly names a network the same as the
container that joins it (keycloak does this today, ordinarily) — so
when the container does not exist yet but the same-named network
already does, the bare form answers with the network's JSON instead
of reporting the container absent, and the template these callers use
(.State.Running) fails to execute against it entirely.
Live on novox tonight: minio's LB container, named the same as its
network ("minio"), could never be created — every apply crashed on
"the container runtime could not say whether minio is here", stuck
since first push, because the check itself never got a clean answer.
Fixed at every call site asking a container's state by name
(containerState, inspectFound, NamesFree, raiseGiteaServer,
containerRunning) by scoping to `docker container inspect`, matching
the type-scoped form this codebase already uses correctly for
networks, volumes and images elsewhere. Also scoped the one image
inspect that was still bare (publish.go), for the same reason.
mesh-host runs as a host-level service (nox-mesh-host.service), not a
Docker module — merging this does not redeploy it. The live novox
failure persists until the service itself is rebuilt and updated.
Two faults that both reported success while being wrong, found while proving
the firewall module actually delivers.
A unit whose job is to apply something and exit — load a rule set, set a
sysctl — is inactive the instant it succeeds. Reading that as stopped made it
permanently unsatisfiable: the host started it, it worked, the host read back
stopped and reported the machine as not doing what it was told, on every apply,
for ever, with the rules correctly in place the whole time. That is what the
firewall has been doing on every machine it was assigned to, and why the
four-machine bed was red.
And a container took its identity from its own fields, not from the files it
reads. A file written in an earlier apply — or before the container declared it
as a dependency — left a process holding a credential the mesh had already
replaced, with everything reporting success (novox/hq 04-ISSUES/045). What a
container reads is now part of what it is, so the comparison is a standing one
rather than a tripwire that fires during one apply and never again.
A schedule: container (ADR 0053) is installed as present state and never run at apply — the
Scheduler fires it later on its cadence. But a service or run-once container only gets its
image as a side effect of docker run, so a scheduled step's image was not pulled until its
first scheduled fire: absent from the node right after a successful apply, so the first run
paid the whole pull latency and tooling that expects the image present after apply found it
missing.
applyContainer now probes the runtime and ensures the pinned image present for a scheduled
step before recording it. A new ensureImage helper inspects the image and pulls it only if
absent, then reads back (ADR 0018). Ensuring an image is not running it: no docker run fires
the container, so the no-run invariant of ADR 0053 holds. The runtime probe, previously
skipped for a schedule, now runs because a pull needs it — the schedule.go comment is updated
to match.
Tests: the install-does-not-run test is extended to allow the image-ensure while asserting no
fire and no needless pull; a new test applies a scheduled container whose image is absent and
asserts it is pulled and still not started. go build, go vet, go test ./... all pass.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The recurring twin of run-once, one modifier over: a container marked
schedule: "<cron>" is run to completion on its cadence, not started as a
service and not run once as a gate.
The gating rule is deliberately reversed. Installing a schedule records it
as present state and reports the node current at once (applySchedule) --
it never runs the container and does not gate what follows. A Scheduler,
held for the life of the daemon and re-established from each applied
declaration (the declaration is the source of truth, ADR 0018), fires the
container off an injected clock. A run that exits non-zero is logged and
never fails the apply or flips the node's state, because it happens
outside the apply and the store entirely. Runs never stack: a run still
going when the next is due is skipped, not started as a second copy.
No new host shape and no new action -- schedule is a string on the
container the host already has, and the host process runs the container
itself rather than installing a system timer (the rejected option 1). A
minimal five-field cron (declaration/cron.go) validates on arrival and
computes the next due minute; time is injected so the scheduler is tested
without the wall clock.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF