Files
hq/02-DECISIONS/0060-the-host-is-built-per-operating-system.md
T
jschoubben c557f99cba Record what testing podman actually showed
0060 said the container runtime was a separate decision. It is now made, and
the reasoning is worth keeping because it is the opposite answer to the same
question one paragraph earlier.

Abstracting service managers is lossy -- systemd and OpenRC are different
models and LoadState has no equivalent. Container runtimes converged on one CLI
deliberately, so almost nothing is lost: checked against podman 6.1.0, run,
rm -f and docker's own template syntax for state and labels all work unchanged.
Only the probe differs. So: a two-entry lookup, not an interface.

The difference that is NOT in the CLI is the one that would have shipped
silently. Podman accepts --restart unless-stopped, records it, and has no
daemon to act on it -- containers do not return after a reboot unless
podman-restart.service is enabled, which by default it is not. Every command
reports success and the effect does not happen.

That belongs in the declaration rather than the host: a node using podman is
told to enable the unit. Which is what made the service shape's missing 'boot'
field visible, and it is now built.
2026-08-27 23:59:05 +02:00

7.1 KiB

status, date, deciders, reconstructed, extends
status date deciders reconstructed extends
accepted 2026-08-28 jochen false 0057-the-host-is-a-root-service-installed-as-a-package.md

60. The host is built per operating system

Context

Three of the host's six shapes need something from the machine: service needs a service manager, package needs a package manager, container needs a container runtime. The other three — file, directory, action — need only a filesystem and the ability to run something.

The host names those capabilities generically and implements them specifically. The detector reports container-runtime, package-manager, service-manager; the appliers call docker, pacman and systemctl. So the design says capability and the code says Arch, and nothing records which of those is the intent.

The question that surfaced it: what happens on a machine that has podman, or one that does not run systemd? And underneath it, a live contradiction — research 012 decides that on conflict during adoption the machine's configuration is kept, so a machine with podman keeps podman, and then the container applier calls docker and fails.

Considered options

  1. Abstract each capability behind an interface. One host, adapters per service manager and package manager. Rejected, and the reason is correctness rather than effort: the service applier reads LoadState to tell not installed apart from stopped, which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both can express, and the lowest common denominator is exactly where that fault lives.
  2. Support one operating system and say so. Honest, and it makes every other machine permanently out of scope rather than not-yet.
  3. A host per operating system. Chosen.

Decision

The host is built for an operating system family, and systemd and pacman are the Arch host's implementation rather than abstractions the mesh has to grow.

mesh-host-arch-x86_64        pacman · systemctl · docker
mesh-host-debian-x86_64      apt    · systemctl · docker      (when there is a machine)
mesh-host-alpine-x86_64      apk    · rc-service · podman     (when there is a machine)

These are not independent choices and treating them as such was the error. A machine has pacman because it is Arch. The package manager, the service manager and the packaging format arrive together, as one decision somebody made when they installed the operating system.

Almost all of it is shared

Not a rewrite per operating system. The declaration vocabulary, the store, the apply loop, the read-back discipline, the refusal model, the bundle and the link are all portable. What differs is two appliers, and the rest is compiled around them.

The control plane names the package, because the host does not decide

A container runtime is docker on Arch and docker.io on Debian. Mapping this node needs a container runtime to a package name is deciding, which ADR 0037 puts outside the host.

It needs no new mechanism: the profile already reports what the machine is, so the declaration a node receives is already tailored to that node. The host receives a package name and installs it.

A host that cannot implement a shape refuses it

The interesting case is not Debian, it is Android — no service manager it will lend us, no package installation, usually no root. Such a host implements file, directory and action, and nothing else.

That needs no new mechanism either. A host already refuses a type it does not know; this host does not implement package is the same refusal with a different reason, and the profile reports which shapes it implements so the control plane never sends one it cannot do.

file, directory and action are the portable floor. They work anywhere there is a filesystem and a way to run something, which makes a partial host a real thing rather than a broken one.

Consequences

  • Each implementation stays as sharp as its operating system allows. The LoadState distinction survives because the Arch host knows it is systemd. Nothing is degraded to fit an interface spanning systems we do not run.

  • Delivery already worked this way, which is the strongest sign this is the right seam: the host ships as a package from the mesh's own repository (ADR 0058), and a .pkg.tar.zst is an Arch artifact. A per-OS binary is consistent with what was already decided rather than an addition to it.

  • The container runtime is a choice within a host, not an OS split — Arch runs docker or podman — and it is detected, not declared, because adoption keeps what the machine already has. Two are supported.

    This is the opposite answer to the one above, for a reason rather than by preference. Abstracting service managers is lossy: systemd and OpenRC are different models, and LoadState has no equivalent. Container runtimes deliberately converged on one CLI, so almost nothing is lost — checked against podman 6.1.0, run, rm -f and docker's own template syntax for reading state and labels all work unchanged. Only the probe differs ({{.ServerVersion}} against {{.Version.Version}}), which makes it a two-entry lookup rather than an interface.

    One difference is not in the CLI and would have shipped silently. Podman accepts --restart unless-stopped and records it, and has no daemon to act on it: containers do not come back after a reboot unless podman-restart.service is enabled, which it is not by default. Every command reports success and the effect does not happen. That belongs in the declaration — a node using podman is told to enable that unit — rather than in the host, which keeps the host dumb and puts the difference where a person can read it.

  • A second operating system is now additive rather than a redesign — write two appliers, ship a package. And it will be designed against a real machine rather than a guess, which is the point of not building the abstraction now.

  • The profile has to report the operating system, and today it reports capabilities without saying which system they belong to. Small, and needed before the control plane can tailor a package name.

  • Nothing states the machine requirements yet. A machine missing a capability fails at apply time rather than being refused up front, even though the host already detects it. That is a gap this record makes visible and does not close.

References

  • ADR 0037 — why the host is handed a package name.
  • ADR 0041 — one static binary, now per system as well as per architecture.
  • ADR 0058 — delivery, which was already per-OS.
  • Research 012 — adoption keeping the machine's configuration, which hardcoding a runtime contradicts.