Files
hq/04-ISSUES/114-should-the-controller-be-a-container-or-a-process-the-host-supervises/00-report.md
T
jschoubben a97feeefe5 Renumber to 114: 113 collided with an issue merged independently to main
Both used the next number available when opened, and the object-store
withdrawal report merged first. No content change beyond the number.
2026-09-24 15:48:56 +02:00

4.4 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-24
mesh-controller module.json
mesh-host internal/apply

114 — Should the controller run as a container, or as a process the host supervises directly?

What was observed

On the control-node, 2026-09-24, over a long session of operating the mesh through mesh-controller's CLI (build, push, plan, status, module moved). Every mutating step reached the binary the same way: docker exec mesh-controller /mesh-controller <command> — because mesh-controller's own manifest declares its one resource as:

{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }

Two things about that declaration are worth naming together, because neither is a problem on its own and the combination is what raises the question:

  • network: host. The controller does not use container network isolation, which is the property a container resource type usually buys over a process one. It runs with the node's own network namespace either way.
  • It is the mesh's single point of coordination. 03-DESIGN/01-to-be/06-the-controller.md is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover; recovery is restore, not failover.

ADR 0005 gives the host — the one thing tier 0 requires to be a real system daemon — exactly this reasoning for refusing to run in a container: "installing the container runtime is a step of the bootstrap, so a host inside a container would need the thing it exists to install." The controller is one tier up and does not install the runtime, but it shares the profile that argument turns on: something the rest of the mesh's operation depends on, sharing fate with a runtime that is not itself.

Why it matters beyond this instance

Practically, tonight: every controller interaction was raw shell into a container (docker exec), not a first-class surface — no logs command beyond docker logs, no systemctl status, and a session permission classifier that (correctly) treats arbitrary shell into a container as needing sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the issue itself.

The actual question is whether type: container is buying the controller anything here besides image-based delivery and a restart policy — both of which ADR 0005's launcher pattern already describes as buildable directly into the host's own supervision (restart on exit, count consecutive failures, roll back after too many, halt after that), for the host's own unit. If the controller were declared type: process instead — still built and versioned through the same delivery pipeline, just executed on the node and supervised by the host the way the host supervises itself — it would stop sharing fate with the container runtime's health (restarts, upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of the mesh is designed to tolerate but nothing is designed to want.

This is squarely a question, not a claim that today's shape is wrong: ADR 0006 already tolerates the controller being down by construction (nodes reconcile from their own last-applied state), which may make the container-runtime coupling moot in practice. Nobody has checked.

Open questions

  • Does mesh-host's process resource type already support the restart/failure-counting semantics ADR 0005 describes for the host's own launcher well enough for something this central — or would this need host-side work first?
  • With network: host already in use, what does type: container provide the controller today that type: process would not?
  • Is there a real circularity risk — the controller's own health depending on the container runtime it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make a controller outage tolerable regardless of which resource type it is?
  • If the answer is "keep it a container," what does that answer, precisely, that this issue asked — so the next person who notices the same asymmetry finds it answered rather than open again?