Files
jochen 82a6badc7c Issue 114: land the controller's container-or-process question, renumbered
Filed 2026-09-24 on a branch of its own and never merged, numbered 113, which is taken. 114 is
free because a sibling branch folded it, so it takes that number and keeps its commit.

Kept separate from issue 117 rather than folded into it. 117 asks the same question of every
module and locates the missing decision; this asks it of the controller, where `network: host`
means container network isolation — the property that resource type usually buys — is not in use.
That observation is this report's own and is nowhere in 117, and folding would lose it.

Its first open question is answered by 117's diagnosis and now says so: the host's `process` shape
is built, applied and tested, restart and run-to-completion semantics included, so deciding this
does not wait on host-side work.
2026-09-25 16:16:21 +02:00

5.2 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-24
mesh-controller module.json
mesh-host internal/apply

114 — Should the controller run as a container, or as a process the host supervises directly?

What was observed

On the control-node, 2026-09-24, over a long session of operating the mesh through mesh-controller's CLI (build, push, plan, status, module moved). Every mutating step reached the binary the same way: docker exec mesh-controller /mesh-controller <command> — because mesh-controller's own manifest declares its one resource as:

{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }

Two things about that declaration are worth naming together, because neither is a problem on its own and the combination is what raises the question:

  • network: host. The controller does not use container network isolation, which is the property a container resource type usually buys over a process one. It runs with the node's own network namespace either way.
  • It is the mesh's single point of coordination. 03-DESIGN/01-to-be/06-the-controller.md is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover; recovery is restore, not failover.

ADR 0005 gives the host — the one thing tier 0 requires to be a real system daemon — exactly this reasoning for refusing to run in a container: "installing the container runtime is a step of the bootstrap, so a host inside a container would need the thing it exists to install." The controller is one tier up and does not install the runtime, but it shares the profile that argument turns on: something the rest of the mesh's operation depends on, sharing fate with a runtime that is not itself.

Why it matters beyond this instance

Practically, tonight: every controller interaction was raw shell into a container (docker exec), not a first-class surface — no logs command beyond docker logs, no systemctl status, and a session permission classifier that (correctly) treats arbitrary shell into a container as needing sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the issue itself.

The actual question is whether type: container is buying the controller anything here besides image-based delivery and a restart policy — both of which ADR 0005's launcher pattern already describes as buildable directly into the host's own supervision (restart on exit, count consecutive failures, roll back after too many, halt after that), for the host's own unit. If the controller were declared type: process instead — still built and versioned through the same delivery pipeline, just executed on the node and supervised by the host the way the host supervises itself — it would stop sharing fate with the container runtime's health (restarts, upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of the mesh is designed to tolerate but nothing is designed to want.

This is squarely a question, not a claim that today's shape is wrong: ADR 0006 already tolerates the controller being down by construction (nodes reconcile from their own last-applied state), which may make the container-runtime coupling moot in practice. Nobody has checked.

Open questions

  • Does mesh-host's process resource type already support the restart/failure-counting semantics ADR 0005 describes for the host's own launcher well enough for something this central — or would this need host-side work first?
  • With network: host already in use, what does type: container provide the controller today that type: process would not?
  • Is there a real circularity risk — the controller's own health depending on the container runtime it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make a controller outage tolerable regardless of which resource type it is?
  • If the answer is "keep it a container," what does that answer, precisely, that this issue asked — so the next person who notices the same asymmetry finds it answered rather than open again?

The general case

Issue 117 is the same question asked of every module rather than of the controller: a module's own code is a container in ADR 0047 and a process in the to-be design, and no record moves it. Its diagnosis answers the first open question above: the host's process shape is built, applied and tested, including the restart and run-to-completion semantics — so this would not need host-side work first.

The two do not collapse into one. The controller is not a code-carrying sidecar, and network: host is what makes the asymmetry visible here and nowhere else.