Issue 113: should the controller be a container or a process the host supervises

Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
This commit is contained in:
2026-09-25 16:15:27 +02:00
committed by jochen
parent 35db2aaa41
commit 10a2b706c6
@@ -0,0 +1,73 @@
---
status: open
opened: 2026-09-24
located-in: [mesh-controller module.json, mesh-host internal/apply]
fixed-by:
amended-design:
---
# 113 — Should the controller run as a container, or as a process the host supervises directly?
## What was observed
On the control-node, 2026-09-24, over a long session of operating the mesh through
`mesh-controller`'s CLI (build, push, plan, status, module moved). Every mutating step reached the
binary the same way: `docker exec mesh-controller /mesh-controller <command>` — because
`mesh-controller`'s own manifest declares its one resource as:
```json
{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }
```
Two things about that declaration are worth naming together, because neither is a problem on its
own and the combination is what raises the question:
- **`network: host`.** The controller does not use container network isolation, which is the
property a `container` resource type usually buys over a `process` one. It runs with the node's
own network namespace either way.
- **It is the mesh's single point of coordination.** [`03-DESIGN/01-to-be/06-the-controller.md`](../../03-DESIGN/01-to-be/06-the-controller.md)
is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover;
recovery is restore, not failover.
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) gives the host — the one thing tier 0 requires
to be a real system daemon — exactly this reasoning for refusing to run in a container: *"installing
the container runtime is a step of the bootstrap, so a host inside a container would need the thing
it exists to install."* The controller is one tier up and does not install the runtime, but it
shares the profile that argument turns on: something the rest of the mesh's operation depends on,
sharing fate with a runtime that is not itself.
## Why it matters beyond this instance
Practically, tonight: every controller interaction was raw shell into a container (`docker exec`),
not a first-class surface — no logs command beyond `docker logs`, no `systemctl status`, and a
session permission classifier that (correctly) treats arbitrary shell into a container as needing
sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the
issue itself.
The actual question is whether `type: container` is buying the controller anything here besides
image-based delivery and a restart policy — both of which [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)'s
launcher pattern already describes as buildable directly into the host's own supervision (restart on
exit, count consecutive failures, roll back after too many, halt after that), for the host's own
unit. If the controller were declared `type: process` instead — still built and versioned through
the same delivery pipeline, just executed on the node and supervised by the host the way the host
supervises itself — it would stop sharing fate with the container runtime's health (restarts,
upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of
the mesh is designed to tolerate but nothing is designed to *want*.
This is squarely a question, not a claim that today's shape is wrong: [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
already tolerates the controller being down by construction (nodes reconcile from their own
last-applied state), which may make the container-runtime coupling moot in practice. Nobody has
checked.
## Open questions
- Does `mesh-host`'s `process` resource type already support the restart/failure-counting semantics
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher well
enough for something this central — or would this need host-side work first?
- With `network: host` already in use, what does `type: container` provide the controller today that
`type: process` would not?
- Is there a real circularity risk — the controller's own health depending on the container runtime
it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make
a controller outage tolerable regardless of which resource type it is?
- If the answer is "keep it a container," what does that answer, precisely, that this issue asked —
so the next person who notices the same asymmetry finds it answered rather than open again?