Issue 114: should the controller be a container or a process the host supervises #100
+73
@@ -0,0 +1,73 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller module.json, mesh-host internal/apply]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 114 — Should the controller run as a container, or as a process the host supervises directly?
|
||||
|
||||
## What was observed
|
||||
|
||||
On the control-node, 2026-09-24, over a long session of operating the mesh through
|
||||
`mesh-controller`'s CLI (build, push, plan, status, module moved). Every mutating step reached the
|
||||
binary the same way: `docker exec mesh-controller /mesh-controller <command>` — because
|
||||
`mesh-controller`'s own manifest declares its one resource as:
|
||||
|
||||
```json
|
||||
{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }
|
||||
```
|
||||
|
||||
Two things about that declaration are worth naming together, because neither is a problem on its
|
||||
own and the combination is what raises the question:
|
||||
|
||||
- **`network: host`.** The controller does not use container network isolation, which is the
|
||||
property a `container` resource type usually buys over a `process` one. It runs with the node's
|
||||
own network namespace either way.
|
||||
- **It is the mesh's single point of coordination.** [`03-DESIGN/01-to-be/06-the-controller.md`](../../03-DESIGN/01-to-be/06-the-controller.md)
|
||||
is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover;
|
||||
recovery is restore, not failover.
|
||||
|
||||
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) gives the host — the one thing tier 0 requires
|
||||
to be a real system daemon — exactly this reasoning for refusing to run in a container: *"installing
|
||||
the container runtime is a step of the bootstrap, so a host inside a container would need the thing
|
||||
it exists to install."* The controller is one tier up and does not install the runtime, but it
|
||||
shares the profile that argument turns on: something the rest of the mesh's operation depends on,
|
||||
sharing fate with a runtime that is not itself.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
Practically, tonight: every controller interaction was raw shell into a container (`docker exec`),
|
||||
not a first-class surface — no logs command beyond `docker logs`, no `systemctl status`, and a
|
||||
session permission classifier that (correctly) treats arbitrary shell into a container as needing
|
||||
sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the
|
||||
issue itself.
|
||||
|
||||
The actual question is whether `type: container` is buying the controller anything here besides
|
||||
image-based delivery and a restart policy — both of which [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)'s
|
||||
launcher pattern already describes as buildable directly into the host's own supervision (restart on
|
||||
exit, count consecutive failures, roll back after too many, halt after that), for the host's own
|
||||
unit. If the controller were declared `type: process` instead — still built and versioned through
|
||||
the same delivery pipeline, just executed on the node and supervised by the host the way the host
|
||||
supervises itself — it would stop sharing fate with the container runtime's health (restarts,
|
||||
upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of
|
||||
the mesh is designed to tolerate but nothing is designed to *want*.
|
||||
|
||||
This is squarely a question, not a claim that today's shape is wrong: [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
|
||||
already tolerates the controller being down by construction (nodes reconcile from their own
|
||||
last-applied state), which may make the container-runtime coupling moot in practice. Nobody has
|
||||
checked.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Does `mesh-host`'s `process` resource type already support the restart/failure-counting semantics
|
||||
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher well
|
||||
enough for something this central — or would this need host-side work first?
|
||||
- With `network: host` already in use, what does `type: container` provide the controller today that
|
||||
`type: process` would not?
|
||||
- Is there a real circularity risk — the controller's own health depending on the container runtime
|
||||
it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make
|
||||
a controller outage tolerable regardless of which resource type it is?
|
||||
- If the answer is "keep it a container," what does that answer, precisely, that this issue asked —
|
||||
so the next person who notices the same asymmetry finds it answered rather than open again?
|
||||
Reference in New Issue
Block a user