From da41c1cc40820f7cfa61f8f842b804dd10601330 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 24 Sep 2026 11:41:55 +0200 Subject: [PATCH] Issue 113: should the controller be a container or a process the host supervises MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Filed after a session where every mesh-controller interaction went through docker exec — its manifest runs it as a container with network: host, using none of the isolation that resource type usually buys, while ADR 0006 makes it the mesh's single point of coordination. Open question, not a claimed defect: does type: container get the controller anything type: process (supervised the way the host supervises its own unit, per ADR 0005) would not. --- .../00-report.md | 73 +++++++++++++++++++ 1 file changed, 73 insertions(+) create mode 100644 04-ISSUES/113-should-the-controller-be-a-container-or-a-process-the-host-supervises/00-report.md diff --git a/04-ISSUES/113-should-the-controller-be-a-container-or-a-process-the-host-supervises/00-report.md b/04-ISSUES/113-should-the-controller-be-a-container-or-a-process-the-host-supervises/00-report.md new file mode 100644 index 0000000..ce0e98c --- /dev/null +++ b/04-ISSUES/113-should-the-controller-be-a-container-or-a-process-the-host-supervises/00-report.md @@ -0,0 +1,73 @@ +--- +status: open +opened: 2026-09-24 +located-in: [mesh-controller module.json, mesh-host internal/apply] +fixed-by: +amended-design: +--- + +# 113 — Should the controller run as a container, or as a process the host supervises directly? + +## What was observed + +On the control-node, 2026-09-24, over a long session of operating the mesh through +`mesh-controller`'s CLI (build, push, plan, status, module moved). Every mutating step reached the +binary the same way: `docker exec mesh-controller /mesh-controller ` — because +`mesh-controller`'s own manifest declares its one resource as: + +```json +{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] } +``` + +Two things about that declaration are worth naming together, because neither is a problem on its +own and the combination is what raises the question: + +- **`network: host`.** The controller does not use container network isolation, which is the + property a `container` resource type usually buys over a `process` one. It runs with the node's + own network namespace either way. +- **It is the mesh's single point of coordination.** [`03-DESIGN/01-to-be/06-the-controller.md`](../../03-DESIGN/01-to-be/06-the-controller.md) + is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover; + recovery is restore, not failover. + +[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) gives the host — the one thing tier 0 requires +to be a real system daemon — exactly this reasoning for refusing to run in a container: *"installing +the container runtime is a step of the bootstrap, so a host inside a container would need the thing +it exists to install."* The controller is one tier up and does not install the runtime, but it +shares the profile that argument turns on: something the rest of the mesh's operation depends on, +sharing fate with a runtime that is not itself. + +## Why it matters beyond this instance + +Practically, tonight: every controller interaction was raw shell into a container (`docker exec`), +not a first-class surface — no logs command beyond `docker logs`, no `systemctl status`, and a +session permission classifier that (correctly) treats arbitrary shell into a container as needing +sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the +issue itself. + +The actual question is whether `type: container` is buying the controller anything here besides +image-based delivery and a restart policy — both of which [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)'s +launcher pattern already describes as buildable directly into the host's own supervision (restart on +exit, count consecutive failures, roll back after too many, halt after that), for the host's own +unit. If the controller were declared `type: process` instead — still built and versioned through +the same delivery pipeline, just executed on the node and supervised by the host the way the host +supervises itself — it would stop sharing fate with the container runtime's health (restarts, +upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of +the mesh is designed to tolerate but nothing is designed to *want*. + +This is squarely a question, not a claim that today's shape is wrong: [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md) +already tolerates the controller being down by construction (nodes reconcile from their own +last-applied state), which may make the container-runtime coupling moot in practice. Nobody has +checked. + +## Open questions + +- Does `mesh-host`'s `process` resource type already support the restart/failure-counting semantics + [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher well + enough for something this central — or would this need host-side work first? +- With `network: host` already in use, what does `type: container` provide the controller today that + `type: process` would not? +- Is there a real circularity risk — the controller's own health depending on the container runtime + it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make + a controller outage tolerable regardless of which resource type it is? +- If the answer is "keep it a container," what does that answer, precisely, that this issue asked — + so the next person who notices the same asymmetry finds it answered rather than open again?