It asked a question rather than reporting a defect, and the question was taken five days later — the mesh's own components are binaries on the machine, and third-party software stays a container because an image is the right way to carry somebody else's build. The delivery of them is issue 142.
105 lines
6.3 KiB
Markdown
105 lines
6.3 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-24
|
|
located-in: [mesh-controller module.json, mesh-host internal/apply]
|
|
fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
|
|
amended-design:
|
|
---
|
|
|
|
# 114 — Should the controller run as a container, or as a process the host supervises directly?
|
|
|
|
## What was observed
|
|
|
|
On the control-node, 2026-09-24, over a long session of operating the mesh through
|
|
`mesh-controller`'s CLI (build, push, plan, status, module moved). Every mutating step reached the
|
|
binary the same way: `docker exec mesh-controller /mesh-controller <command>` — because
|
|
`mesh-controller`'s own manifest declares its one resource as:
|
|
|
|
```json
|
|
{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }
|
|
```
|
|
|
|
Two things about that declaration are worth naming together, because neither is a problem on its
|
|
own and the combination is what raises the question:
|
|
|
|
- **`network: host`.** The controller does not use container network isolation, which is the
|
|
property a `container` resource type usually buys over a `process` one. It runs with the node's
|
|
own network namespace either way.
|
|
- **It is the mesh's single point of coordination.** [`03-DESIGN/01-to-be/06-the-controller.md`](../../03-DESIGN/01-to-be/06-the-controller.md)
|
|
is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover;
|
|
recovery is restore, not failover.
|
|
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) gives the host — the one thing tier 0 requires
|
|
to be a real system daemon — exactly this reasoning for refusing to run in a container: *"installing
|
|
the container runtime is a step of the bootstrap, so a host inside a container would need the thing
|
|
it exists to install."* The controller is one tier up and does not install the runtime, but it
|
|
shares the profile that argument turns on: something the rest of the mesh's operation depends on,
|
|
sharing fate with a runtime that is not itself.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
Practically, tonight: every controller interaction was raw shell into a container (`docker exec`),
|
|
not a first-class surface — no logs command beyond `docker logs`, no `systemctl status`, and a
|
|
session permission classifier that (correctly) treats arbitrary shell into a container as needing
|
|
sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the
|
|
issue itself.
|
|
|
|
The actual question is whether `type: container` is buying the controller anything here besides
|
|
image-based delivery and a restart policy — both of which [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)'s
|
|
launcher pattern already describes as buildable directly into the host's own supervision (restart on
|
|
exit, count consecutive failures, roll back after too many, halt after that), for the host's own
|
|
unit. If the controller were declared `type: process` instead — still built and versioned through
|
|
the same delivery pipeline, just executed on the node and supervised by the host the way the host
|
|
supervises itself — it would stop sharing fate with the container runtime's health (restarts,
|
|
upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of
|
|
the mesh is designed to tolerate but nothing is designed to *want*.
|
|
|
|
This is squarely a question, not a claim that today's shape is wrong: [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
|
|
already tolerates the controller being down by construction (nodes reconcile from their own
|
|
last-applied state), which may make the container-runtime coupling moot in practice. Nobody has
|
|
checked.
|
|
|
|
## Open questions
|
|
|
|
- Does `mesh-host`'s `process` resource type already support the restart/failure-counting semantics
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher well
|
|
enough for something this central — or would this need host-side work first?
|
|
- With `network: host` already in use, what does `type: container` provide the controller today that
|
|
`type: process` would not?
|
|
- Is there a real circularity risk — the controller's own health depending on the container runtime
|
|
it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make
|
|
a controller outage tolerable regardless of which resource type it is?
|
|
- If the answer is "keep it a container," what does that answer, precisely, that this issue asked —
|
|
so the next person who notices the same asymmetry finds it answered rather than open again?
|
|
|
|
## The general case
|
|
|
|
[Issue 117](../117-a-modules-own-code-is-a-container-and-a-process/00-report.md) is the same
|
|
question asked of every module rather than of the controller: a module's own code is a `container`
|
|
in [ADR 0047](../../02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md)
|
|
and a `process` in the to-be design, and no record moves it. Its
|
|
[diagnosis](../117-a-modules-own-code-is-a-container-and-a-process/01-diagnosis.md) answers the
|
|
first open question above: the host's `process` shape is built, applied and tested, including the
|
|
restart and run-to-completion semantics — so this would not need host-side work first.
|
|
|
|
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
|
|
is what makes the asymmetry visible here and nowhere else.
|
|
|
|
## Answered
|
|
|
|
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
|
|
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
|
|
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
|
|
vault — are **binaries on the machine**, delivered by the mechanism
|
|
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
|
|
third-party software (the store, the registry, the broker) stays a container because an image is the
|
|
right way to carry somebody else's build.
|
|
|
|
So the operating experience this record was written from — every mutating command reached through
|
|
`docker exec mesh-controller` — is answered, and answered against the container.
|
|
|
|
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
|
|
component travels yet; that is
|
|
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
|
|
|