The controller's private network writes insecure-registries into daemon.json, a file this
module owns. The runtime's module states it instead, through ${seat:mesh-artifact-store:reach}
(hq ADR 0222), so the controller can stop generating its registry-trust resources.
166 lines
9.8 KiB
Markdown
166 lines
9.8 KiB
Markdown
# docker
|
||
|
||
The container runtime as a module (novox/hq to-be 42 phase 1, item 8; research 027/01–02; ADR 0166,
|
||
ADR 0207). It claims the node seat `node-container-runtime`. That seat carries no verbs yet: its verbs,
|
||
and the host creating containers through its holder, wait on ADR 0166's acceptance. Until then the
|
||
tools below are the module's own.
|
||
|
||
## What it declares
|
||
|
||
| resource | what | the host's rule |
|
||
|---|---|---|
|
||
| `package` | `docker` | installed if absent; never uninstalled when the module goes |
|
||
| `buildx` | `docker-buildx` | the same. Only the build machine has it today; `docker build` needs it for BuildKit everywhere |
|
||
| `socket` | `docker.socket` running, enabled at boot | given back as found when the module goes (ADR 0118) |
|
||
| `prune-service`, `prune-timer` | `/etc/systemd/system/docker-prune.{service,timer}`, written whole | removed with the module |
|
||
| `prune` | `docker-prune.timer` running, enabled at boot; restarted when either file changes | stopped and disabled with the module (the mesh made the unit) |
|
||
| `daemon` | `live-restore` and the mesh's registry in `insecure-registries` written into `/etc/docker/daemon.json`, beside other keys | each key given back as found when the module goes; only the member this module added leaves the list |
|
||
| `runtime` | `docker.service` running, enabled at boot; reloaded, never restarted, when `daemon` changes | given back as found (ADR 0118) |
|
||
|
||
The weekly prune takes **dangling images and build cache unused for a week, and nothing else**. It
|
||
takes no volume, no container and no image a container uses, so it never touches a container the mesh
|
||
holds. It runs at idle priority, at a random point in the hour after the weekly mark. A run missed
|
||
while the machine was off happens at the next boot.
|
||
|
||
**Capabilities:** `package-manager`, `service-manager`, `privileged`. It does not declare
|
||
`container-runtime`: under ADR 0165, which is still proposed, that word means a running daemon, and
|
||
the module that installs the daemon cannot require it.
|
||
|
||
## The runtime's own file and service (issue 190, hq ADR 0196, ADR 0222)
|
||
|
||
This module writes two keys into `/etc/docker/daemon.json` (`into: json`, ADR 0102): `live-restore`
|
||
and `insecure-registries`. `dnsmasq` used to write `live-restore`, beside `dns`; it no longer writes
|
||
either. Under ADR 0196 a container copies its machine's resolvers, so no module writes `dns`. The
|
||
controller's private network wrote `insecure-registries`; under ADR 0222 the controller writes
|
||
nothing into this file, and this module states the registry itself (the order below).
|
||
|
||
- `daemon`: `{"live-restore": true, "insecure-registries": ["${seat:mesh-artifact-store:reach}"]}`,
|
||
merged into the file beside the keys others write.
|
||
- `${seat:mesh-artifact-store:reach}` is where this machine reaches the mesh's artifact store
|
||
(host:port), filled in by the controller: no binding, no credential, the same address the mesh
|
||
composes into every image it built. Trusting it in the clear is ADR 0082's decision: every path to
|
||
it is inside the private network's encryption. While no machine on the network holds the store the
|
||
answer is empty, and the controller drops the empty member, so the list gets nothing.
|
||
- `insecure-registries` is a list, and the host adds to it rather than replacing it: a machine's own
|
||
trusted registries stay, and undeclaring takes out only the member this module added.
|
||
- `runtime`: `docker.service` running, enabled at boot, and **reloaded, never restarted**, when
|
||
`daemon` changes. A restart stops every container. A reload turns `live-restore` on and takes the
|
||
trusted registries, and with `live-restore` on a later restart keeps every container running.
|
||
|
||
In the apply that moves the key, the host first gives back `dnsmasq`'s resources, then applies this
|
||
module's: `live-restore` is set again in the same apply, and the daemon is reloaded once.
|
||
|
||
`dns` is removed from the file then, but the daemon reads it only at its next start, and a running
|
||
container keeps the resolvers it was created with. Each container pinned to a machine's own resolver
|
||
is restarted before that machine's `dnsmasq` goes (ADR 0194, step 4).
|
||
|
||
**The order it lands in.** The controller that fills `${seat:…:reach}` is deployed first: one that
|
||
does not know the placeholder would send it through unfilled. Then this module. Then the controller
|
||
stops generating the private network's `registry-trust` and `registry-trust-reload` and refuses a
|
||
generated resource that collides with a module's (issue 190, steps 2 and 5). In the apply that moves
|
||
the member, the host removes the private network's record first (the member leaves the list) and
|
||
then applies this module's (it is added back, recorded as this module's); the daemon is reloaded
|
||
once, for `daemon`. The address is the same one, so the runtime's trust does not change.
|
||
|
||
Still elsewhere:
|
||
|
||
- **Nobody** writes log rotation. One machine has `log-driver` and `log-opts` by hand; they are left
|
||
as they are until a size is chosen for every machine.
|
||
|
||
## What it does not declare yet, and why
|
||
|
||
One thing this module should own is still declared elsewhere. The controller refuses two modules
|
||
on one node that declare the same `path`, `unit`, `name` or `package` (`checkResources`,
|
||
mesh-controller `internal/catalogue/resolve.go`), so declaring it here would make the module
|
||
unassignable everywhere. The refusals were checked against the controller's own
|
||
check:
|
||
|
||
```
|
||
zsh and docker both declare the name "${machine:account}"
|
||
```
|
||
|
||
### The operator account's membership of the `docker` group
|
||
|
||
The right shape is the host's `user` shape. Its `groups` are additive: the host runs
|
||
`usermod --append` and never takes a group away.
|
||
|
||
```json
|
||
{"id": "group", "type": "user", "name": "${machine:account}", "groups": ["docker"]}
|
||
```
|
||
|
||
`zsh` already declares a `user` resource for the same account (its login shell). The controller
|
||
compares `name` across modules, so the two collide.
|
||
|
||
**The change proposed (mesh-controller, `checkResources`):** judge a `user` resource by the fields it
|
||
sets, not by its name:
|
||
|
||
- `shell` and `home` stay single-owner;
|
||
- `groups` may be declared by any number of modules, because the host only adds them.
|
||
|
||
Then this module declares the resource above, and no module has to carry another's group.
|
||
|
||
Today the operator account is in the group on every machine, by hand. Nothing is lost while it waits.
|
||
|
||
## The bootstrap's runtime
|
||
|
||
On the machine the mesh was first installed on, the foundation bundle declared `package docker`
|
||
(`container-runtime`) and `docker.service` running and enabled (`container-runtime-running`). ADR 0207
|
||
§5 exempts them.
|
||
|
||
- The host records them under their bare ids, with origin *carried*. A mesh declaration's orphan pass
|
||
never sees them (mesh-host `store.go`).
|
||
- So `docker.package` here is a **second record of the same package**. The apply says "already
|
||
installed", and neither record ever uninstalls it.
|
||
- `docker.runtime` is likewise a second record of `docker.service`. Its found state is *running*,
|
||
because genesis started it, so undeclaring this module leaves the daemon running.
|
||
|
||
## Tools
|
||
|
||
The tools run as the operator account. If the daemon's socket refuses that account, a call is asked
|
||
again through `sudo -n` (a process keeps the groups it started with). Every call has a 20 s bound.
|
||
A failure is an error naming how it failed, never an empty answer.
|
||
|
||
**Every container on the machine is in scope.** A container the mesh holds carries the host's label
|
||
`mesh-host.id` (its value names the assignment), and every answer says `mesh_held`.
|
||
|
||
| tool | | what |
|
||
|---|---|---|
|
||
| `docker_list` | r | every container: image, state, health, restarts, ports, mounts, compose project, `mesh_held`; filter by owner, state or name |
|
||
| `docker_inspect` | r | one container whole, **environment values left out** (names kept) |
|
||
| `docker_logs` | r | the last lines of both streams, merged in order, with timestamps (default 200, at most 2000) |
|
||
| `docker_stats` | r | CPU, memory, I/O and process count per running container, heaviest first |
|
||
| `docker_start` / `docker_stop` / `docker_restart` | a | one container. On a mesh-held one, the answer says the host restores its declared state at its next apply |
|
||
| `docker_top` | r | the processes inside one container |
|
||
| `docker_images` | r | images, largest first, with the containers using each; `dangling`, `unused` or `used` |
|
||
| `docker_prune` | a | dangling images and build cache, and stopped containers the mesh does not hold if `containers` is true. **A dry run unless `dry_run` is false. Never a volume** |
|
||
| `docker_disk_usage` | r | `docker system df -v`: total, active and reclaimable per kind, with the largest of each |
|
||
| `docker_networks` | r | networks, subnets, and the containers on each |
|
||
| `docker_volumes` | r | volumes, who mounts each, whether the mesh holds one of them, anonymous or not, and sizes if asked |
|
||
| `docker_events` | r | the runtime's events over a window ending now (default 60 min, at most 24 h), without exec noise |
|
||
| `docker_daemon_config` | r | `daemon.json` as on disk, `docker info`'s essentials, and keys the daemon has not taken yet |
|
||
| `docker_unlabelled` | r | the containers the mesh does not hold: the cleanup list |
|
||
| `docker_problems` | r | unhealthy, restarting, dead, killed for memory, failed, or restarted five times or more |
|
||
| `docker_ports` | r | every published port, and the containers on the host's network |
|
||
|
||
## Tests
|
||
|
||
```
|
||
go test ./...
|
||
```
|
||
|
||
The tests run against a fake runner and cover:
|
||
|
||
- escalation through `sudo -n` on a refused socket, and never as root;
|
||
- each failure named by its cause;
|
||
- a name or id never read as an option;
|
||
- mesh-held marking;
|
||
- the environment left out of `inspect`;
|
||
- the restore note on a mesh-held act;
|
||
- prune being a dry run by default and never reaching a volume, a mesh container or `--volumes`;
|
||
- the log merge;
|
||
- size parsing;
|
||
- what the daemon has not yet taken;
|
||
- event filtering;
|
||
- volume ownership;
|
||
- that the tools served are exactly the manifest's `tools`.
|