mesh/merge-gate pass: builds docker → ace, g14, novox, shanks; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery superseded: a newer delivery to the same trunk took over its walk
Images pulled by digest are not dangling, so the weekly prune never takes an old version and every machine keeps every image it ever ran. docker_prune_images takes them, keeps what the declaration names (asked of the controller; no answer, nothing removed) and the one before, and is a dry run unless asked with a why.
248 lines
16 KiB
Markdown
248 lines
16 KiB
Markdown
# docker
|
||
|
||
The container runtime as a module (novox/hq to-be 42 phase 1, item 8; research 027/01–02; ADR 0166,
|
||
ADR 0207). It claims the node seat `node-container-runtime`. That seat carries no verbs yet: its verbs,
|
||
and the host creating containers through its holder, wait on ADR 0166's acceptance. Until then the
|
||
tools below are the module's own.
|
||
|
||
## What it declares
|
||
|
||
| resource | what | the host's rule |
|
||
|---|---|---|
|
||
| `package` | `docker` | installed if absent; never uninstalled when the module goes |
|
||
| `buildx` | `docker-buildx` | the same. Only the build machine has it today; `docker build` needs it for BuildKit everywhere |
|
||
| `socket` | `docker.socket` running, enabled at boot | given back as found when the module goes (ADR 0118) |
|
||
| `prune-service`, `prune-timer` | `/etc/systemd/system/docker-prune.{service,timer}`, written whole | removed with the module |
|
||
| `prune` | `docker-prune.timer` running, enabled at boot; restarted when either file changes | stopped and disabled with the module (the mesh made the unit) |
|
||
| `daemon` | `live-restore` and the mesh's registry in `insecure-registries` written into `/etc/docker/daemon.json`, beside other keys | each key given back as found when the module goes; only the member this module added leaves the list |
|
||
| `runtime` | `docker.service` running, enabled at boot; reloaded, never restarted, when `daemon` changes | given back as found (ADR 0118) |
|
||
|
||
The weekly prune takes **dangling images and build cache unused for a week, and nothing else**. It
|
||
takes no volume, no container and no image a container uses, so it never touches a container the mesh
|
||
holds. It runs at idle priority, at a random point in the hour after the weekly mark. A run missed
|
||
while the machine was off happens at the next boot.
|
||
|
||
**Capabilities:** `package-manager`, `service-manager`, `privileged`. It does not declare
|
||
`container-runtime`: under ADR 0165, which is still proposed, that word means a running daemon, and
|
||
the module that installs the daemon cannot require it.
|
||
|
||
## The runtime's own file and service (issue 190, hq ADR 0196, ADR 0222)
|
||
|
||
This module writes two keys into `/etc/docker/daemon.json` (`into: json`, ADR 0102): `live-restore`
|
||
and `insecure-registries`. `dnsmasq` used to write `live-restore`, beside `dns`; it no longer writes
|
||
either. Under ADR 0196 a container copies its machine's resolvers, so no module writes `dns`. The
|
||
controller's private network wrote `insecure-registries`; under ADR 0222 the controller writes
|
||
nothing into this file, and this module states the registry itself (the order below).
|
||
|
||
- `daemon`: `{"live-restore": true, "insecure-registries": ["${seat:mesh-artifact-store:reach}"]}`,
|
||
merged into the file beside the keys others write.
|
||
- `${seat:mesh-artifact-store:reach}` is where this machine reaches the mesh's artifact store
|
||
(host:port), filled in by the controller: no binding, no credential, the same address the mesh
|
||
composes into every image it built. Trusting it in the clear is ADR 0082's decision: every path to
|
||
it is inside the private network's encryption. While no machine on the network holds the store the
|
||
answer is empty, and the controller drops the empty member, so the list gets nothing.
|
||
- `insecure-registries` is a list, and the host adds to it rather than replacing it: a machine's own
|
||
trusted registries stay, and undeclaring takes out only the member this module added.
|
||
- `runtime`: `docker.service` running, enabled at boot, and **reloaded, never restarted**, when
|
||
`daemon` changes. A restart stops every container. A reload turns `live-restore` on and takes the
|
||
trusted registries, and with `live-restore` on a later restart keeps every container running.
|
||
|
||
In the apply that moves the key, the host first gives back `dnsmasq`'s resources, then applies this
|
||
module's: `live-restore` is set again in the same apply, and the daemon is reloaded once.
|
||
|
||
`dns` is removed from the file then, but the daemon reads it only at its next start, and a running
|
||
container keeps the resolvers it was created with. Each container pinned to a machine's own resolver
|
||
is restarted before that machine's `dnsmasq` goes (ADR 0194, step 4).
|
||
|
||
**The order it lands in.** The controller that fills `${seat:…:reach}` is deployed first: one that
|
||
does not know the placeholder would send it through unfilled. Then this module. Then the controller
|
||
stops generating the private network's `registry-trust` and `registry-trust-reload` and refuses a
|
||
generated resource that collides with a module's (issue 190, steps 2 and 5). In the apply that moves
|
||
the member, the host removes the private network's record first (the member leaves the list) and
|
||
then applies this module's (it is added back, recorded as this module's); the daemon is reloaded
|
||
once, for `daemon`. The address is the same one, so the runtime's trust does not change.
|
||
|
||
Still elsewhere:
|
||
|
||
- **Nobody** writes log rotation. One machine has `log-driver` and `log-opts` by hand; they are left
|
||
as they are until a size is chosen for every machine.
|
||
|
||
## What it does not declare yet, and why
|
||
|
||
One thing this module should own is still declared elsewhere. The controller refuses two modules
|
||
on one node that declare the same `path`, `unit`, `name` or `package` (`checkResources`,
|
||
mesh-controller `internal/catalogue/resolve.go`), so declaring it here would make the module
|
||
unassignable everywhere. The refusals were checked against the controller's own
|
||
check:
|
||
|
||
```
|
||
zsh and docker both declare the name "${machine:account}"
|
||
```
|
||
|
||
### The operator account's membership of the `docker` group
|
||
|
||
The right shape is the host's `user` shape. Its `groups` are additive: the host runs
|
||
`usermod --append` and never takes a group away.
|
||
|
||
```json
|
||
{"id": "group", "type": "user", "name": "${machine:account}", "groups": ["docker"]}
|
||
```
|
||
|
||
`zsh` already declares a `user` resource for the same account (its login shell). The controller
|
||
compares `name` across modules, so the two collide.
|
||
|
||
**The change proposed (mesh-controller, `checkResources`):** judge a `user` resource by the fields it
|
||
sets, not by its name:
|
||
|
||
- `shell` and `home` stay single-owner;
|
||
- `groups` may be declared by any number of modules, because the host only adds them.
|
||
|
||
Then this module declares the resource above, and no module has to carry another's group.
|
||
|
||
Today the operator account is in the group on every machine, by hand. Nothing is lost while it waits.
|
||
|
||
## The bootstrap's runtime
|
||
|
||
On the machine the mesh was first installed on, the foundation bundle declared `package docker`
|
||
(`container-runtime`) and `docker.service` running and enabled (`container-runtime-running`). ADR 0207
|
||
§5 exempts them.
|
||
|
||
- The host records them under their bare ids, with origin *carried*. A mesh declaration's orphan pass
|
||
never sees them (mesh-host `store.go`).
|
||
- So `docker.package` here is a **second record of the same package**. The apply says "already
|
||
installed", and neither record ever uninstalls it.
|
||
- `docker.runtime` is likewise a second record of `docker.service`. Its found state is *running*,
|
||
because genesis started it, so undeclaring this module leaves the daemon running.
|
||
|
||
## Tools
|
||
|
||
The tools run as the operator account. If the daemon's socket refuses that account, a call is asked
|
||
again through `sudo -n` (a process keeps the groups it started with). Every call has a 20 s bound.
|
||
A failure is an error naming how it failed, never an empty answer.
|
||
|
||
**Every container on the machine is in scope.** A container the mesh holds carries the host's label
|
||
`mesh-host.id` (its value names the assignment), and every answer says `mesh_held`.
|
||
|
||
| tool | | what |
|
||
|---|---|---|
|
||
| `docker_list` | r | every container: image, state, health, restarts, ports, mounts, compose project, `mesh_held`; filter by owner, state or name |
|
||
| `docker_inspect` | r | one container whole, **environment values left out** (names kept), and its command line redacted as an exec's is |
|
||
| `docker_logs` | r | the last lines of both streams, merged in order, with timestamps (default 200, at most 2000); **a secret the container printed is shown as `[redacted: <name>]`** |
|
||
| `docker_secrets_in_logs` | r | which containers printed a secret they were given, **by name, never by value** (below) |
|
||
| `docker_secrets_in_events` | r | which secrets exec command lines carried, as the runtime recorded them in its events, **by name, never by value** (below) |
|
||
| `docker_stats` | r | CPU, memory, I/O and process count per running container, heaviest first |
|
||
| `docker_start` / `docker_stop` / `docker_restart` | a | one container. On a mesh-held one, the answer says the host restores its declared state at its next apply |
|
||
| `docker_top` | r | the processes inside one container |
|
||
| `docker_images` | r | images, largest first, with the containers using each; `dangling`, `unused` or `used` |
|
||
| `docker_prune` | a | dangling images and build cache, and stopped containers the mesh does not hold if `containers` is true. **A dry run unless `dry_run` is false. Never a volume** |
|
||
| `docker_prune_images` | a | named images no container and no declaration uses, keeping the previous version of each line (below). **A dry run unless `dry_run` is false, which needs `why`. Never forced** |
|
||
| `docker_disk_usage` | r | `docker system df -v`: total, active and reclaimable per kind, with the largest of each |
|
||
| `docker_networks` | r | networks, subnets, and the containers on each |
|
||
| `docker_volumes` | r | volumes, who mounts each, whether the mesh holds one of them, anonymous or not, and sizes if asked |
|
||
| `docker_events` | r | the runtime's events over a window ending now (default 60 min, at most 24 h), without exec noise unless asked; **an exec's command line is shown with any secret it carried as `[redacted: <what it was>]`** |
|
||
| `docker_daemon_config` | r | `daemon.json` as on disk, `docker info`'s essentials, and keys the daemon has not taken yet |
|
||
| `docker_unlabelled` | r | the containers the mesh does not hold: the cleanup list |
|
||
| `docker_problems` | r | unhealthy, restarting, dead, killed for memory, failed, or restarted five times or more |
|
||
| `docker_ports` | r | every published port, and the containers on the host's network |
|
||
|
||
## Secrets in a container's own log (hq issue 268)
|
||
|
||
Software prints what it is given: a server announcing its password as it starts, a startup script
|
||
echoing the database URI it connects with. The container's log is then a copy of the secret, held by
|
||
whoever reads it — this bundle's `docker_logs` among them. `docker_secrets_in_logs` reads the last
|
||
lines of each container's log (the mesh's by default, 5000 lines each, at most 50000) and compares
|
||
them with:
|
||
|
||
- the values of the container's environment whose names say they are secrets (`PASSWORD`, `SECRET`,
|
||
`TOKEN`, `API_KEY`, …; not a path, a URL, a number or a switch), as given and URL-encoded;
|
||
- the password inside any URI its environment holds;
|
||
- the shape `scheme://user:password@`, anywhere in a line, whatever the source — a password a program
|
||
already masked (`***`) is not one.
|
||
|
||
A finding names the container, the module, the assignment and the secret's variable, with how many
|
||
lines carry it and the first and last time. **It never carries the value or the line.** A secret
|
||
delivered only as a mounted file, never in the environment, is not known here — the bundle runs as the
|
||
operator account, which cannot read the host's 0600 files — and is caught only inside a URI.
|
||
|
||
`docker_logs` redacts the same values before it answers, because what it answers is read by agents
|
||
and kept in their transcripts. Its answer says how many it redacted, and points here.
|
||
|
||
A finding is a secret to rotate once the program stops printing it; recreating the container drops
|
||
its old log (the runtime's file goes with the container).
|
||
|
||
## Secrets on an exec's command line (hq issue 282)
|
||
|
||
The runtime records the command line of every exec — a `docker exec`, and a health check, which is one
|
||
— in its event stream (`exec_create: <argv joined by spaces>`). A program that passes a password to a
|
||
tool as an argument has given it to everyone who may ask the runtime what happened, for as long as the
|
||
runtime keeps its events, and to every transcript of a `docker_events` call. The mosquitto module did
|
||
that with the broker's admin password on every administrative call, until it handed it over on stdin.
|
||
|
||
`docker_events` redacts, in every exec's command line, before it answers:
|
||
|
||
- the values of that container's environment named like a secret, and the passwords in its URIs —
|
||
the same values `docker_logs` redacts;
|
||
- any URI carrying a password;
|
||
- by shape, whatever the source: the word after a flag that takes a password (`-P`, `--password`,
|
||
`--secret-key`, `--token`, …; `-a` for `redis-cli`; `-p` for `mosquitto_ctrl`, whose connect `-p`
|
||
is a port and is left alone as a number), a `NAME=value` whose name says secret, and the password a
|
||
`mosquitto_ctrl dynsec` command sets as an argument (`setClientPassword`, `init`).
|
||
|
||
`docker_secrets_in_events` reads the exec events of a window ending now (60 minutes by default, at most
|
||
24 hours) and names each secret found by container, module, what it was and the program, with how many
|
||
execs carried it and the first and last time — **never the value or the command line**. A finding is
|
||
code to change first (the secret handed over as a file or on stdin), then a secret to rotate. The
|
||
runtime keeps a bounded number of events, so on a busy machine a long window reads only what it still
|
||
holds — which is also why a leaked value ages out of the event stream quickly, and not out of a
|
||
transcript that already copied it.
|
||
|
||
`docker_inspect` shows a container's own command line (`Path`/`Args`, `Cmd`, `Entrypoint`) redacted
|
||
the same way.
|
||
|
||
## Removing the images nothing uses (hq ADR 0251 §5)
|
||
|
||
The weekly prune takes dangling images only, and an image pulled by digest is not dangling, so every
|
||
version of every module a machine ever ran stays on it. `docker_prune_images` takes those, and keeps:
|
||
|
||
- every image a container on the machine uses, in any state;
|
||
- every image the machine's declaration names — what the mesh would send it now and what it was last
|
||
sent — asked of the controller's `images` verb (the manifest's `invokes`). **No answer, an error, or
|
||
no machine name (`MESH_NODE`) means nothing is removed**: every image is answered as kept,
|
||
`declaration-unknown`;
|
||
- every image younger than `older_than_days` (seven by default, the weekly prune's week);
|
||
- in each **line** holding an image kept for one of the first two reasons, the newest other image: the
|
||
previous version, so going back needs no pull. A line is the images sharing a repository name,
|
||
joined across names when one image carries several (a build's local tag and the store's name).
|
||
|
||
A declared reference is matched against the runtime's own names for an image (`docker.io/` and
|
||
`library/` left out, a bare name tagged `latest`), by the whole reference and then by its digest.
|
||
|
||
Dangling images are not taken here; they stay the weekly prune's. A real run removes each image by
|
||
every name it carries, one image at a time and never with force, so the runtime refuses an image a
|
||
container uses; a refusal is reported for that image and the rest go on. The bytes it states are each
|
||
image's size summed, an upper bound, because images share layers.
|
||
|
||
## Tests
|
||
|
||
```
|
||
go test ./...
|
||
```
|
||
|
||
The tests run against a fake runner and cover:
|
||
|
||
- escalation through `sudo -n` on a refused socket, and never as root;
|
||
- each failure named by its cause;
|
||
- a name or id never read as an option;
|
||
- mesh-held marking;
|
||
- the environment left out of `inspect`;
|
||
- the restore note on a mesh-held act;
|
||
- prune being a dry run by default and never reaching a volume, a mesh container or `--volumes`;
|
||
- image pruning: each reason an image is kept, a two-name image being one line, the previous being the newest other image, nothing removed without the controller's answer or in a dry run, removal never forced, a real run refused without `why`;
|
||
- the log merge;
|
||
- a printed secret found by name and never answered by value, in the scan and in `docker_logs`;
|
||
- size parsing;
|
||
- what the daemon has not yet taken;
|
||
- event filtering;
|
||
- volume ownership;
|
||
- that the tools served are exactly the manifest's `tools`.
|