diff --git a/modules/docker/README.md b/modules/docker/README.md new file mode 100644 index 0000000..1838114 --- /dev/null +++ b/modules/docker/README.md @@ -0,0 +1,166 @@ +# docker + +The container runtime as a module (novox/hq to-be 42 phase 1, item 8; research 027/01–02; ADR 0166, +ADR 0207). It claims the node seat `node-container-runtime`. That seat carries no verbs yet: its verbs, +and the host creating containers through its holder, wait on ADR 0166's acceptance. Until then the +tools below are the module's own. + +## What it declares + +| resource | what | the host's rule | +|---|---|---| +| `package` | `docker` | installed if absent; never uninstalled when the module goes | +| `buildx` | `docker-buildx` | the same. Only the build machine has it today; `docker build` needs it for BuildKit everywhere | +| `socket` | `docker.socket` running, enabled at boot | given back as found when the module goes (ADR 0118) | +| `prune-service`, `prune-timer` | `/etc/systemd/system/docker-prune.{service,timer}`, written whole | removed with the module | +| `prune` | `docker-prune.timer` running, enabled at boot; restarted when either file changes | stopped and disabled with the module (the mesh made the unit) | + +The weekly prune takes **dangling images and build cache unused for a week, and nothing else**. It +takes no volume, no container and no image a container uses, so it never touches a container the mesh +holds. It runs at idle priority, at a random point in the hour after the weekly mark. A run missed +while the machine was off happens at the next boot. + +**Capabilities:** `package-manager`, `service-manager`, `privileged`. It does not declare +`container-runtime`: under ADR 0165, which is still proposed, that word means a running daemon, and +the module that installs the daemon cannot require it. + +## What it does not declare yet, and why + +Three things this module should own are already declared by other modules on every machine. The +controller refuses two modules on one node that declare the same `path`, `unit`, `name` or `package` +(`checkResources`, mesh-controller `internal/catalogue/resolve.go`). Declaring any of them here would +make the module unassignable everywhere. The refusals were checked against the controller's own +check: + +``` +zsh and docker both declare the name "${machine:account}" +dnsmasq and docker both declare the path "/etc/docker/daemon.json" +dnsmasq and docker both declare the unit "docker.service" +``` + +### 1. `/etc/docker/daemon.json` and `docker.service` (issue 190) + +Today the file has three writers. Each writes into it (`into: json`, ADR 0102) and reloads the +service: + +- **`dnsmasq`** writes `dns` and `live-restore`, through `dnsmasq.runtime-dns` and `dnsmasq.runtime`. +- **The private network**, generated by the controller (`internal/overlay/generator.go`), writes + `insecure-registries`. The collision check does not see generated resources. +- **Nobody** writes log rotation. One machine has `log-driver` and `log-opts` by hand. + +**The change proposed, in one merge:** + +1. `dnsmasq` drops its `runtime-dns` and `runtime` resources. +2. `docker` adds the two resources below: + + ```json + {"id": "daemon", "type": "file", "path": "/etc/docker/daemon.json", "mode": "0644", "into": "json", + "content": "{\"dns\": [\"${machine:address}\"], \"live-restore\": true, \"log-driver\": \"json-file\", \"log-opts\": {\"max-size\": \"100m\", \"max-file\": \"5\"}}\n"}, + {"id": "runtime", "type": "service", "unit": "docker.service", "state": "running", "boot": "enabled", "reload-on": ["daemon"]} + ``` + + The service is **reloaded, never restarted**: a restart stops every container. The daemon reads + `live-restore` on a reload. It reads `dns`, `log-driver` and `log-opts` only at its next start, so + they apply then (to containers created afterwards, for the log keys). With `live-restore` on, that + start keeps every container running. + +**Why one merge, and only after this module is on every machine:** + +- In one apply, the host first gives back the resources that are no longer declared, then applies + the new ones (mesh-host `apply.go`). +- `dnsmasq` gives back `dns` and `live-restore` to what they held before it, and `docker` sets them + again in the same apply. The daemon is reloaded once, after both steps. +- A machine pushed the new `dnsmasq` *without* this module would keep its pre-mesh values for both + keys. On one machine that is `live-restore: false`, and the next daemon restart there would stop + every container. + +**Later:** the controller hands the registry to this module as a value, and the overlay stops +generating its two resources (issue 190, steps 2 and 5). Until then the overlay keeps writing its one +key beside this module's. The host merges disjoint keys correctly; the mesh-host `into.go` record is +per resource. + +### 2. The operator account's membership of the `docker` group + +The right shape is the host's `user` shape. Its `groups` are additive: the host runs +`usermod --append` and never takes a group away. + +```json +{"id": "group", "type": "user", "name": "${machine:account}", "groups": ["docker"]} +``` + +`zsh` already declares a `user` resource for the same account (its login shell). The controller +compares `name` across modules, so the two collide. + +**The change proposed (mesh-controller, `checkResources`):** judge a `user` resource by the fields it +sets, not by its name: + +- `shell` and `home` stay single-owner; +- `groups` may be declared by any number of modules, because the host only adds them. + +Then this module declares the resource above, and no module has to carry another's group. + +Today the operator account is in the group on every machine, by hand. Nothing is lost while it waits. + +## The bootstrap's runtime + +On the machine the mesh was first installed on, the foundation bundle declared `package docker` +(`container-runtime`) and `docker.service` running and enabled (`container-runtime-running`). ADR 0207 +§5 exempts them. + +- The host records them under their bare ids, with origin *carried*. A mesh declaration's orphan pass + never sees them (mesh-host `store.go`). +- So `docker.package` here is a **second record of the same package**. The apply says "already + installed", and neither record ever uninstalls it. +- This module does not declare `docker.service` today, so nothing overlaps there. The proposed step + 1 would add a second record of that unit. Its found state is *running*, because genesis started + it, so undeclaring this module would leave the daemon running. + +## Tools + +The tools run as the operator account. If the daemon's socket refuses that account, a call is asked +again through `sudo -n` (a process keeps the groups it started with). Every call has a 20 s bound. +A failure is an error naming how it failed, never an empty answer. + +**Every container on the machine is in scope.** A container the mesh holds carries the host's label +`mesh-host.id` (its value names the assignment), and every answer says `mesh_held`. + +| tool | | what | +|---|---|---| +| `docker_list` | r | every container: image, state, health, restarts, ports, mounts, compose project, `mesh_held`; filter by owner, state or name | +| `docker_inspect` | r | one container whole, **environment values left out** (names kept) | +| `docker_logs` | r | the last lines of both streams, merged in order, with timestamps (default 200, at most 2000) | +| `docker_stats` | r | CPU, memory, I/O and process count per running container, heaviest first | +| `docker_start` / `docker_stop` / `docker_restart` | a | one container. On a mesh-held one, the answer says the host restores its declared state at its next apply | +| `docker_top` | r | the processes inside one container | +| `docker_images` | r | images, largest first, with the containers using each; `dangling`, `unused` or `used` | +| `docker_prune` | a | dangling images and build cache, and stopped containers the mesh does not hold if `containers` is true. **A dry run unless `dry_run` is false. Never a volume** | +| `docker_disk_usage` | r | `docker system df -v`: total, active and reclaimable per kind, with the largest of each | +| `docker_networks` | r | networks, subnets, and the containers on each | +| `docker_volumes` | r | volumes, who mounts each, whether the mesh holds one of them, anonymous or not, and sizes if asked | +| `docker_events` | r | the runtime's events over a window ending now (default 60 min, at most 24 h), without exec noise | +| `docker_daemon_config` | r | `daemon.json` as on disk, `docker info`'s essentials, and keys the daemon has not taken yet | +| `docker_unlabelled` | r | the containers the mesh does not hold: the cleanup list | +| `docker_problems` | r | unhealthy, restarting, dead, killed for memory, failed, or restarted five times or more | +| `docker_ports` | r | every published port, and the containers on the host's network | + +## Tests + +``` +go test ./... +``` + +The tests run against a fake runner and cover: + +- escalation through `sudo -n` on a refused socket, and never as root; +- each failure named by its cause; +- a name or id never read as an option; +- mesh-held marking; +- the environment left out of `inspect`; +- the restore note on a mesh-held act; +- prune being a dry run by default and never reaching a volume, a mesh container or `--volumes`; +- the log merge; +- size parsing; +- what the daemon has not yet taken; +- event filtering; +- volume ownership; +- that the tools served are exactly the manifest's `tools`. diff --git a/modules/docker/cmd/docker-tools/docker.go b/modules/docker/cmd/docker-tools/docker.go new file mode 100644 index 0000000..9f08014 --- /dev/null +++ b/modules/docker/cmd/docker-tools/docker.go @@ -0,0 +1,1064 @@ +package main + +// The container runtime's own command line, asked by the operator account the node's tool runtime +// runs as (novox/hq ADR 0175 §4). That account is in the docker group on every machine, but a +// process keeps the groups it started with: a runtime started before the account joined the group +// is refused by the daemon's socket. So a refused socket is asked again through `sudo -n`, as the +// service manager's and the packet filter's acts are, and a refusal there is named by how it failed. +// +// A container the mesh holds carries the host's label (MeshLabel); its value names the assignment +// (`.`). Every answer says which containers are the mesh's, because the host +// restores a held container to what its declaration says at its next apply (novox/hq ADR 0166). + +import ( + "context" + "encoding/json" + "fmt" + "os" + "regexp" + "sort" + "strconv" + "strings" + "time" +) + +// MeshLabel is the label the host puts on every container it creates (mesh-host internal/apply). +const MeshLabel = "mesh-host.id" + +// DaemonFile is the runtime's configuration file. +const DaemonFile = "/etc/docker/daemon.json" + +// Client asks the runtime. +type Client struct { + Run Runner + UID int + ReadFile func(string) ([]byte, error) + Now func() time.Time +} + +// NewClient is the client the bundle serves with. +func NewClient() *Client { + return &Client{Run: ExecRunner, UID: os.Getuid(), ReadFile: os.ReadFile, Now: time.Now} +} + +var socketRefused = regexp.MustCompile(`(?i)permission denied.*docker.*sock|docker\.sock.*permission denied`) + +// docker runs one docker command and answers its stdout, or an error naming what went wrong. +func (c *Client) docker(ctx context.Context, args ...string) (string, error) { + r := c.Run(ctx, "docker", args...) + program := "docker" + if r.Status != 0 && r.Err == "" && c.UID != 0 && socketRefused.MatchString(r.Stderr) { + program = "sudo" + r = c.Run(ctx, "sudo", append([]string{"-n", "docker"}, args...)...) + } + if r.Status == 0 && r.Err == "" { + return r.Stdout, nil + } + return "", failure(args, program, r) +} + +func failure(args []string, program string, r Ran) error { + verb := "docker" + if len(args) > 0 { + verb += " " + args[0] + } + said := strings.TrimSpace(r.Stderr + "\n" + r.Stdout) + switch { + case r.Err == "ENOENT" && program == "sudo": + return fmt.Errorf("the runtime's socket refused this account, and sudo is not installed here to escalate with") + case r.Err == "ENOENT": + return fmt.Errorf("docker is not installed on this machine") + case r.Err != "": + return fmt.Errorf("%s did not answer: %s (the daemon may still be doing it)", verb, r.Err) + case program == "sudo" && regexp.MustCompile(`(?m)^sudo:`).MatchString(said): + return fmt.Errorf("the runtime's socket refused this account (not in the docker group, or not since the tool runtime started) and it may not escalate without a prompt: %s", firstLine(said)) + case strings.Contains(said, "Cannot connect to the Docker daemon"): + return fmt.Errorf("the docker daemon is not answering on this machine: %s", firstLine(said)) + case strings.Contains(said, "No such container") || strings.Contains(said, "No such object"): + return fmt.Errorf("%s", firstLine(said)) + } + if l := firstLine(said); l != "" { + return fmt.Errorf("%s failed (%d): %s", verb, r.Status, l) + } + return fmt.Errorf("%s failed with status %d", verb, r.Status) +} + +var refPattern = regexp.MustCompile(`^[A-Za-z0-9][A-Za-z0-9_.:/@+-]*$`) + +// Ref is a container's, image's, network's or volume's name or id as an argument: never something +// docker would read as an option, which under sudo would be root's option. +func Ref(s string) (string, error) { + s = strings.TrimSpace(s) + if !refPattern.MatchString(s) || len(s) > 256 { + return "", fmt.Errorf("%q is not a name or id docker knows things by", s) + } + return s, nil +} + +// lines splits output into its non-empty lines. +func lines(s string) []string { + var out []string + for _, l := range strings.Split(s, "\n") { + if l = strings.TrimSpace(l); l != "" { + out = append(out, l) + } + } + return out +} + +// jsonLines decodes one JSON object per line, as `--format '{{json .}}'` prints them. +func jsonLines[T any](s string) ([]T, error) { + out := []T{} + for _, l := range lines(s) { + var v T + if err := json.Unmarshal([]byte(l), &v); err != nil { + return nil, fmt.Errorf("docker answered a line that is not JSON: %.120s", l) + } + out = append(out, v) + } + return out, nil +} + +// Bytes reads docker's human sizes ("55.63GB", "33.2MiB", "0B"): decimal units as docker's own +// HumanSize prints them, binary ones where it says so. -1 when it is not a size. +func Bytes(s string) int64 { + s = strings.TrimSpace(s) + if i := strings.Index(s, " ("); i > 0 { + s = s[:i] + } + m := regexp.MustCompile(`^([0-9.]+)\s*([A-Za-z]*)$`).FindStringSubmatch(s) + if m == nil { + return -1 + } + n, err := strconv.ParseFloat(m[1], 64) + if err != nil { + return -1 + } + units := map[string]float64{"": 1, "B": 1, "kB": 1e3, "KB": 1e3, "MB": 1e6, "GB": 1e9, "TB": 1e12, "PB": 1e15, + "KiB": 1 << 10, "MiB": 1 << 20, "GiB": 1 << 30, "TiB": 1 << 40} + u, ok := units[m[2]] + if !ok { + return -1 + } + return int64(n * u) +} + +// ---- containers ------------------------------------------------------------------------------- + +type inspected struct { + ID string `json:"Id"` + Name string + Created string + Image string + Config struct { + Image string + Labels map[string]string + } + State struct { + Status string + Running bool + Restarting bool + OOMKilled bool + ExitCode int + StartedAt string + FinishedAt string + Health *struct{ Status string } + } + RestartCount int + HostConfig struct { + RestartPolicy struct{ Name string } + NetworkMode string + Privileged bool + } + NetworkSettings struct { + Ports map[string][]struct { + HostIP string `json:"HostIp"` + HostPort string + } + } + Mounts []struct { + Type string + Name string + Source string + Destination string + RW bool + } +} + +// Mount is one thing a container mounts. +type Mount struct { + Type string `json:"type"` + Name string `json:"name,omitempty"` + Source string `json:"source,omitempty"` + Destination string `json:"destination"` + ReadOnly bool `json:"read_only,omitempty"` +} + +// Container is one container as every tool here answers it. +type Container struct { + ID string `json:"id"` + Name string `json:"name"` + Image string `json:"image"` + ImageID string `json:"image_id"` + State string `json:"state"` + Health string `json:"health,omitempty"` + ExitCode int `json:"exit_code"` + OOMKilled bool `json:"oom_killed,omitempty"` + Created string `json:"created"` + StartedAt string `json:"started_at,omitempty"` + FinishedAt string `json:"finished_at,omitempty"` + Restarts int `json:"restarts"` + RestartPolicy string `json:"restart_policy,omitempty"` + Network string `json:"network_mode,omitempty"` + Privileged bool `json:"privileged,omitempty"` + MeshHeld bool `json:"mesh_held"` + HeldBy string `json:"held_by,omitempty"` + Module string `json:"module,omitempty"` + Compose string `json:"compose_project,omitempty"` + ComposeDir string `json:"compose_dir,omitempty"` + Ports []string `json:"ports"` + Mounts []Mount `json:"mounts"` +} + +func summary(i inspected) Container { + c := Container{ + ID: shortID(i.ID), Name: strings.TrimPrefix(i.Name, "/"), Image: i.Config.Image, ImageID: i.Image, + State: i.State.Status, ExitCode: i.State.ExitCode, OOMKilled: i.State.OOMKilled, Created: i.Created, + Restarts: i.RestartCount, RestartPolicy: i.HostConfig.RestartPolicy.Name, Network: i.HostConfig.NetworkMode, + Privileged: i.HostConfig.Privileged, Ports: []string{}, Mounts: []Mount{}, + } + if !zeroTime(i.State.StartedAt) { + c.StartedAt = i.State.StartedAt + } + if !zeroTime(i.State.FinishedAt) && !i.State.Running { + c.FinishedAt = i.State.FinishedAt + } + if i.State.Health != nil { + c.Health = i.State.Health.Status + } + if v, ok := i.Config.Labels[MeshLabel]; ok { + c.MeshHeld, c.HeldBy = true, v + c.Module, _, _ = strings.Cut(v, ".") + } + c.Compose = i.Config.Labels["com.docker.compose.project"] + c.ComposeDir = i.Config.Labels["com.docker.compose.project.working_dir"] + for port, binds := range i.NetworkSettings.Ports { + for _, b := range binds { + c.Ports = append(c.Ports, fmt.Sprintf("%s:%s->%s", b.HostIP, b.HostPort, port)) + } + } + sort.Strings(c.Ports) + for _, m := range i.Mounts { + mm := Mount{Type: m.Type, Destination: m.Destination, ReadOnly: !m.RW} + if m.Type == "volume" { + mm.Name = m.Name + } else { + mm.Source = m.Source + } + c.Mounts = append(c.Mounts, mm) + } + return c +} + +func zeroTime(s string) bool { return s == "" || strings.HasPrefix(s, "0001-01-01") } + +func shortID(id string) string { + id = strings.TrimPrefix(id, "sha256:") + if len(id) > 12 { + return id[:12] + } + return id +} + +// inspectAll is every container on the machine, the mesh's and every other. +func (c *Client) inspectAll(ctx context.Context) ([]inspected, error) { + ids, err := c.docker(ctx, "ps", "--all", "--quiet", "--no-trunc") + if err != nil { + return nil, err + } + all := lines(ids) + if len(all) == 0 { + return []inspected{}, nil + } + out, err := c.docker(ctx, append([]string{"container", "inspect"}, all...)...) + if err != nil { + // A container removed between the two calls fails the whole inspect; ask once more. + if strings.Contains(err.Error(), "No such") { + return c.inspectAll(ctx) + } + return nil, err + } + var got []inspected + if err := json.Unmarshal([]byte(out), &got); err != nil { + return nil, fmt.Errorf("docker inspect answered something that is not JSON: %v", err) + } + return got, nil +} + +// Containers is every container, narrowed as asked. +func (c *Client) Containers(ctx context.Context, held, state, match string) ([]Container, error) { + all, err := c.inspectAll(ctx) + if err != nil { + return nil, err + } + out := []Container{} + for _, i := range all { + s := summary(i) + switch held { + case "", "all": + case "mesh": + if !s.MeshHeld { + continue + } + case "other": + if s.MeshHeld { + continue + } + default: + return nil, fmt.Errorf("held %q: \"all\", \"mesh\" or \"other\"", held) + } + if state != "" && s.State != state { + continue + } + if match != "" && !strings.Contains(s.Name, match) && !strings.Contains(s.Image, match) { + continue + } + out = append(out, s) + } + sort.Slice(out, func(a, b int) bool { return out[a].Name < out[b].Name }) + return out, nil +} + +// Inspect is one container whole, as docker inspects it, with every environment variable's value +// left out (names kept): a container's environment is where its passwords are. +func (c *Client) Inspect(ctx context.Context, ref string) (map[string]any, error) { + ref, err := Ref(ref) + if err != nil { + return nil, err + } + out, err := c.docker(ctx, "container", "inspect", ref) + if err != nil { + return nil, err + } + var got []map[string]any + if err := json.Unmarshal([]byte(out), &got); err != nil || len(got) != 1 { + return nil, fmt.Errorf("docker inspect answered something that is not one container") + } + obj := got[0] + if cfg, ok := obj["Config"].(map[string]any); ok { + if env, ok := cfg["Env"].([]any); ok { + names := []string{} + for _, e := range env { + name, _, _ := strings.Cut(fmt.Sprint(e), "=") + names = append(names, name) + } + cfg["Env"] = names + cfg["EnvValues"] = "left out: a container's environment holds its secrets" + } + if labels, ok := cfg["Labels"].(map[string]any); ok { + if v, ok := labels[MeshLabel]; ok { + obj["mesh_held"], obj["held_by"] = true, v + } else { + obj["mesh_held"] = false + } + } + } + return obj, nil +} + +// Logs is the last lines a container wrote, both streams in the order they were written. +func (c *Client) Logs(ctx context.Context, ref string, tail int, since string) (map[string]any, error) { + ref, err := Ref(ref) + if err != nil { + return nil, err + } + args := []string{"logs", "--timestamps", "--tail", strconv.Itoa(tail)} + if since != "" { + if !regexp.MustCompile(`^[0-9]+[smhd]?$|^\d{4}-\d{2}-\d{2}`).MatchString(since) { + return nil, fmt.Errorf("since %q: a duration such as 30m or 2h, or a time such as 2026-10-04T10:00:00", since) + } + args = append(args, "--since", since) + } + args = append(args, ref) + r := c.Run(ctx, "docker", args...) + program := "docker" + if r.Status != 0 && r.Err == "" && c.UID != 0 && socketRefused.MatchString(r.Stderr) { + program = "sudo" + r = c.Run(ctx, "sudo", append([]string{"-n", "docker"}, args...)...) + } + if r.Status != 0 || r.Err != "" { + return nil, failure(args, program, r) + } + // Both streams carry the container's lines, each led by its timestamp, so they merge in order. + all := append(lines(r.Stdout), lines(r.Stderr)...) + sort.SliceStable(all, func(a, b int) bool { return all[a] < all[b] }) + if len(all) > tail { + all = all[len(all)-tail:] + } + const most = 4096 + for i, l := range all { + if len(l) > most { + all[i] = l[:most] + "…" + } + } + return map[string]any{"container": ref, "lines": all, "count": len(all)}, nil +} + +// Stat is one container's use of the machine now. +type Stat struct { + Name string `json:"name"` + ID string `json:"id"` + CPU string `json:"cpu"` + Memory string `json:"memory"` + MemPerc string `json:"memory_percent"` + MemBytes int64 `json:"memory_bytes"` + NetIO string `json:"net_io"` + BlockIO string `json:"block_io"` + PIDs string `json:"pids"` +} + +// Stats is a snapshot of every running container's use, or one's; the heaviest first. +func (c *Client) Stats(ctx context.Context, ref string) ([]Stat, error) { + args := []string{"stats", "--no-stream", "--format", "{{json .}}"} + if ref != "" { + r, err := Ref(ref) + if err != nil { + return nil, err + } + args = append(args, r) + } + out, err := c.docker(ctx, args...) + if err != nil { + return nil, err + } + raw, err := jsonLines[map[string]string](out) + if err != nil { + return nil, err + } + stats := []Stat{} + for _, s := range raw { + used, _, _ := strings.Cut(s["MemUsage"], " / ") + stats = append(stats, Stat{Name: s["Name"], ID: s["ID"], CPU: s["CPUPerc"], Memory: s["MemUsage"], MemPerc: s["MemPerc"], + MemBytes: Bytes(used), NetIO: s["NetIO"], BlockIO: s["BlockIO"], PIDs: s["PIDs"]}) + } + sort.Slice(stats, func(a, b int) bool { return stats[a].MemBytes > stats[b].MemBytes }) + return stats, nil +} + +// Act starts, stops or restarts one container and answers its state after. A container the mesh +// holds is acted on too — the host restores what its declaration says at its next apply, and the +// answer says so (novox/hq ADR 0166). +func (c *Client) Act(ctx context.Context, verb, ref string) (map[string]any, error) { + ref, err := Ref(ref) + if err != nil { + return nil, err + } + args := []string{verb} + switch verb { + case "start": + case "stop", "restart": + // Ten seconds to stop before it is killed: inside the call's own bound. + args = append(args, "--time", "10") + default: + return nil, fmt.Errorf("%q is not start, stop or restart", verb) + } + if _, err := c.docker(ctx, append(args, ref)...); err != nil { + return nil, err + } + out, err := c.docker(ctx, "container", "inspect", ref) + if err != nil { + return nil, err + } + var got []inspected + if err := json.Unmarshal([]byte(out), &got); err != nil || len(got) != 1 { + return nil, fmt.Errorf("docker inspect answered something that is not one container") + } + s := summary(got[0]) + answer := map[string]any{"container": s.Name, "verb": verb, "ok": true, "state": s.State, "mesh_held": s.MeshHeld} + if s.MeshHeld { + answer["held_by"] = s.HeldBy + answer["note"] = fmt.Sprintf("the mesh holds this container (%s): the host restores what its declaration says at its next apply", s.HeldBy) + } + return answer, nil +} + +// Top is the processes running in one container. +func (c *Client) Top(ctx context.Context, ref string) (map[string]any, error) { + ref, err := Ref(ref) + if err != nil { + return nil, err + } + out, err := c.docker(ctx, "top", ref, "-eo", "pid,user,etime,pcpu,rss,args") + if err != nil { + return nil, err + } + rows := lines(out) + procs := []map[string]string{} + for _, l := range rows[min(1, len(rows)):] { + f := strings.Fields(l) + if len(f) < 6 { + continue + } + procs = append(procs, map[string]string{"pid": f[0], "user": f[1], "elapsed": f[2], "cpu": f[3], "rss_kb": f[4], "command": strings.Join(f[5:], " ")}) + } + return map[string]any{"container": ref, "processes": procs}, nil +} + +// Problems is every container that is not well: unhealthy, restarting, dead, killed for memory, +// or exited with a failure. +func (c *Client) Problems(ctx context.Context) ([]map[string]any, error) { + all, err := c.Containers(ctx, "", "", "") + if err != nil { + return nil, err + } + out := []map[string]any{} + for _, s := range all { + var why []string + if s.Health == "unhealthy" { + why = append(why, "unhealthy") + } + if s.State == "restarting" { + why = append(why, "restarting") + } + if s.State == "dead" { + why = append(why, "dead") + } + if s.OOMKilled { + why = append(why, "killed for memory") + } + if s.State == "exited" && s.ExitCode != 0 { + why = append(why, fmt.Sprintf("exited %d", s.ExitCode)) + } + if s.Restarts >= 5 { + why = append(why, fmt.Sprintf("restarted %d times", s.Restarts)) + } + if len(why) == 0 { + continue + } + out = append(out, map[string]any{"name": s.Name, "state": s.State, "health": s.Health, "why": why, + "mesh_held": s.MeshHeld, "held_by": s.HeldBy, "image": s.Image, "finished_at": s.FinishedAt}) + } + return out, nil +} + +// Ports is every port the containers publish on the machine. +func (c *Client) Ports(ctx context.Context) ([]map[string]any, error) { + all, err := c.Containers(ctx, "", "", "") + if err != nil { + return nil, err + } + out := []map[string]any{} + for _, s := range all { + for _, p := range s.Ports { + out = append(out, map[string]any{"published": p, "container": s.Name, "state": s.State, "mesh_held": s.MeshHeld}) + } + if s.Network == "host" && s.State == "running" { + out = append(out, map[string]any{"published": "host network: every port it listens on", "container": s.Name, "state": s.State, "mesh_held": s.MeshHeld}) + } + } + return out, nil +} + +// Unlabelled is every container the mesh does not hold: the cleanup list, with what each is +// likely to be (a compose project, a one-off). +func (c *Client) Unlabelled(ctx context.Context) (map[string]any, error) { + all, err := c.Containers(ctx, "other", "", "") + if err != nil { + return nil, err + } + running := 0 + for _, s := range all { + if s.State == "running" { + running++ + } + } + return map[string]any{"count": len(all), "running": running, "containers": all, + "note": "the mesh holds none of these; it never removes them. docker_prune with containers=true removes the stopped ones"}, nil +} + +// ---- images ---------------------------------------------------------------------------------- + +// Image is one image with what uses it. +type Image struct { + ID string `json:"id"` + Repository string `json:"repository"` + Tag string `json:"tag"` + Created string `json:"created"` + Size string `json:"size"` + SizeBytes int64 `json:"size_bytes"` + Dangling bool `json:"dangling"` + UsedBy []string `json:"used_by"` + MeshUsed bool `json:"used_by_mesh"` +} + +// Images is the images on the machine, the largest first, with the containers using each. +func (c *Client) Images(ctx context.Context, filter, match string, limit int) (map[string]any, error) { + out, err := c.docker(ctx, "image", "ls", "--no-trunc", "--format", "{{json .}}") + if err != nil { + return nil, err + } + raw, err := jsonLines[map[string]string](out) + if err != nil { + return nil, err + } + containers, err := c.inspectAll(ctx) + if err != nil { + return nil, err + } + users, meshUsers := map[string][]string{}, map[string]bool{} + for _, i := range containers { + s := summary(i) + users[i.Image] = append(users[i.Image], s.Name) + meshUsers[i.Image] = meshUsers[i.Image] || s.MeshHeld + } + images := []Image{} + var total int64 + for _, r := range raw { + img := Image{ID: r["ID"], Repository: r["Repository"], Tag: r["Tag"], Created: r["CreatedAt"], Size: r["Size"], + SizeBytes: Bytes(r["Size"]), Dangling: r["Repository"] == "" && r["Tag"] == ""} + img.UsedBy = append([]string{}, users[img.ID]...) + img.MeshUsed = meshUsers[img.ID] + switch filter { + case "", "all": + case "dangling": + if !img.Dangling { + continue + } + case "unused": + if len(img.UsedBy) > 0 { + continue + } + case "used": + if len(img.UsedBy) == 0 { + continue + } + default: + return nil, fmt.Errorf("filter %q: all, dangling, unused or used", filter) + } + if match != "" && !strings.Contains(img.Repository+":"+img.Tag, match) { + continue + } + total += max(img.SizeBytes, 0) + images = append(images, img) + } + sort.Slice(images, func(a, b int) bool { return images[a].SizeBytes > images[b].SizeBytes }) + count := len(images) + if len(images) > limit { + images = images[:limit] + } + return map[string]any{"count": count, "shown": len(images), "size_bytes_summed": total, + "note": "sizes are summed per image; images share layers, so the disk they take together is docker_disk_usage's", "images": images}, nil +} + +// ---- prune ----------------------------------------------------------------------------------- + +// PruneAsk is what a prune is asked to take. Volumes are never among them: a volume is data. +type PruneAsk struct { + Images bool + BuildCache bool + Containers bool + OlderThanH int + DryRun bool +} + +// Prune removes, or with DryRun only lists, dangling images, unused build cache and — only when +// asked — stopped containers the mesh does not hold. Never a volume, never a container the mesh +// holds, never an image a container uses (the daemon refuses that itself). +func (c *Client) Prune(ctx context.Context, a PruneAsk) (map[string]any, error) { + answer := map[string]any{"dry_run": a.DryRun, "volumes": "never pruned: a volume is data"} + until := []string{} + if a.OlderThanH > 0 { + until = []string{"--filter", fmt.Sprintf("until=%dh", a.OlderThanH)} + answer["older_than_hours"] = a.OlderThanH + } + if a.Images { + args := append([]string{"image", "ls", "--no-trunc", "--filter", "dangling=true", "--format", "{{json .}}"}, until...) + out, err := c.docker(ctx, args...) + if err != nil { + return nil, err + } + raw, err := jsonLines[map[string]string](out) + if err != nil { + return nil, err + } + var sum int64 + ids := []string{} + for _, r := range raw { + sum += max(Bytes(r["Size"]), 0) + ids = append(ids, shortID(r["ID"])) + } + if len(ids) > 50 { + ids = ids[:50] + } + section := map[string]any{"dangling": len(raw), "size_bytes_summed": sum, "ids_first_50": ids} + if !a.DryRun && len(raw) > 0 { + out, err := c.docker(ctx, append([]string{"image", "prune", "--force"}, until...)...) + if err != nil { + return nil, err + } + section["reclaimed"] = reclaimed(out) + } + answer["images"] = section + } + if a.BuildCache { + section := map[string]any{} + if a.DryRun { + out, err := c.docker(ctx, "system", "df", "--format", "{{json .}}") + if err != nil { + return nil, err + } + rows, err := jsonLines[map[string]string](out) + if err != nil { + return nil, err + } + for _, r := range rows { + if r["Type"] == "Build Cache" { + section["entries"], section["size"], section["reclaimable"] = r["TotalCount"], r["Size"], r["Reclaimable"] + } + } + section["note"] = "an estimate: a real prune takes the cache nothing refers to, which can be less than reclaimable" + } else { + out, err := c.docker(ctx, append([]string{"builder", "prune", "--force"}, until...)...) + if err != nil { + return nil, err + } + section["reclaimed"] = reclaimed(out) + } + answer["build_cache"] = section + } + if a.Containers { + all, err := c.Containers(ctx, "other", "", "") + if err != nil { + return nil, err + } + cutoff := c.Now().Add(-time.Duration(a.OlderThanH) * time.Hour) + stopped := []string{} + for _, s := range all { + if s.State != "exited" && s.State != "created" && s.State != "dead" { + continue + } + if a.OlderThanH > 0 { + when := s.FinishedAt + if when == "" { + when = s.Created + } + if t, err := time.Parse(time.RFC3339Nano, when); err == nil && t.After(cutoff) { + continue + } + } + stopped = append(stopped, s.Name) + } + section := map[string]any{"stopped_not_held": stopped} + if !a.DryRun && len(stopped) > 0 { + // By name, and without --volumes: what they mounted stays. + if _, err := c.docker(ctx, append([]string{"container", "rm"}, stopped...)...); err != nil { + return nil, err + } + section["removed"] = stopped + } + answer["containers"] = section + } + if a.DryRun { + answer["note"] = "nothing was removed: call again with dry_run false to prune" + } + return answer, nil +} + +func reclaimed(out string) string { + for _, l := range lines(out) { + if strings.HasPrefix(l, "Total reclaimed space:") || strings.HasPrefix(l, "Total:") { + return strings.TrimSpace(l[strings.Index(l, ":")+1:]) + } + } + return "0B" +} + +// ---- disk, networks, volumes, events, daemon --------------------------------------------------- + +// DiskUsage is `docker system df -v`, summed per kind, with the largest of each. +func (c *Client) DiskUsage(ctx context.Context, top int) (map[string]any, error) { + out, err := c.docker(ctx, "system", "df", "--format", "{{json .}}") + if err != nil { + return nil, err + } + rows, err := jsonLines[map[string]any](out) + if err != nil { + return nil, err + } + summaryRows := []map[string]any{} + for _, r := range rows { + summaryRows = append(summaryRows, map[string]any{"type": r["Type"], "total": r["TotalCount"], "active": r["Active"], + "size": r["Size"], "size_bytes": Bytes(fmt.Sprint(r["Size"])), "reclaimable": r["Reclaimable"]}) + } + verbose, err := c.docker(ctx, "system", "df", "--verbose", "--format", "{{json .}}") + if err != nil { + return nil, err + } + var v map[string][]map[string]any + if err := json.Unmarshal([]byte(verbose), &v); err != nil { + return nil, fmt.Errorf("docker system df -v answered something that is not JSON: %v", err) + } + largest := map[string]any{} + for kind, items := range v { + sort.Slice(items, func(a, b int) bool { return Bytes(fmt.Sprint(items[a]["Size"])) > Bytes(fmt.Sprint(items[b]["Size"])) }) + if len(items) > top { + items = items[:top] + } + trimmed := []map[string]any{} + for _, it := range items { + t := map[string]any{"size": it["Size"]} + for _, k := range []string{"Repository", "Tag", "ID", "Names", "Name", "Containers", "Links", "Image", "State", "Description", "InUse", "LastUsedSince", "UniqueSize"} { + if val, ok := it[k]; ok && val != nil && val != "" { + t[strings.ToLower(k)] = val + } + } + trimmed = append(trimmed, t) + } + largest[kind] = trimmed + } + return map[string]any{"summary": summaryRows, "largest": largest}, nil +} + +// Networks is every network with its addressing and the running containers on it. +func (c *Client) Networks(ctx context.Context) ([]map[string]any, error) { + ids, err := c.docker(ctx, "network", "ls", "--quiet", "--no-trunc") + if err != nil { + return nil, err + } + if len(lines(ids)) == 0 { + return []map[string]any{}, nil + } + out, err := c.docker(ctx, append([]string{"network", "inspect"}, lines(ids)...)...) + if err != nil { + return nil, err + } + var got []struct { + Name, Driver, Scope string + Internal bool + IPAM struct { + Config []struct{ Subnet, Gateway string } + } + Containers map[string]struct{ Name, IPv4Address string } + Labels map[string]string + } + if err := json.Unmarshal([]byte(out), &got); err != nil { + return nil, fmt.Errorf("docker network inspect answered something that is not JSON: %v", err) + } + held := map[string]bool{} + if all, err := c.Containers(ctx, "", "", ""); err == nil { + for _, s := range all { + held[s.Name] = s.MeshHeld + } + } + res := []map[string]any{} + for _, n := range got { + subnets := []string{} + for _, cfg := range n.IPAM.Config { + subnets = append(subnets, strings.TrimSpace(cfg.Subnet+" gw "+cfg.Gateway)) + } + members := []map[string]any{} + for _, m := range n.Containers { + members = append(members, map[string]any{"name": m.Name, "address": m.IPv4Address, "mesh_held": held[m.Name]}) + } + sort.Slice(members, func(a, b int) bool { return fmt.Sprint(members[a]["name"]) < fmt.Sprint(members[b]["name"]) }) + res = append(res, map[string]any{"name": n.Name, "driver": n.Driver, "scope": n.Scope, "internal": n.Internal, + "subnets": subnets, "containers": members, "compose_project": n.Labels["com.docker.compose.project"]}) + } + sort.Slice(res, func(a, b int) bool { return fmt.Sprint(res[a]["name"]) < fmt.Sprint(res[b]["name"]) }) + return res, nil +} + +// Volumes is every volume with the containers mounting it, whether the mesh holds any of them, and +// — when asked, which is slower — its size. +func (c *Client) Volumes(ctx context.Context, unmountedOnly, sizes bool) (map[string]any, error) { + names, err := c.docker(ctx, "volume", "ls", "--quiet") + if err != nil { + return nil, err + } + containers, err := c.Containers(ctx, "", "", "") + if err != nil { + return nil, err + } + mountedBy := map[string][]map[string]any{} + for _, s := range containers { + for _, m := range s.Mounts { + if m.Type == "volume" { + mountedBy[m.Name] = append(mountedBy[m.Name], map[string]any{"container": s.Name, "state": s.State, "mesh_held": s.MeshHeld}) + } + } + } + size := map[string]string{} + if sizes { + out, err := c.docker(ctx, "system", "df", "--verbose", "--format", "{{json .}}") + if err != nil { + return nil, err + } + var v struct{ Volumes []map[string]any } + if err := json.Unmarshal([]byte(out), &v); err == nil { + for _, vol := range v.Volumes { + size[fmt.Sprint(vol["Name"])] = fmt.Sprint(vol["Size"]) + } + } + } + vols := []map[string]any{} + if len(lines(names)) > 0 { + out, err := c.docker(ctx, append([]string{"volume", "inspect"}, lines(names)...)...) + if err != nil { + return nil, err + } + var got []struct { + Name, Driver, CreatedAt, Mountpoint string + Labels map[string]string + } + if err := json.Unmarshal([]byte(out), &got); err != nil { + return nil, fmt.Errorf("docker volume inspect answered something that is not JSON: %v", err) + } + for _, v := range got { + by := mountedBy[v.Name] + if unmountedOnly && len(by) > 0 { + continue + } + mesh := false + for _, b := range by { + mesh = mesh || b["mesh_held"].(bool) + } + _, anonymous := v.Labels["com.docker.volume.anonymous"] + entry := map[string]any{"name": v.Name, "driver": v.Driver, "created": v.CreatedAt, "anonymous": anonymous, + "compose_project": v.Labels["com.docker.compose.project"], "mounted_by": append([]map[string]any{}, by...), "mesh_held": mesh} + if sizes { + entry["size"] = size[v.Name] + } + vols = append(vols, entry) + } + } + return map[string]any{"count": len(vols), "volumes": vols, + "note": "mesh_held: a container the mesh holds mounts it. Nothing here removes a volume; docker_prune never does"}, nil +} + +// Events is what the runtime did in a window ending now, the latest last. +func (c *Client) Events(ctx context.Context, minutes int, kind string, limit int, execs bool) (map[string]any, error) { + args := []string{"events", "--since", fmt.Sprintf("%dm", minutes), "--until", "0s", "--format", "{{json .}}"} + if kind != "" { + switch kind { + case "container", "image", "network", "volume", "daemon", "plugin", "builder": + default: + return nil, fmt.Errorf("type %q: container, image, network, volume, daemon, plugin or builder", kind) + } + args = append(args, "--filter", "type="+kind) + } + out, err := c.docker(ctx, args...) + if err != nil { + return nil, err + } + raw, err := jsonLines[struct { + Type, Action string + Actor struct { + ID string + Attributes map[string]string + } + TimeNano int64 `json:"timeNano"` + }](out) + if err != nil { + return nil, err + } + events := []map[string]any{} + for _, e := range raw { + if !execs && strings.HasPrefix(e.Action, "exec_") { + continue + } + _, held := e.Actor.Attributes[MeshLabel] + ev := map[string]any{"time": time.Unix(0, e.TimeNano).UTC().Format(time.RFC3339), "type": e.Type, "action": e.Action, + "id": shortID(e.Actor.ID), "name": e.Actor.Attributes["name"]} + if e.Type == "container" { + ev["mesh_held"] = held + ev["image"] = e.Actor.Attributes["image"] + if code, ok := e.Actor.Attributes["exitCode"]; ok { + ev["exit_code"] = code + } + } + events = append(events, ev) + } + total := len(events) + if len(events) > limit { + events = events[len(events)-limit:] + } + return map[string]any{"minutes": minutes, "count": total, "shown": len(events), "events": events}, nil +} + +// restartOnly are the daemon keys the runtime reads only when it starts: a reload leaves them as +// they were (dockerd's reloadable set is debug, labels, live-restore, registries, max-concurrent-*, +// default-runtime, runtimes, shutdown-timeout and a few more). +var restartOnly = []string{"dns", "dns-opts", "dns-search", "data-root", "storage-driver", "log-driver", "log-opts", "bip", + "default-address-pools", "iptables", "ip6tables", "ipv6", "userns-remap", "exec-opts", "allow-direct-routing", "mtu"} + +// DaemonConfig is the runtime's file and what the daemon runs with now, and where the two differ. +func (c *Client) DaemonConfig(ctx context.Context) (map[string]any, error) { + answer := map[string]any{"file": DaemonFile} + var file map[string]any + raw, err := c.ReadFile(DaemonFile) + switch { + case os.IsNotExist(err): + answer["file_state"] = "absent: the daemon runs on its defaults" + case err != nil: + answer["file_state"] = "unreadable: " + err.Error() + default: + if err := json.Unmarshal(raw, &file); err != nil { + answer["file_state"] = "not JSON — the daemon refuses to start with it: " + err.Error() + } else { + answer["file_state"] = "read" + answer["keys"] = file + } + } + out, err := c.docker(ctx, "info", "--format", "{{json .}}") + if err != nil { + return nil, err + } + var info map[string]any + if err := json.Unmarshal([]byte(out), &info); err != nil { + return nil, fmt.Errorf("docker info answered something that is not JSON: %v", err) + } + essentials := map[string]any{} + for _, k := range []string{"ServerVersion", "Driver", "LoggingDriver", "CgroupDriver", "CgroupVersion", "LiveRestoreEnabled", + "DockerRootDir", "Containers", "ContainersRunning", "ContainersPaused", "ContainersStopped", "Images", "KernelVersion", + "OperatingSystem", "NCPU", "MemTotal", "DefaultRuntime", "FirewallBackend", "SecurityOptions", "Warnings", "Debug"} { + if v, ok := info[k]; ok { + essentials[k] = v + } + } + if reg, ok := info["RegistryConfig"].(map[string]any); ok { + insecure := []string{} + if idx, ok := reg["IndexConfigs"].(map[string]any); ok { + for name, cfg := range idx { + if m, ok := cfg.(map[string]any); ok && m["Secure"] == false { + insecure = append(insecure, name) + } + } + } + sort.Strings(insecure) + essentials["InsecureRegistries"] = insecure + essentials["InsecureRegistryCIDRs"] = reg["InsecureRegistryCIDRs"] + } + answer["daemon"] = essentials + + // Where the file says one thing and the daemon runs another: a key changed since the last reload, + // or one only a restart takes. + pending := []string{} + if file != nil { + if v, ok := file["live-restore"].(bool); ok && info["LiveRestoreEnabled"] != v { + pending = append(pending, fmt.Sprintf("live-restore is %v in the file and %v in the daemon: a reload takes it", v, info["LiveRestoreEnabled"])) + } + if v, ok := file["log-driver"].(string); ok && info["LoggingDriver"] != v { + pending = append(pending, fmt.Sprintf("log-driver is %s in the file and %v in the daemon: only a restart takes it", v, info["LoggingDriver"])) + } + present := []string{} + for _, k := range restartOnly { + if _, ok := file[k]; ok { + present = append(present, k) + } + } + answer["read_only_at_start"] = present + } + answer["pending"] = pending + answer["note"] = "keys read only at start take effect at the daemon's next restart; with live-restore on, a restart keeps every container running" + return answer, nil +} diff --git a/modules/docker/cmd/docker-tools/docker_test.go b/modules/docker/cmd/docker-tools/docker_test.go new file mode 100644 index 0000000..03d1b98 --- /dev/null +++ b/modules/docker/cmd/docker-tools/docker_test.go @@ -0,0 +1,360 @@ +package main + +import ( + "context" + "encoding/json" + "errors" + "io/fs" + "os" + "reflect" + "strings" + "testing" + "time" +) + +type call struct { + name string + args []string +} + +// fake answers each command by the first rule whose prefix matches "name arg arg…". +type fake struct { + rules []rule + calls []call +} + +type rule struct { + prefix string + ran Ran +} + +func (f *fake) on(prefix string, r Ran) *fake { f.rules = append(f.rules, rule{prefix, r}); return f } + +func (f *fake) run(_ context.Context, name string, args ...string) Ran { + f.calls = append(f.calls, call{name, args}) + line := strings.Join(append([]string{name}, args...), " ") + for _, r := range f.rules { + if strings.HasPrefix(line, r.prefix) { + return r.ran + } + } + return Ran{Status: 1, Stderr: "unexpected: " + line} +} + +func (f *fake) ran(prefix string) bool { + for _, c := range f.calls { + if strings.HasPrefix(strings.Join(append([]string{c.name}, c.args...), " "), prefix) { + return true + } + } + return false +} + +func client(f *fake, uid int) *Client { + return &Client{Run: f.run, UID: uid, ReadFile: func(string) ([]byte, error) { return nil, fs.ErrNotExist }, + Now: func() time.Time { return time.Date(2026, 10, 4, 12, 0, 0, 0, time.UTC) }} +} + +const held = `{"Id":"aaaaaaaaaaaaaaaa","Name":"/mesh-web","Created":"2026-10-01T00:00:00Z","Image":"sha256:img1", + "Config":{"Image":"web:1","Labels":{"mesh-host.id":"hello-web.server","mesh-host.spec":"x"},"Env":["PASSWORD=hunter2","PATH=/bin"]}, + "State":{"Status":"running","Running":true,"StartedAt":"2026-10-01T00:00:01Z","FinishedAt":"0001-01-01T00:00:00Z","Health":{"Status":"healthy"}}, + "HostConfig":{"RestartPolicy":{"Name":"unless-stopped"},"NetworkMode":"bridge"}, + "NetworkSettings":{"Ports":{"80/tcp":[{"HostIp":"0.0.0.0","HostPort":"8080"}]}}, + "Mounts":[{"Type":"volume","Name":"webdata","Destination":"/data","RW":true}]}` + +const stray = `{"Id":"bbbbbbbbbbbbbbbb","Name":"/dev-db","Created":"2026-09-01T00:00:00Z","Image":"sha256:img2", + "Config":{"Image":"postgres:16","Labels":{"com.docker.compose.project":"dev","com.docker.compose.project.working_dir":"/home/op/dev"}}, + "State":{"Status":"exited","ExitCode":1,"FinishedAt":"2026-09-02T00:00:00Z"}, + "HostConfig":{"RestartPolicy":{"Name":"no"}},"NetworkSettings":{"Ports":{}}, + "Mounts":[{"Type":"volume","Name":"dbdata","Destination":"/var/lib/postgresql/data","RW":true}]}` + +func machine() *fake { + return (&fake{}). + on("docker ps --all --quiet --no-trunc", Ran{Stdout: "aaaaaaaaaaaaaaaa\nbbbbbbbbbbbbbbbb\n"}). + on("docker container inspect aaaaaaaaaaaaaaaa bbbbbbbbbbbbbbbb", Ran{Stdout: "[" + held + "," + stray + "]"}). + on("docker container inspect mesh-web", Ran{Stdout: "[" + held + "]"}). + on("docker container inspect dev-db", Ran{Stdout: "[" + stray + "]"}) +} + +func TestARefusedSocketIsAskedAgainThroughSudoWithoutAPromptUnlessThisIsRoot(t *testing.T) { + denied := Ran{Status: 1, Stderr: "permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock: Get ...: dial unix /var/run/docker.sock: connect: permission denied\n"} + f := (&fake{}).on("docker ", denied).on("sudo -n docker info", Ran{Stdout: "{}"}) + if _, err := client(f, 1000).docker(context.Background(), "info", "--format", "{{json .}}"); err != nil { + t.Fatal(err) + } + if !f.ran("sudo -n docker info --format") { + t.Fatalf("not escalated: %+v", f.calls) + } + f = (&fake{}).on("docker ", denied) + if _, err := client(f, 0).docker(context.Background(), "info"); err == nil || f.ran("sudo") { + t.Fatalf("root escalated or answered: %v %+v", err, f.calls) + } +} + +func TestFailuresAreNamedByHowTheyFailed(t *testing.T) { + denied := Ran{Status: 1, Stderr: "permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock\n"} + cases := map[string]*fake{ + "may not escalate without a prompt": (&fake{}).on("docker ", denied).on("sudo ", Ran{Status: 1, Stderr: "sudo: a password is required\n"}), + "sudo is not installed": (&fake{}).on("docker ", denied).on("sudo ", Ran{Status: 127, Err: "ENOENT"}), + "docker is not installed": (&fake{}).on("docker ", Ran{Status: 127, Err: "ENOENT"}), + "daemon is not answering": (&fake{}).on("docker ", Ran{Status: 1, Stderr: "Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?\n"}), + "did not answer: no answer within": (&fake{}).on("docker ", Ran{Status: 124, Err: "no answer within 20 s"}), + "docker info failed (3): boom": (&fake{}).on("docker ", Ran{Status: 3, Stderr: "boom\n"}), + } + for want, f := range cases { + _, err := client(f, 1000).docker(context.Background(), "info") + if err == nil || !strings.Contains(err.Error(), want) { + t.Errorf("want %q, got %v", want, err) + } + } +} + +func TestANameIsNeverAnOption(t *testing.T) { + for _, bad := range []string{"--help", "-v", "", "a b", "x;y"} { + if _, err := Ref(bad); err == nil { + t.Errorf("%q accepted", bad) + } + } + for _, good := range []string{"mesh-web", "aaaaaaaaaaaa", "registry.mesh.internal:5100/x@sha256:abc", "dev_db.1"} { + if _, err := Ref(good); err != nil { + t.Errorf("%q refused: %v", good, err) + } + } + f := machine() + for _, verb := range []string{"start", "stop", "restart"} { + if _, err := client(f, 1000).Act(context.Background(), verb, "--rm"); err == nil { + t.Errorf("%s took an option", verb) + } + } + if len(f.calls) != 0 { + t.Fatalf("docker was called: %+v", f.calls) + } +} + +func TestEveryContainerIsListedAndTheMeshsAreMarked(t *testing.T) { + c := client(machine(), 1000) + all, err := c.Containers(context.Background(), "", "", "") + if err != nil || len(all) != 2 { + t.Fatalf("%v %+v", err, all) + } + web, db := all[1], all[0] + if !web.MeshHeld || web.HeldBy != "hello-web.server" || web.Module != "hello-web" || web.Health != "healthy" { + t.Errorf("held: %+v", web) + } + if !reflect.DeepEqual(web.Ports, []string{"0.0.0.0:8080->80/tcp"}) || web.Mounts[0].Name != "webdata" { + t.Errorf("ports/mounts: %+v", web) + } + if db.MeshHeld || db.Compose != "dev" || db.ComposeDir != "/home/op/dev" || db.FinishedAt == "" { + t.Errorf("stray: %+v", db) + } + mesh, _ := c.Containers(context.Background(), "mesh", "", "") + other, _ := c.Containers(context.Background(), "other", "", "") + if len(mesh) != 1 || mesh[0].Name != "mesh-web" || len(other) != 1 || other[0].Name != "dev-db" { + t.Errorf("held filter: %+v / %+v", mesh, other) + } + if _, err := c.Containers(context.Background(), "mine", "", ""); err == nil { + t.Error("an unknown held filter was accepted") + } +} + +func TestNoContainersIsAnEmptyListAndAFailureIsAnError(t *testing.T) { + got, err := client((&fake{}).on("docker ps", Ran{}), 1000).Containers(context.Background(), "", "", "") + if err != nil || got == nil || len(got) != 0 { + t.Fatalf("%v %v", got, err) + } + if _, err := client((&fake{}).on("docker ps", Ran{Status: 1, Stderr: "Cannot connect to the Docker daemon\n"}), 1000).Containers(context.Background(), "", "", ""); err == nil { + t.Fatal("a daemon that does not answer read as no containers") + } +} + +func TestInspectLeavesTheEnvironmentsValuesOut(t *testing.T) { + got, err := client(machine(), 1000).Inspect(context.Background(), "mesh-web") + if err != nil { + t.Fatal(err) + } + b, _ := json.Marshal(got) + if strings.Contains(string(b), "hunter2") || !strings.Contains(string(b), `"PASSWORD"`) || got["mesh_held"] != true { + t.Fatalf("%s", b) + } +} + +func TestActingOnAMeshContainerSaysTheHostRestoresIt(t *testing.T) { + f := machine().on("docker stop", Ran{}).on("docker start", Ran{}) + got, err := client(f, 1000).Act(context.Background(), "stop", "mesh-web") + if err != nil { + t.Fatal(err) + } + if !f.ran("docker stop --time 10 mesh-web") || got["mesh_held"] != true || !strings.Contains(got["note"].(string), "host restores") { + t.Fatalf("%v %+v", got, f.calls) + } + got, _ = client(f, 1000).Act(context.Background(), "start", "dev-db") + if _, noted := got["note"]; noted || got["mesh_held"] != false { + t.Fatalf("a stray was noted: %v", got) + } +} + +func TestPruneIsADryRunByDefaultAndNeverTouchesAVolumeOrAMeshContainer(t *testing.T) { + f := machine(). + on("docker image ls --no-trunc --filter dangling=true", Ran{Stdout: `{"ID":"sha256:dead","Size":"1.5GB"}` + "\n"}). + on("docker system df --format", Ran{Stdout: `{"Type":"Build Cache","TotalCount":"3","Size":"2GB","Reclaimable":"1GB"}` + "\n"}). + on("docker image prune", Ran{Stdout: "Deleted Images:\nx\n\nTotal reclaimed space: 1.5GB\n"}). + on("docker builder prune", Ran{Stdout: "Total:\t1GB\n"}). + on("docker container rm", Ran{}) + c := client(f, 1000) + got, err := c.Prune(context.Background(), PruneAsk{Images: true, BuildCache: true, Containers: true, DryRun: true}) + if err != nil { + t.Fatal(err) + } + if f.ran("docker image prune") || f.ran("docker builder prune") || f.ran("docker container rm") { + t.Fatalf("a dry run removed something: %+v", f.calls) + } + if got["images"].(map[string]any)["dangling"] != 1 || !reflect.DeepEqual(got["containers"].(map[string]any)["stopped_not_held"], []string{"dev-db"}) { + t.Fatalf("%v", got) + } + got, err = c.Prune(context.Background(), PruneAsk{Images: true, BuildCache: true, Containers: true, OlderThanH: 24}) + if err != nil { + t.Fatal(err) + } + if !f.ran("docker image prune --force --filter until=24h") || !f.ran("docker builder prune --force --filter until=24h") || !f.ran("docker container rm dev-db") { + t.Fatalf("not pruned: %+v", f.calls) + } + for _, c := range f.calls { + line := strings.Join(c.args, " ") + if strings.Contains(line, "volume") || strings.Contains(line, "mesh-web") && c.args[0] != "container" || strings.Contains(line, "--volumes") || strings.Contains(line, "--all") && c.args[0] != "ps" { + t.Errorf("prune reached too far: %s", line) + } + } + if got["images"].(map[string]any)["reclaimed"] != "1.5GB" || got["build_cache"].(map[string]any)["reclaimed"] != "1GB" { + t.Errorf("reclaimed: %v", got) + } +} + +func TestLogsMergeBothStreamsInOrderAndKeepTheTail(t *testing.T) { + f := (&fake{}).on("docker logs", Ran{Stdout: "2026-10-04T10:00:01Z out one\n2026-10-04T10:00:03Z out two\n", Stderr: "2026-10-04T10:00:02Z err one\n"}) + got, err := client(f, 1000).Logs(context.Background(), "web", 2, "30m") + if err != nil { + t.Fatal(err) + } + if !reflect.DeepEqual(got["lines"], []string{"2026-10-04T10:00:02Z err one", "2026-10-04T10:00:03Z out two"}) { + t.Fatalf("%v", got["lines"]) + } + if !f.ran("docker logs --timestamps --tail 2 --since 30m web") { + t.Fatalf("%+v", f.calls) + } + if _, err := client(f, 1000).Logs(context.Background(), "web", 2, "--follow"); err == nil { + t.Fatal("since took an option") + } + f = (&fake{}).on("docker logs", Ran{Status: 1, Stderr: "Error response from daemon: No such container: nope\n"}) + if _, err := client(f, 1000).Logs(context.Background(), "nope", 2, ""); err == nil { + t.Fatal("a missing container read as no lines") + } +} + +func TestSizesAreReadAsDockerPrintsThem(t *testing.T) { + for in, want := range map[string]int64{"0B": 0, "55.63GB": 55630000000, "33.2MiB": 34812723, "1.5kB": 1500, "12MB (34%)": 12000000, "N/A": -1} { + if got := Bytes(in); got != want { + t.Errorf("%s: %d, want %d", in, got, want) + } + } +} + +func TestDaemonConfigSaysWhatTheDaemonHasNotTakenYet(t *testing.T) { + f := (&fake{}).on("docker info", Ran{Stdout: `{"ServerVersion":"29.8.2","LiveRestoreEnabled":false,"LoggingDriver":"json-file","RegistryConfig":{"IndexConfigs":{"docker.io":{"Secure":true},"registry.mesh.internal:5100":{"Secure":false}}}}`}) + c := client(f, 1000) + c.ReadFile = func(string) ([]byte, error) { + return []byte(`{"live-restore": true, "dns": ["10.0.0.1"], "log-driver": "local"}`), nil + } + got, err := c.DaemonConfig(context.Background()) + if err != nil { + t.Fatal(err) + } + pending := strings.Join(got["pending"].([]string), "\n") + if !strings.Contains(pending, "live-restore is true in the file and false") || !strings.Contains(pending, "log-driver is local") { + t.Errorf("pending: %s", pending) + } + if !reflect.DeepEqual(got["daemon"].(map[string]any)["InsecureRegistries"], []string{"registry.mesh.internal:5100"}) { + t.Errorf("registries: %v", got["daemon"]) + } + if !reflect.DeepEqual(got["read_only_at_start"], []string{"dns", "log-driver"}) { + t.Errorf("start-only: %v", got["read_only_at_start"]) + } + c.ReadFile = func(string) ([]byte, error) { return nil, os.ErrNotExist } + got, _ = c.DaemonConfig(context.Background()) + if !strings.HasPrefix(got["file_state"].(string), "absent") { + t.Errorf("absent: %v", got["file_state"]) + } + c.ReadFile = func(string) ([]byte, error) { return nil, errors.New("permission denied") } + got, _ = c.DaemonConfig(context.Background()) + if !strings.HasPrefix(got["file_state"].(string), "unreadable") { + t.Errorf("unreadable: %v", got["file_state"]) + } +} + +func TestEventsAreABoundedWindowWithoutExecNoise(t *testing.T) { + out := `{"Type":"container","Action":"exec_start: pg_isready","Actor":{"ID":"aaaaaaaaaaaaaaaa","Attributes":{"name":"db"}},"timeNano":1} +{"Type":"container","Action":"die","Actor":{"ID":"aaaaaaaaaaaaaaaa","Attributes":{"name":"web","mesh-host.id":"hello-web.server","exitCode":"137"}},"timeNano":2} +` + f := (&fake{}).on("docker events", Ran{Stdout: out}) + got, err := client(f, 1000).Events(context.Background(), 30, "container", 10, false) + if err != nil { + t.Fatal(err) + } + evs := got["events"].([]map[string]any) + if len(evs) != 1 || evs[0]["action"] != "die" || evs[0]["mesh_held"] != true || evs[0]["exit_code"] != "137" { + t.Fatalf("%v", evs) + } + if !f.ran("docker events --since 30m --until 0s --format {{json .}} --filter type=container") { + t.Fatalf("%+v", f.calls) + } + if _, err := client(f, 1000).Events(context.Background(), 30, "secret", 10, false); err == nil { + t.Fatal("an unknown type was accepted") + } +} + +func TestVolumesSayWhoMountsThemAndWhetherTheMeshDoes(t *testing.T) { + f := machine(). + on("docker volume ls --quiet", Ran{Stdout: "webdata\ndbdata\nloose\n"}). + on("docker volume inspect", Ran{Stdout: `[{"Name":"webdata","Driver":"local"},{"Name":"dbdata","Driver":"local"},{"Name":"loose","Driver":"local","Labels":{"com.docker.volume.anonymous":""}}]`}) + got, err := client(f, 1000).Volumes(context.Background(), false, false) + if err != nil { + t.Fatal(err) + } + vols := got["volumes"].([]map[string]any) + if vols[0]["mesh_held"] != true || vols[1]["mesh_held"] != false || len(vols[2]["mounted_by"].([]map[string]any)) != 0 || vols[2]["anonymous"] != true { + t.Fatalf("%v", vols) + } + got, _ = client(f, 1000).Volumes(context.Background(), true, false) + if got["count"] != 1 { + t.Fatalf("unmounted: %v", got) + } +} + +func TestImagesNameTheirUsers(t *testing.T) { + f := machine().on("docker image ls", Ran{Stdout: `{"ID":"sha256:img1","Repository":"web","Tag":"1","Size":"100MB"} +{"ID":"sha256:img3","Repository":"","Tag":"","Size":"2GB"} +`}) + got, err := client(f, 1000).Images(context.Background(), "", "", 10) + if err != nil { + t.Fatal(err) + } + imgs := got["images"].([]Image) + if imgs[0].ID != "sha256:img3" || !imgs[0].Dangling || imgs[1].UsedBy[0] != "mesh-web" || !imgs[1].MeshUsed { + t.Fatalf("%+v", imgs) + } + got, _ = client(f, 1000).Images(context.Background(), "unused", "", 10) + if got["count"] != 1 { + t.Fatalf("unused: %v", got) + } +} + +func TestProblemsNameWhyAndUnlabelledIsTheCleanupList(t *testing.T) { + c := client(machine(), 1000) + p, err := c.Problems(context.Background()) + if err != nil || len(p) != 1 || p[0]["name"] != "dev-db" || p[0]["why"].([]string)[0] != "exited 1" { + t.Fatalf("%v %v", p, err) + } + u, err := c.Unlabelled(context.Background()) + if err != nil || u["count"] != 1 { + t.Fatalf("%v %v", u, err) + } +} diff --git a/modules/docker/cmd/docker-tools/main.go b/modules/docker/cmd/docker-tools/main.go new file mode 100644 index 0000000..6d1755b --- /dev/null +++ b/modules/docker/cmd/docker-tools/main.go @@ -0,0 +1,279 @@ +// docker's Go tools bundle (novox/hq ADR 0188, ADR 0193): a process the node's tool runtime launches +// and speaks MCP over stdio to, through the Go SDK. It answers for every container on this machine — +// the mesh's and every other — and for the runtime's images, networks, volumes, events and +// configuration. It runs as the operator account (ADR 0175 §4); docker.go says how it reaches the +// daemon's socket. The host applies the module's resources; these tools answer about the runtime. +package main + +import ( + "context" + "fmt" + "math" + "os" + "strings" + + stdio "git.novox.be/novox/mesh-sdk/go" +) + +func main() { + // An empty name serves as the module the runtime names (MESH_SERVED_MODULE): docker. + if err := stdio.Serve("", tools(NewClient())); err != nil { + fmt.Fprintln(os.Stderr, err) + os.Exit(1) + } +} + +var containerArg = map[string]any{"type": "string", "description": "the container's name or id"} + +func tools(c *Client) []stdio.Tool { + ctx := context.Background() + act := func(verb, description string) stdio.Tool { + return stdio.Tool{ + Name: "docker_" + verb, Description: description, + Input: map[string]any{"container": containerArg}, + Run: func(args map[string]any) (any, error) { + ref, err := text(args, "container") + if err != nil { + return nil, err + } + return c.Act(ctx, verb, ref) + }, + } + } + return []stdio.Tool{ + { + Name: "docker_list", + Description: "Every container on this machine — the mesh's and every other — with its image, state, health, restarts, " + + "published ports, mounts, compose project, and mesh_held/held_by (the assignment that holds it).", + Input: map[string]any{ + "held": map[string]any{"type": "string", "enum": []string{"all", "mesh", "other"}, "description": "whose: all (default), the mesh's, or the others"}, + "state": map[string]any{"type": "string", "description": "only containers in this state (running, exited, created, restarting, paused, dead)"}, + "match": map[string]any{"type": "string", "description": "only containers whose name or image contains this"}, + }, + Run: func(args map[string]any) (any, error) { + list, err := c.Containers(ctx, optional(args, "held"), optional(args, "state"), optional(args, "match")) + if err != nil { + return nil, err + } + return map[string]any{"count": len(list), "containers": list}, nil + }, + }, + { + Name: "docker_inspect", + Description: "One container whole, as docker inspects it, with mesh_held; its environment's values are left out (names kept), because that is where a container's secrets are.", + Input: map[string]any{"container": containerArg}, + Run: func(args map[string]any) (any, error) { + ref, err := text(args, "container") + if err != nil { + return nil, err + } + return c.Inspect(ctx, ref) + }, + }, + { + Name: "docker_logs", + Description: "The last lines one container wrote, both streams merged in order, each with its timestamp (default 200, at most 2000 lines; a line is cut at 4 KiB).", + Input: map[string]any{ + "container": containerArg, + "lines": map[string]any{"type": "integer", "description": "how many lines from the end (default 200, at most 2000)"}, + "since": map[string]any{"type": "string", "description": "only lines since then: a duration such as 30m or 2h, or a time"}, + }, + Run: func(args map[string]any) (any, error) { + ref, err := text(args, "container") + if err != nil { + return nil, err + } + n, err := bounded(args, "lines", 200, 2000) + if err != nil { + return nil, err + } + return c.Logs(ctx, ref, n, optional(args, "since")) + }, + }, + { + Name: "docker_stats", + Description: "What the running containers use now — CPU, memory, network and disk I/O, processes — the heaviest by memory first; or one container's.", + Input: map[string]any{"container": map[string]any{"type": "string", "description": "one container (optional)"}}, + Run: func(args map[string]any) (any, error) { + stats, err := c.Stats(ctx, optional(args, "container")) + if err != nil { + return nil, err + } + return map[string]any{"count": len(stats), "containers": stats}, nil + }, + }, + act("start", "Start one container. A container the mesh holds is started too, and the answer says the host restores what its declaration says at its next apply."), + act("stop", "Stop one container (ten seconds, then killed). For a container the mesh holds, the answer says the host will start it again at its next apply if its declaration says running."), + act("restart", "Restart one container (ten seconds to stop, then killed); the answer says whether the mesh holds it."), + { + Name: "docker_top", + Description: "The processes running inside one container: pid, user, elapsed time, CPU, resident memory and command.", + Input: map[string]any{"container": containerArg}, + Run: func(args map[string]any) (any, error) { + ref, err := text(args, "container") + if err != nil { + return nil, err + } + return c.Top(ctx, ref) + }, + }, + { + Name: "docker_images", + Description: "The images on this machine, the largest first, each with its size and the containers using it (and whether one of them is the mesh's). " + + "filter: all, dangling, unused or used.", + Input: map[string]any{ + "filter": map[string]any{"type": "string", "enum": []string{"all", "dangling", "unused", "used"}, "description": "which images (default all)"}, + "match": map[string]any{"type": "string", "description": "only images whose repository:tag contains this"}, + "limit": map[string]any{"type": "integer", "description": "how many to show (default 100, at most 1000); count says how many matched"}, + }, + Run: func(args map[string]any) (any, error) { + n, err := bounded(args, "limit", 100, 1000) + if err != nil { + return nil, err + } + return c.Images(ctx, optional(args, "filter"), optional(args, "match"), n) + }, + }, + { + Name: "docker_prune", + Description: "Reclaim space: dangling images and unused build cache, and — only when containers is true — stopped containers the mesh does not hold. " + + "Never a volume, never a container the mesh holds, never an image a container uses. A dry run by default: it lists what would go; dry_run false removes it.", + Input: map[string]any{ + "dry_run": map[string]any{"type": "boolean", "description": "list only (default true)"}, + "images": map[string]any{"type": "boolean", "description": "dangling images (default true)"}, + "build_cache": map[string]any{"type": "boolean", "description": "build cache nothing refers to (default true)"}, + "containers": map[string]any{"type": "boolean", "description": "stopped containers the mesh does not hold (default false); what they mounted is kept"}, + "older_than_hours": map[string]any{"type": "integer", "description": "only what is older than this many hours (default 0: any age)"}, + }, + Run: func(args map[string]any) (any, error) { + older := 0 + if v, ok := args["older_than_hours"]; ok && v != nil && v != float64(0) { + n, err := bounded(args, "older_than_hours", 0, 24*365) + if err != nil { + return nil, err + } + older = n + } + return c.Prune(ctx, PruneAsk{ + DryRun: flag(args, "dry_run", true), Images: flag(args, "images", true), BuildCache: flag(args, "build_cache", true), + Containers: flag(args, "containers", false), OlderThanH: older, + }) + }, + }, + { + Name: "docker_disk_usage", + Description: "What the runtime takes on disk (docker system df -v): per kind — images, containers, volumes, build cache — the total, the active and the reclaimable, and the largest of each.", + Input: map[string]any{"top": map[string]any{"type": "integer", "description": "how many of the largest per kind (default 10, at most 100)"}}, + Run: func(args map[string]any) (any, error) { + n, err := bounded(args, "top", 10, 100) + if err != nil { + return nil, err + } + return c.DiskUsage(ctx, n) + }, + }, + { + Name: "docker_networks", + Description: "Every network the runtime has: driver, scope, subnets and gateway, and the running containers on it with their addresses and whether the mesh holds them.", + Run: func(map[string]any) (any, error) { return c.Networks(ctx) }, + }, + { + Name: "docker_volumes", + Description: "Every volume with the containers mounting it, whether the mesh holds any of them, whether it is anonymous, its compose project, and — when sizes is true (slower) — its size.", + Input: map[string]any{ + "unmounted": map[string]any{"type": "boolean", "description": "only volumes no container mounts (default false)"}, + "sizes": map[string]any{"type": "boolean", "description": "measure each volume (default false: it walks every volume)"}, + }, + Run: func(args map[string]any) (any, error) { + return c.Volumes(ctx, flag(args, "unmounted", false), flag(args, "sizes", false)) + }, + }, + { + Name: "docker_events", + Description: "What the runtime did in a window ending now (default the last 60 minutes, at most 24 hours): containers created, started, died, health changes, images pulled — with mesh_held. Exec events are left out unless asked.", + Input: map[string]any{ + "minutes": map[string]any{"type": "integer", "description": "how far back (default 60, at most 1440)"}, + "type": map[string]any{"type": "string", "description": "only one kind: container, image, network, volume, daemon, plugin or builder"}, + "limit": map[string]any{"type": "integer", "description": "the latest this many (default 200, at most 2000)"}, + "execs": map[string]any{"type": "boolean", "description": "include exec_* events (default false: health checks make many)"}, + }, + Run: func(args map[string]any) (any, error) { + minutes, err := bounded(args, "minutes", 60, 1440) + if err != nil { + return nil, err + } + limit, err := bounded(args, "limit", 200, 2000) + if err != nil { + return nil, err + } + return c.Events(ctx, minutes, optional(args, "type"), limit, flag(args, "execs", false)) + }, + }, + { + Name: "docker_daemon_config", + Description: "The runtime's configuration: /etc/docker/daemon.json as it is on disk, the daemon's essentials as it runs now (docker info: version, storage and logging drivers, " + + "live restore, root directory, insecure registries, warnings), and where the two differ — keys a reload or only a restart would take.", + Run: func(map[string]any) (any, error) { return c.DaemonConfig(ctx) }, + }, + { + Name: "docker_unlabelled", + Description: "The containers the mesh does not hold — the cleanup list — each with its image, state, compose project and directory, ports and mounts.", + Run: func(map[string]any) (any, error) { return c.Unlabelled(ctx) }, + }, + { + Name: "docker_problems", + Description: "Every container that is not well: unhealthy, restarting, dead, killed for memory, exited with a failure, or restarted five times or more — with whether the mesh holds it.", + Run: func(map[string]any) (any, error) { + p, err := c.Problems(ctx) + if err != nil { + return nil, err + } + return map[string]any{"count": len(p), "containers": p}, nil + }, + }, + { + Name: "docker_ports", + Description: "Every port the containers publish on this machine (address:port -> container port), and the containers on the host's network, which publish whatever they listen on.", + Run: func(map[string]any) (any, error) { + p, err := c.Ports(ctx) + if err != nil { + return nil, err + } + return map[string]any{"count": len(p), "ports": p}, nil + }, + }, + } +} + +func text(args map[string]any, key string) (string, error) { + s, _ := args[key].(string) + if s = strings.TrimSpace(s); s == "" { + return "", fmt.Errorf("%s is required", key) + } + return s, nil +} + +func optional(args map[string]any, key string) string { + s, _ := args[key].(string) + return strings.TrimSpace(s) +} + +func flag(args map[string]any, key string, def bool) bool { + if b, ok := args[key].(bool); ok { + return b + } + return def +} + +// bounded is a whole number argument, defaulted when absent and held to a ceiling. +func bounded(args map[string]any, key string, def, most int) (int, error) { + v, ok := args[key] + if !ok || v == nil { + return def, nil + } + f, ok := v.(float64) + if !ok || f != math.Trunc(f) || f < 1 { + return 0, fmt.Errorf("%s must be a whole number of at least 1", key) + } + return int(math.Min(f, float64(most))), nil +} diff --git a/modules/docker/cmd/docker-tools/runner.go b/modules/docker/cmd/docker-tools/runner.go new file mode 100644 index 0000000..01ed874 --- /dev/null +++ b/modules/docker/cmd/docker-tools/runner.go @@ -0,0 +1,59 @@ +package main + +import ( + "bytes" + "context" + "errors" + "fmt" + "os/exec" + "strings" + "time" +) + +// Ran is what a command did: its output, its exit status, and why it never ran to an answer. +type Ran struct { + Stdout string + Stderr string + Status int + // Err is "ENOENT" when the program is not installed, or says it was ended for taking too long. + Err string +} + +// Runner runs one command, so every tool can be tested without a daemon. +type Runner func(ctx context.Context, name string, args ...string) Ran + +// CallTimeout is how long one docker command may take: below the runtime's thirty-second call +// limit, so a daemon that hangs is answered as such rather than as a call the runtime gave up on. +const CallTimeout = 20 * time.Second + +// ExecRunner runs a command on this machine, bounded by CallTimeout. +func ExecRunner(ctx context.Context, name string, args ...string) Ran { + ctx, cancel := context.WithTimeout(ctx, CallTimeout) + defer cancel() + cmd := exec.CommandContext(ctx, name, args...) + var out, errb bytes.Buffer + cmd.Stdout, cmd.Stderr = &out, &errb + err := cmd.Run() + r := Ran{Stdout: out.String(), Stderr: errb.String()} + var exitErr *exec.ExitError + switch { + case errors.Is(ctx.Err(), context.DeadlineExceeded): + r.Status, r.Err = 124, fmt.Sprintf("no answer within %d s", int(CallTimeout/time.Second)) + case errors.Is(err, exec.ErrNotFound): + r.Status, r.Err = 127, "ENOENT" + case errors.As(err, &exitErr): + r.Status = exitErr.ExitCode() + case err != nil: + r.Status, r.Err = 1, err.Error() + } + return r +} + +func firstLine(s string) string { + for _, l := range strings.Split(s, "\n") { + if l = strings.TrimSpace(l); l != "" { + return l + } + } + return "" +} diff --git a/modules/docker/cmd/docker-tools/tools_test.go b/modules/docker/cmd/docker-tools/tools_test.go new file mode 100644 index 0000000..d29f054 --- /dev/null +++ b/modules/docker/cmd/docker-tools/tools_test.go @@ -0,0 +1,60 @@ +package main + +import ( + "encoding/json" + "os" + "reflect" + "sort" + "strings" + "testing" +) + +func TestTheToolsServedAreTheToolsTheManifestNames(t *testing.T) { + raw, err := os.ReadFile("../../module.json") + if err != nil { + t.Fatal(err) + } + var m struct { + Tools []string `json:"tools"` + } + if err := json.Unmarshal(raw, &m); err != nil { + t.Fatal(err) + } + served := []string{} + for _, tool := range tools(client(&fake{}, 1000)) { + if !strings.HasPrefix(tool.Name, "docker_") || tool.Description == "" || tool.Run == nil { + t.Errorf("tool %q", tool.Name) + } + served = append(served, tool.Name) + } + sort.Strings(served) + listed := append([]string{}, m.Tools...) + sort.Strings(listed) + if !reflect.DeepEqual(served, listed) { + t.Fatalf("served %v, manifest %v", served, listed) + } +} + +func TestNumbersAreDefaultedAndBounded(t *testing.T) { + if n, _ := bounded(map[string]any{}, "lines", 200, 2000); n != 200 { + t.Error(n) + } + if n, _ := bounded(map[string]any{"lines": float64(99999)}, "lines", 200, 2000); n != 2000 { + t.Error(n) + } + for _, bad := range []any{float64(0), float64(-1), float64(1.5), "10"} { + if _, err := bounded(map[string]any{"lines": bad}, "lines", 200, 2000); err == nil { + t.Errorf("%v accepted", bad) + } + } +} + +func TestAStopFromTheToolNeedsAContainer(t *testing.T) { + for _, tool := range tools(client(&fake{}, 1000)) { + if tool.Name == "docker_stop" { + if _, err := tool.Run(map[string]any{}); err == nil { + t.Fatal("a stop without a container was accepted") + } + } + } +} diff --git a/modules/docker/go.mod b/modules/docker/go.mod new file mode 100644 index 0000000..5faa536 --- /dev/null +++ b/modules/docker/go.mod @@ -0,0 +1,5 @@ +module docker + +go 1.22 + +require git.novox.be/novox/mesh-sdk/go v0.1.6 diff --git a/modules/docker/go.sum b/modules/docker/go.sum new file mode 100644 index 0000000..0dd6061 --- /dev/null +++ b/modules/docker/go.sum @@ -0,0 +1,2 @@ +git.novox.be/novox/mesh-sdk/go v0.1.6 h1:9qzdYONYbJdWcu6sxQcq9v1LI0JxcfkiKYkMUzJSkVQ= +git.novox.be/novox/mesh-sdk/go v0.1.6/go.mod h1:GFuZUElBZ9A++mxgIKo97aXXo+kV0uJ/UkbhQPPIbrY= diff --git a/modules/docker/module.json b/modules/docker/module.json new file mode 100644 index 0000000..56d12a2 --- /dev/null +++ b/modules/docker/module.json @@ -0,0 +1,94 @@ +{ + "module": "docker", + "version": "1", + "capabilities": [ + "package-manager", + "service-manager", + "privileged" + ], + "claims": [ + { + "name": "node-container-runtime", + "scope": "node" + } + ], + "tools": [ + "docker_list", + "docker_inspect", + "docker_logs", + "docker_stats", + "docker_start", + "docker_stop", + "docker_restart", + "docker_top", + "docker_images", + "docker_prune", + "docker_disk_usage", + "docker_networks", + "docker_volumes", + "docker_events", + "docker_daemon_config", + "docker_unlabelled", + "docker_problems", + "docker_ports" + ], + "resources": [ + { + "id": "package", + "type": "package", + "package": "docker" + }, + { + "id": "buildx", + "type": "package", + "package": "docker-buildx" + }, + { + "id": "socket", + "type": "service", + "unit": "docker.socket", + "state": "running", + "boot": "enabled" + }, + { + "id": "prune-service", + "type": "file", + "path": "/etc/systemd/system/docker-prune.service", + "mode": "0644", + "content": "# Generated by the mesh. Do not edit — module docker writes this file and replaces it at every push.\n[Unit]\nDescription=Prune dangling images and unused build cache (the mesh's docker module)\n# Never volumes, never a container, never an image a container uses: dangling\n# images and build cache nothing refers to, unused for a week. What a person\n# prunes beyond that is docker_prune's, by hand.\nAfter=docker.service\nConditionPathExists=/run/docker.sock\n\n[Service]\nType=oneshot\nNice=19\nIOSchedulingClass=idle\nExecStart=/usr/bin/docker image prune --force --filter until=168h\nExecStart=/usr/bin/docker builder prune --force --filter until=168h\n" + }, + { + "id": "prune-timer", + "type": "file", + "path": "/etc/systemd/system/docker-prune.timer", + "mode": "0644", + "content": "# Generated by the mesh. Do not edit — module docker writes this file and replaces it at every push.\n[Unit]\nDescription=Weekly prune of dangling images and unused build cache (the mesh's docker module)\n\n[Timer]\nOnCalendar=weekly\nRandomizedDelaySec=1h\nPersistent=true\n\n[Install]\nWantedBy=timers.target\n" + }, + { + "id": "prune", + "type": "service", + "unit": "docker-prune.timer", + "state": "running", + "boot": "enabled", + "restart-on": [ + "prune-service", + "prune-timer" + ] + } + ], + "build": { + "artifacts": [ + { + "name": "tools", + "kind": "bundle", + "language": "go", + "system": "arch", + "from": "cmd/docker-tools", + "binary": "docker-tools", + "loads": [ + "docker-tools" + ] + } + ] + } +}