mesh-controller#259 fixes while-stopped to name the container as the machine
knows it — `distribution.store`, not `store`. Until that controller is the
one composing, novox refuses its whole declaration and takes nothing at all.
The step comes out; deletion stays on, already applied and harmless on its
own. A collect step without its window would be worse than none: garbage
collection against a live registry can sweep a blob a build is pushing.
Put back once the fixed controller is deployed and stays.
The distribution packages both, so they are installed as packages rather than cloned or
vendored (novox/hq ADR 0205). Each contributes the line that loads the package's own
copy, from the path the Arch package installs, to a slot of the login shell's block
(ADR 0204). Syntax highlighting goes in last, as its upstream asks.
The distribution does not package the theme, and the predecessor cloned whatever
upstream's default branch held the day a hook ran (novox/hq ADR 0205). So upstream's
v1.20.0 release is vendored verbatim, with its licence, and shipped as an archive the
host unpacks under the account's home and checks by digest.
The prompt's configuration is today's ~/.p10k.zsh byte for byte, as a second archive.
Inline, its 86 KB would ride in every declaration and be unreviewable JSON. The zsh code
that loads both is a contribution to the normal slot (ADR 0204). Instant prompt stays off,
as it is today.
The seat is now the mesh's node-login-shell, which a shell module claims rather than
declares (novox/hq ADR 0204), and the environment is one module's that every module
contributes to (ADR 0203). Per hq to-be 41 WP3:
- no seat declaration; the claim is node-login-shell serving execute;
- EDITOR, VISUAL, XDG_CONFIG_HOME and the three PATH entries are an environment
contribution, not exports in the block;
- a ~/.zshenv block sources ~/.config/mesh/environment.sh, so a script, a login and
execute all see the environment;
- the ~/.zshrc block goes at the start, so the operator's lines run after it, and holds
today's shared defaults between the first, normal and last slots. The prompt, the
plugins and the operator's own lines are no longer in it;
- execute is bounded below the runtime's call limit (20 s default, 25 s at most), kills its
whole process group on timeout, cuts each stream at 256 KiB and says so, runs in the
account's home without the mesh's words, with the account's session words. The dead
runuser branch is gone, because the runtime is the account;
- zsh_config shows both files with their block line counts;
- the README lists the one-off migration (ADR 0182).
Every module contributes variables and PATH entries as facts, and one holder of
node-environment places them (novox/hq ADR 0203, to-be 41 WP3). This is that holder: no
package, no process, no tools — the directories it owns under the home and two files the
controller fills, the POSIX file at the path the seat fixes (~/.config/mesh/environment.sh,
sourced by the login shell) and environment.d's 50-mesh.conf for the account's service
manager and graphical session.
The node tools runtime runs as the operator account, not root, and gives its bundles no
session words (novox/hq ADR 0175, 0188, 0193). So, per hq to-be 41 WP4:
- system-scope start/stop/restart/enable/disable go through sudo -n when not root, as the
packet filter and intrusion prevention do, and a refusal is named by how it failed;
- user scope is plain --user with XDG_RUNTIME_DIR and the session bus of /run/user/<uid>;
the dead --machine branches are gone;
- a failed systemctl or journalctl is an error, and an unreachable user manager is said
even when systemctl exits 0; systemd_failed reports it beside the other manager's answer
instead of claiming nothing failed;
- status says whether the mesh declares the unit: its loaded unit file begins with the
header the host writes for a module's process. Only such a unit carries the restore note;
- the package resource goes: the service manager is always present, and it collided with
systemd-networkd's identical declaration;
- calls are bounded below the runtime's call limit, a unit name is never an option, and
the runner is injected so the tests use a fake one.
One key per registration in the module's servers bucket — all.<server> or
<node>.<server> — watched by every node, so a node assigned after a
registration takes it at start, which the mcp.registered event could not do.
Also narrows apply()'s refusal by hand: the builder compiles without strict,
where the discriminated union does not narrow and the build failed.
The bundles refactor took ADR 0188 on main, so minio's comments cite 0201.
The sdk is 0.1.7 after the same rebase, and minio needs the `derived` field
it carries.
REGISTRY_STORAGE_DELETE_ENABLED on the server — the door already accepts a
push — and a scheduled step running the registry's own collector over the
volume at 03:30 with the server held still. Plain garbage-collect: what the
mesh keeps is still a manifest, so --delete-untagged is not needed and would
delete images machines are running.
serves.s3-bucket.bucket is ${consumer:as:dns}; the provisioner uses what it
is given. nextcloud, invoicing and photos ask for ${bound:s3-bucket:bucket}
instead of naming mesh-novox-* literals, which also named this node.
bucketFor and the long-dead accessKeyFor are gone.
The holder of the seat the controller seeds under novox/hq ADR 0177. Eight
verbs under the seat's name — units, status, start, stop, restart, enable,
disable, journal — each taking an optional scope, "system" by default or
"user" for the operator account's own manager, reached as
`systemctl --user --machine=<account>@` when the runtime is not that account.
One tool of its own, systemd_failed, for every failed unit in both scopes.
A package, a claim and a bundle; no container, no process: served by the node
tools runtime (ADR 0175) once it exists. `module check` passes against a
controller that carries the seat; the tools type-check against the SDK.
The first module of the operator's environment (novox/hq to-be 37 §1, ADR 0173,
0176). A package, the mesh's default configuration as a block inside the
account's ~/.zshrc so the operator's own lines around it survive every push
(ADR 0174 as the host's `into: block` realises it), a `user` shape that makes
zsh the account's login shell, the `login-shell` seat declared with its one
verb, and a tools bundle: `execute` under the seat's name, `zsh_config` under
the module's. No container, no process: the tools are served by the node tools
runtime (ADR 0175), which does not exist yet — the bundle builds and the
manifest registers ahead of it. `module check` passes; the tools type-check
against the SDK.
Two things the manifest cannot yet say, left for the controller: the `user`
shape applies wherever the module is assigned, not only where it holds the
seat; and the runtime learns the account from MESH_OPERATOR_ACCOUNT, which
nothing sets yet.
Events carry what happened and no secret; tokens travel on requests (design 32 §10). The manager's
licence.rotated/switched events make the module ask anthropic-licence-manager.current; at start it
asks once to catch up. A refresh token appearing in the credentials file is a login: it is pushed to
the manager's adopt at once, sealed to the manager's key — the one time a refresh token travels. A
switch replaces the old licence's grant whole, removes the API key and its helper, and rewrites
oauthAccount in ~/.claude.json. New tools register and unregister MCP servers on this node, or with
nodes: all / a list via an mcp.registered event every node consumes; called for one node, the
answer names the other nodes running claude-code. 26 tests.
Node 25 refuses an IP address as the TLS server name, and the module reaches its server on
loopback, so every connection failed on the live machines. The certificate is trusted either way.
The test subscribed '**', which its in-memory broker never matched, while the module subscribes
'#'; and it still expected the module-qualified type from before event names became local.
The mesh-mssql container goes with its Dockerfile (and the sqlcmd it fetched), build bases and bus credential. The client speaks TDS through the mssql driver its package.json names, inlined into the bundle by the builder (ADR 0198 §4): one session per call as one sqlcmd invocation was, FOR JSON rendering rows exactly as before, the consumer's password checked as a bound parameter. A caller's statement still runs only as the reader login (issue 193); the one-line rule and -x guarded against sqlcmd's own commands and variable substitution, which no longer stand between the caller and the server. The server is reached on loopback at the port the machine published (${port:1433}). The reader test drives a fake session in place of a fake sqlcmd.
The mesh-mongodb container goes with its Dockerfile, build bases and bus credential. Its client shelled out to mongosh, which no machine's system carries, so it now speaks to the server through the official mongodb driver its package.json names, inlined into the bundle by the builder (ADR 0198 §4); the tools answer exactly as before (relaxed Extended JSON). The server is reached on loopback at the port the machine published (${port:27017}). The root secret was owned by the mongo image's user (secrets-owner 999:999), which the runtime's account cannot read; the module's own copy is now the runtime's, and the server is given its own 999-owned copy rendered from the same secret.
The mesh-catalog container goes with its Dockerfile, build bases, bus credential and mesh-state directory. Its words are the database URL file where the mesh writes it. `pg` is a dependency in its package.json, which the builder now installs and inlines into the bundle (mesh-controller: a TypeScript bundle installs its module's own packages). `prepares: true` needs a container running the module's own artifact, so it becomes what ADR 0198 §3 says it is: prepare/index.js run by node as a run-once process, with the same words and no bus, before the runtime is started with the version that needs it, and again when the database URL changes.
The module's code moved out of its container and took mosquitto_ctrl from a host package. A
machine whose package index is stale cannot install it (hq issue 205), so the tools failed. The
broker's own image carries the tool at the broker's version: the tools exec into the running
broker, and the bootstrap seeds from a throwaway container of the same image.
Both containers go with the Dockerfile, build bases, bus credential and state directory. apply needs no bus and runs every five minutes as a process on the machine at the host paths the container mounted. usage emitted by spawning the runtime image's own emit command with the module's credential, which exists nowhere now, so it is loaded by the node's runtime instead: it emits through the SDK as this module and reads on the cadence the schedule gave it, once at start and every five minutes. That is the one code change.
The mesh-openai-consumer-apply container goes with its Dockerfile and build bases. The same entrypoint runs every five minutes as a process on the machine, reading the binding and writing the credentials at the host paths the container used to mount.
The mesh-route-adapter container goes with its Dockerfile and build bases. The step runs node on the bundle as a run-once process, reading what the mesh contributed and its config where the mesh writes them and writing the proxy's dynamic directory at the path the container used to mount; it still runs again when a route or its config changes.
The mesh-lab container goes with its Dockerfile, build bases, bus credential and state directory. What the image installed — git, make, python, file, iproute2, sudo, npm, go and the incus client — are packages of the machine, and docker and incus are reached through their sockets as the runtime's account. The forge is an operator's setting, which reaches a file and never a bundle's words, so the tools read it from the env-file the mesh already fills, at each call; that is the one code change.
The mesh-mailu container goes with its Dockerfile, the mesh-tools build bases and its bus credential; automx keeps its own image. The code reached the admin API by its name on the mailu network, which a process on the machine cannot, so the admin container publishes 8080 to this machine only and the bundle reaches it on loopback at that port. Mail is still read through docker exec into mailu-imap, so the runtime's account needs the docker socket as nextcloud's does.
The records container goes with its Dockerfile, build bases and bus credential. The checkout, the config file and the origin file are read where the mesh writes them, and git comes from the machine's git package instead of the image's apt layer.
The mesh-vault container goes with its Dockerfile, build bases, bus credential and state directory; its env was already host paths, so it becomes the bundle's words unchanged.
The mesh-gitea container goes with its Dockerfile, build bases and bus credential; its env becomes the bundle's words with mount targets folded back to host paths: the config file, the admin password and the kept-token state directory are read where the mesh writes them.
The mesh-audit-logger container goes with its Dockerfile, build bases and bus credential: its one entrypoint is a load of one bundle, which subscribes to every event through the runtime and writes the trail at the host path the container used to mount.
ADR 0193 says the node's runtime launches every served bundle over MCP stdio and knows no
language, and ADR 0188 says one module may carry several bundles in any language. Nothing in
the catalogue shows both at once: every tools bundle is TypeScript, and the only Go bundle is
the runtime itself. netcheck is the smallest real module that does — read-only checks from a
machine, worth having on their own:
- tools-go (Go SDK go/v0.1.6): netcheck_tcp (one connect, nothing sent) and netcheck_dns
(A/AAAA/CNAME/TXT/MX through the machine's resolver).
- tools-typescript (@novox/mesh-sdk): netcheck_http (HEAD or GET, body neither sent nor read,
redirects reported not followed, anything but http(s) refused).
Both say loads; the module lists its tools. No container, no image, no env: nothing to be
given, so the runtime's own words suffice.
So the controller's ownership check refuses a second module owning either. ~/.claude is the
operator's at 0700 (it was 0755 on the workstations); of what is inside, the module owns only what it
writes, and the host keeps a directory that is not empty when the module goes (hq ADR 0182).
Every bundle is now a child speaking MCP over stdio, so stdout is the channel: the module logs on
stderr. The managed CLAUDE.md teaches mesh_search, mesh_describe, mesh_call, mesh_overview and
mesh_machine with addresses (<seat>.<verb>, <node>/<module>.<tool>) instead of flat tool names.
A hand-over is applied whatever the trailing render says; a failed render is reported beside it.
Proven over stdio as the runtime drives it: five tools listed, a key made on first use, a sealed
switch writing an access-token-only 0600 credentials file that keeps unknown keys.
The module owns /etc/claude-code: managed-mcp.json lists the console as `mesh` over HTTP on
loopback plus the servers in its mcp_servers setting (exclusive, by the operator's choice — the
https rule of managedMcpServers refuses a loopback console); managed-settings.json carries the
attribution convention, keeps claude.ai connectors, and adds the key-helper only for an API-key
licence; CLAUDE.md says how a session here works. Rendered whenever the runtime collects the
tools, written only on change, through the operator account's sudo. Under the home, only the
credentials file, only on a hand-over. Nothing declared under a home or /etc; the console's
port comes from node-tools' mcp-endpoint (mesh-tools #34).
The parts of the agent module that hold whichever way the console is registered: X25519 +
HKDF + AES-GCM from Node's own library so the bundle carries no dependency; the predecessor's
lineage rule (rotation only if newer, a re-issue adopted, a switch regardless) with its
incidents as tests; an atomic 0600 write that strips any refresh token and keeps keys it does not
know; the account read from the agent's own state file. Manifest and renderer follow.
The mesh-minio container goes with its Dockerfile, build bases, bus credential and state directory. The client reaches minio on the published port, runs the minio-client package's mcli instead of the image's mc, and keeps mc's config, which holds the root alias, in the module's own state directory rather than a shared /tmp.
The mesh-nextcloud container goes with its Dockerfile, build bases and bus credential. occ still runs through docker exec, now with the host's own docker CLI and socket.
The mesh-nodered container goes with its Dockerfile, build bases and bus credential. The mqtt step runs node on the bundle and reads the binding and settings files where the mesh writes them, from an env-file the mesh fills because a process's env is not given ${port:…}.