The received grants file names each pair secret by its HOST path. A sidecar
that mounts the placed grants directory at another path inside its
container sees mesh.json and not the secrets beside it — mosquitto's
provisioner reported ENOENT for a file that was there, and nodered's mqtt
step failed against a login that was never created. Mounted at its own
path now, as postgres does.
The tool runtime's own client, mesh serve, started by the mesh on the credential it sealed to the
machine (novox/hq ADR 0152, design 34): invokes every tool, listens from the machine only, holds no
state. Checked with module check before it was ever registered (hq issue 148).
They came from `requires: secret`, minted by the vault — right for a fresh
setup (the image's INIT_* variables read them once), wrong for an instance
that already exists: setup is skipped, the minted values match nothing,
and the provisioner holds a token the server never issued (issue 100).
InfluxDB will not take a chosen token value, so the operator token must be
accepted from the instance (`secret accept ace influxdb admin-token`); the
password can be either. As own secrets both are minted for a fresh install
exactly as before, and accepted where the data already knows them.
Found migrating ace's influxdb.
A program that checks its resolver validates — Mailu's admin does, at
start — could not use the mesh's resolver, and the one it ships instead
knows no mesh name (novox/hq issue 171). proxy-dnssec copies the AD bit
from 1.1.1.1 and 8.8.8.8, both of which validate.
admin is the one container here that reaches something by a mesh name:
the database, bound in database.env. A mesh name is answered by the
machine's resolver, which the runtime hands every container that does
not name its own (novox/hq ADR 0148); Mailu's unbound knows no mesh
name, and since the mesh stopped copying names into containers admin
could not find its database. The others keep unbound — rspamd needs a
validating resolver for DNSBL lookups and asks for no mesh name.
dnsmasq admits a query by the interface it arrives on when told interface=,
and by the address it is sent to when told listen-address=. A container's
query is sent to the machine's private address but arrives on docker0, so
interface=mesh0 dropped it silently on every machine (novox/hq issue 110).
Name the address, not the interface.
novox/hq 04-ISSUES/110. The module writes the runtime's `dns` key and
deliberately ordered no restart, because a restart stops every container.
It also ordered no reload, and the key holds only for containers created
after the runtime next starts. On two of four machines the runtime
predated the file — one since August — so every container there was
handed a public resolver and no mesh name resolved, while everything read
as fine. The third machine works only because its runtime happened to
restart later.
The file now also sets live-restore, which the runtime reads on a reload,
and the module declares the runtime reloaded when the file changes (ADR
0102: reload, don't restart). A reload still does not make `dns` take
effect; what it does is make the one restart that key needs keep every
container running. That restart stays the operator's, once per machine,
and is harmless from the second time on.
The mesh restarts nothing here. Undeclared, the runtime's unit goes back
to the state it was found in (ADR 0118).
MESH_PROVISION_POSTGRES named 127.0.0.1:5432 literally; the seat twin
(${seat:mesh-store:5432}) corrected it only on the machine holding the
mesh-store seat. On any other machine — ace, where the module provides
postgres-database without the seat — the twin is empty and the sidecar
would have dialled whatever else holds 5432 (HAL's postgres). ${port:5432}
is the machine port on every node; on novox the twin still says the same
6852, so nothing moves there.
sonarr, radarr, lidarr and bookshelf reached jackett as http://jackett:9117 (a HAL
container name) or https://indexers.zurag.be (its public route), typed into each app by
hand. The mesh has neither: an app now requires jackett-api and its downloads step writes
the bound address into the app.
Serves scheme, port and url-base; `at` and the machine port come from the binding. Mesh
scope, like sonarr-api: an indexer proxy shares no files with its consumers.
The pair credential is jackett's one API key. The mesh cannot mint it, so the operator
accepts it per consumer pair (ADR 0092), as #156 does for sonarr-api; a consumer's step
refuses a minted value and names the accept.
Node-RED's one broker node pointed at zurag.be:1884, where nothing listens. nodered now requires
mqtt-topic (asking for every topic: flows follow the devices' own) and a run-once `mqtt` step —
declared last, restarted when the binding, credential or settings change — points the mesh's broker
nodes at the bound broker through Node-RED's admin API with the module's api-token: the node the
step makes itself when none is named, or the ones an assignment names in `mqtt.brokers`. Only host,
port, TLS and the login change; the broker is asked first whether it takes the login; the deploy is
against the revision read ("nodes", so only that node restarts) and a digest makes a rerun a no-op.
A broker node nobody named is never touched. settings.js keeps `mqtt` and `topics` out of Node-RED.
mqtt-topic served nothing: with two listens the mesh could not say which port a consumer dials, so a
consumer had to type 1883 into its config. It now serves the MQTT listener's port (the machine's,
once assigned) and the scheme, so `${bound:mqtt-topic:port}` fills.
The provisioner confined every consumer to `<as>/#`, which leaves nothing for the consumers the
broker exists for: Home Assistant discovers under homeassistant/# and tasmota/discovery/#, and
Node-RED's flows follow the devices' own topics. A consumer now contributes `topics` (MQTT topic
filters) to its mqtt-topic requirement and is granted exactly those; with none, its own subtree as
before. Settings merge into contributions, so an operator narrows a grant per assignment. The role
is brought to exactly the wanted ACLs (stale ones removed), `holds` checks the ACLs too, and an
invalid list is refused, never quietly narrowed. Only the role named for the consumer is touched:
a client carried from the predecessor's password file keeps its own.
grafana's data source and Node-RED's influxdb nodes reached ace's
InfluxDB by a LAN IP or a public name nobody routes, with a credential
somebody made by hand. Now a consumer requires influxdb-api and is told
where it is, which org and default bucket it serves, and signs in with
the password the mesh minted for the pair.
The credential is a v1-compatibility authorization, made per grant by
the new provisioner: InfluxDB 2.x generates API tokens itself and
ignores one the caller sends, so a v2 token could only be accepted by
hand per pair; a v1 authorization takes a caller-chosen password (8-72
characters, the mesh mints 40) and reads/writes every bucket as a
database of its name over InfluxQL and line protocol. A consumer
contributes `access` (read, write, read-write) and, for writing, the
buckets; a missing bucket is made and never deleted. Only
authorizations named mesh_* and marked [mesh] are ever changed or
removed; anything else of that name is refused and left alone.
The org and default bucket are served facts the assignment's settings
set, reaching both the consumers and the provisioner's config.json.
The module named /var/lib/letta, a layout no definition may carry (ADR
0112); state is now a placed directory holding the bindings and secrets.
letta had no route, while the server it replaces is reached by its public
name (a workflow calls it there). It now contributes one for its `web`
endpoint; reach is the assignment's.
The server needs an OpenAI key for agents on OpenAI models, and nothing
gave it one: `openai-api-key` is an own-secret, accepted from the
operator (a key someone chose, not one the mesh can mint). The
server-password is accepted the same way where a server already has
clients. Both, and the database password inside LETTA_PG_URI, stay in the
server's environment: letta 0.6.x reads settings from the environment
only, and its startup.sh starts an embedded PostgreSQL unless
LETTA_PG_URI is set - the declared reason now says so.
The runtime's tools never authenticated: the client sent only a Bearer
token, and 0.6.x's --secure mode checks X-BARE-PASSWORD ("password <it>")
and answers 401 otherwise. The client now sends both. Its password comes
from the runtime config file (the key client.ts reads first) instead of an
env-file, so the runtime container no longer carries a secret in its
environment.
Image: the same 0.6.8 image, now pinned by the index digest ace runs
rather than its amd64 manifest.
Verified: catalogue tests with MESH_CATALOGUE set; tsc -p tsconfig.json in
the mesh-tools build image. Throwaway containers: a fresh 0.6.8 with its
embedded PG and two blocks made through the API; stopped, copied, dumped
from the copy; restored (schema letta + vector pre-made by the superuser,
--no-owner --role, search_path set on the database as the original had
it) into a grant-shaped database on the postgres module's pgvector image
(PG17); started in this shape: alembic finds nothing to do, both blocks
are there, a wrong password is refused with 401, and the patched client
lists agents. Test containers and data removed.
The manifest stated /services/redis/data and /var/lib/redis-module - novox's
old layout, paths no definition may carry (ADR 0112). State is now the
assignment's root, grants and data are placed, and the config file, the
secret file, receives and grants all name them as ${dir:...}. Paths inside
the sidecar are its own view and are unchanged.
The data directory and config are owned 999:1000: the image's redis user is
uid 999 in gid 1000 (checked in both builds), which is who owns ace's data
today; 999:999 named a group the image does not use.
Image pinned to the 7.4.11-alpine build ace runs (2026-09-17); the old pin was
the same version, built in August. Older-than-running is never the pin.
Nothing is assigned it anywhere today, so no machine changes.
Verified: catalogue tests pass with MESH_CATALOGUE on this tree; the
declaration composes for ace with every path under /var/lib/redis. The pinned
image ran as a throwaway with a 0600 999:1000 config and a 0700 data dir:
unauthenticated PING is refused (NOAUTH), authenticated SET/GET works,
appendonly is on, the server runs as redis.
The manifest stated /var/lib/mssql, its grants and its SA file by path, and
declared it listens on 4848 - the port one machine's predecessor published,
which is an assignment's fact (ADR 0112, 0138). ace is moving its own
SQL Server (80 GB of work databases) onto the mesh, so the module has to be
the same on every machine.
- state is the assignment's root (place "."), grants is placed, and every
reference (sa.env, the env-file, the SA and grants mounts, receives,
grants, own-secrets) names them as ${dir:...}.
- the database endpoint listens on 1433, the port SQL Server uses; a machine
that must keep an older number pins it in its assignment.
- the image pin is unchanged: it is the digest ace runs today (CU27,
16.0.4295), the same as novox.
novox is untouched: rendered with novox's own setting ({"ports":{"1433":4848}})
through the controller's Declaration and Rules, every resource - paths,
container names, volumes, env-file, the 4848:1433 mapping, owners, modes - is
byte-identical to what main renders; the only difference is the firewall
rule's comment text (still port 4848, from the mesh).
Verified: catalogue tests pass with MESH_CATALOGUE on this tree. The pinned
image ran as a throwaway on a 0700 10001:0 data dir with a root-owned 0600
env-file (dummy SA), answered sqlcmd as sa; a scratch database stopped,
copied with cp -a, checksummed and started on the copy kept its rows and
CHECKSUM_AGG.
The sidecar runs on the host network and dialled 127.0.0.1:1880, the
software's port; the mesh publishes nodered on a machine port it assigns,
so the tools reached whatever else holds 1880, or nothing (hq 088).
The manifest named /services/influxdb and /var/lib/influxdb-module — one
machine's paths — and passed the admin password and token through the
environment. ace is moving its 2022 instance onto the mesh, so the module
has to be what it is on any machine.
- data, config and state are placed directories; the data keeps 1000:1000,
the image's influxdb user, which is who owns ace's data today.
- the init secrets reach the image through its own
DOCKER_INFLUXDB_INIT_{PASSWORD,ADMIN_TOKEN}_FILE; the vault's files are
mounted read-only. secrets-in-environment is gone.
- the sidecar reads its token from the same file (MESH_INFLUXDB_TOKEN_FILE,
added to client.ts) and reaches the server at its assigned machine port
(${port:8086}) instead of assuming 8086. The unused config-dir mount,
which held the CLI's copy of the admin token, is dropped.
- the api endpoint contributes a route: the web UI is how people use it,
and reach is the assignment's to say.
Verified: catalogue tests pass with MESH_CATALOGUE pointed at this tree.
The pinned 2.9.1 image, run on a scratch copy of ace's 2.4.0 data, opens
it, runs its metadata migrations (backing up the pre-upgrade bolt/sqlite)
and hashes the two stored tokens; /health passes. A fresh setup through
the _FILE variables, with dummy secrets as root-owned 0600 files, accepts
the token (200 on /api/v2/buckets) and the password (204 on /signin).
client.ts typechecks strict and reads the token file, tolerating the
endpoints key in its config.
The manifest named /services/jackett/config and /var/lib/mesh/jackett/config.json — host
paths ADR 0112 takes out of definitions. The config dir is now a pathless directory
(${dir:config}) and the runtime's config and route binding live in a placed state dir, as
searxng does.
The image is pinned to v0.24.2627-ls34, the digest ace runs today; the old pin
(v0.24.2517-ls16) was older than the running version.
The runtime reached jackett at a fixed 127.0.0.1:9117; it now uses ${port:9117}, the
machine port the mesh actually assigned.
The tools never loaded: the client needed an API key nobody set. Like sonarr/radarr read
config.xml, it now reads APIKey from Jackett's own ServerConfig.json (the config dir is
already mounted read-only), so no secret goes into an assignment. jackett_indexers called
/api/v2.0/indexers, which is the web UI's endpoint and answers an API key with a redirect;
it now reads the Torznab t=indexers feed, and treats Torznab's 200-with-<error> as a
failure.
Verified: catalogue key tests (MESH_CATALOGUE set, not skipped); a throwaway container of
the pinned image on a fresh 0700 1000:1000 config dir serves its UI; the client discovers
the key from the generated ServerConfig.json, lists 617 indexers through the Torznab feed,
searches via /results, and a wrong key is refused; client.ts typechecks under --strict.
The catalogue ran the image's defaults: no adminAuth, so a routed Node-RED
editor (which runs arbitrary code) was open to anyone who reached it, and
the module's own tools had no token to present to an install that was locked.
- settings.js (fixed, 0600, uid 1000) carries adminAuth: user admin checked
against the admin secret -- a minted password, or the bcrypt hash an
existing install held (accepted), so current logins keep working -- and a
static bearer token (api-token) the sidecar presents. It loads settings.json
beside it, the one mergeable file; endpoints is dropped there, and an
optional timeZone sets process.env.TZ (assignments cannot set env).
- The sidecar's runtime config is no longer merged; it carries the token.
- Directories are placed (state, data), the route binds into state.
- Image pinned to 5.0.7 (a649dd71), what ace runs; the old pin was 5.0.6.
- deployFlows asks for API v2: v1 answers 204 with no body, which the client
tried to parse as JSON.
Verified: catalogue tests pass against this tree. A throwaway 5.0.7 container
started with the generated files: anonymous /flows 401, bearer api-token 200,
bad token 401, password grant 200/403 with a minted password and with a
bcrypt-hash-accepted one; endpoints and timeZone do not reach /settings;
timeZone Europe/Brussels overrides TZ=Etc/UTC; v1 deploy 204, v2 deploy
answers {rev}.
The module stated /services/unifi/data and fixed machine ports (8443:8443 and
eight more) — one installation's layout and numbers, which a definition may not
carry (ADR 0038, 0112). Data is now a placed directory (${dir:data}, 1000:1000,
0700), the container publishes its own ports and the mesh assigns the machine
side; an assignment pins them where devices already know them. The L2 endpoint
names the port the software uses (1900), not the one a machine published it on.
The sidecar dialled https://127.0.0.1:8443, true only while the machine port
equals the container's; it now asks for ${port:8443}. Its controller password
was a setting (plaintext in the mesh DB); it is now an own-secret written into
the one mergeable file (ADR 0086). The username stays a setting.
The web UI is contributed as a route to the "web" endpoint over https with
insecure upstream (the controller's own self-signed tls), as mailu's web-tls —
what HAL's hand-written traefik file for unifi does today.
Image pinned to the manifest list ace runs (8.0.24-ls221); the old pin was its
amd64 child, so the image is unchanged.
Verified: catalogue tests pass with MESH_CATALOGUE set; a throwaway container
of the pinned image on a fresh 1000:1000/0700 data dir answers /status (8.0.24,
up) and /inform; the sidecar client built from a config.json carrying site,
password, username and an endpoints key reaches it and is refused only on the
dummy credentials.
The image seds ICECAST_*_PASSWORD from the environment into /etc/icecast.xml;
ADR 0086 wants secrets as files. icecast starts as root, reads its config, then
drops to uid 100, so a root-owned 0600 icecast.xml rendered with ${secret:...}
and mounted read-only works and the entrypoint's seds never fire (no env set).
The "secrets-in-environment" exemption and server.env are gone.
Also: directories are placed (state, logs owned 100:101 so the image's VOLUME
/var/log/icecast is not an anonymous volume per container, as HAL learned);
the server and sidecar share a module network, so the sidecar reaches
http://icecast:8000 instead of assuming machine port 8000 on the host; the
stream endpoint is routed (label "icecast"), as HAL served it via traefik.
Secrets remain mesh-vault grants (requires secret), now under ${dir:state}.
Verified: catalogue tests with MESH_CATALOGUE pointed at this tree; a
throwaway container of the pinned digest (the one ace runs) with the rendered
file (dummy secrets, root 0600, :ro): runs as icecast, status-json 200,
admin 401 without / 200 with the admin secret, a source PUT with the source
secret mounts, a listener receives it, a wrong source password gets 401, logs
land in the uid-100 directory.
The module stated /var/lib/mosquitto-module and /services/mosquitto/data —
novox's layout, a path no definition may carry (ADR 0112). State, grants and
data are now placed directories (${dir:state}, ${dir:grants}, ${dir:data}),
the admin secret lives beside the broker account under the mesh's own state,
and the receives/grants maps follow the grants directory. Paths inside the
sidecar are its own view and are unchanged.
Image pinned to the 2.1.2 build ace's predecessor runs (2026-09-17); the old
pin was the same version, built in June.
Found preparing ace, whose broker carries a password-file user (an IoT switch
and home-assistant). Carrying it is a data step, not a manifest one: the
migration repo has scripts/mosquitto-pwdfile-to-dynsec.py, which moves $7$
PBKDF2 entries into the dynsec store hash-for-hash (tested end to end).
The route binding still named /var/lib/searxng-module, the directory the
previous commit placed elsewhere — the host would have written it into a
directory nothing declares. Same shape as gitea and nextcloud.
The module ran searxng on the image's built-in settings, which serve html
only — so the module's own search tool (format=json) was refused by the
software it fronts. And there was no way to configure it per machine: the
only file settings reach was the sidecar's.
settings.yml is now the module's one mergeable file (JSON is YAML): generic
defaults in the manifest (json format on, limiter and image proxy off,
valkey wired), and whatever differs per machine — base_url, method,
autocomplete, suspended times — set as the assignment's settings. The
secret key is filled on the machine through ${secret:secret}, so the
secrets-in-environment exception and the env file go. Directories are
placed. Image pinned to 2026.9.20, what ace runs today (the old pin was
older, 2026.9.1).
The sidecar's config.json is no longer mergeable: settings merge into every
mergeable file of a module, and the sidecar would have received searxng's
keys. It only ever read an optional url, which its env already carries.
Verified on ace: the pinned image serves html and json from a read-only,
root-owned 0600 JSON settings.yml.
novox/hq ADR 0147, issue 129. Every internal HTTPS name fails verification
on every machine: the certificates are genuine and nothing on a machine has
ever been told what issued them. The proxy's fetch answers for the proxy and
for nothing else — a browser, git over HTTPS and every module calling another
by an internal name read the machine's own trust store.
The module requires internal-acme-ca, fetches the root over the mesh's own
network (no prior trust to have; that is what this establishes), installs it
among the machine's anchors and refreshes the extracted bundles. Being
unassigned stops the unit, and stopping it takes the anchor away and
refreshes them again.
Arch's layout is named out loud: a machine that keeps anchors elsewhere fails
visibly rather than writing a file nothing reads.
What was in the catalogue was the first thing I built, not the thing ADR 0146
describes. It dialled raw ports on machine addresses from one hosting form and
emitted nothing, so findings would have sat in a file on the machine — the exact
thing issue 145 is about. It was never registered, never assigned, and never ran.
0146 says names per hosting form, fetched over TLS with the certificate verified,
and machines discovered over the bus. That shares nothing with this but the word
checker, so it goes rather than being bent into shape. Recorded as work to be
analysed and built deliberately.
Connectivity is checked by hand in the meantime, against the services the mesh
already runs.
The mesh asserts three things are callable (ADR 0144) — what runs on the same
machine, another machine's service exposed to the private network, and another
machine's service exposed publicly — and has never checked any of them. The first
was broken for eleven hours while the mesh reported every machine healthy.
This runs on every machine, on the cadence the mesh already has, in its own
container: the same position every other module calls from. Not the host and not
the control plane, both of which reach these addresses by paths no ordinary caller
uses and would have passed throughout that outage.
**Its probe is its own endpoint, and that is the point.** Declared reachable over
the private network like any other service, so it is admitted by exactly the rule
that governs every internally-exposed service and fails when that rule is wrong.
The tempting target is a service every machine has, and those are the ones never
closed — ssh above all — which would have passed while the thing that actually
broke was a service exposed to the private network.
It resolves before it dials and says which failed, because a name that does not
resolve and a port that does not answer have different owners. One failure is not
a fault: a machine rebooting is ordinary, so a path is broken after consecutive
runs and the count travels with the result. It reports and repairs nothing.
novox/hq ADR 0145. Eight tests; the consecutive-failure logic proved by reverting
it once. Not yet registered or assigned.