The manifest named /services/radarr/config and /var/lib/mesh/radarr/route.json, host
paths ADR 0112 takes out of definitions. The config dir is now a pathless ${dir:config}
(0700, 1000:1000) mounted into the server and, read-only, into the runtime that reads the
API key from config.xml; the route binding lives in a placed state dir, as jackett and
searxng do. /var/lib/mesh/radarr stays for the broker secret.
The image is pinned to 6.4.4.10685-ls317, the digest ace runs today. The old pin
(6.3.0.10514-ls314) was older than the running version, and Radarr migrates its database
forward on start, so pointing the older build at ace's data is not safe.
The container mount points stay /movies and /downloads: Radarr's database stores its root
folder and the download clients' reported paths under exactly those names, and there are
no remote path mappings to absorb a change. The generic access paths stay; ace's
(/storage/media/movies, /storage/downloads) and its media owner 1001:2000 wait on hq 153.
Based on feat/servarr-api-provision (#156), which makes radarr provide radarr-api.
Verified: catalogue key tests with MESH_CATALOGUE set (73 manifests parsed, not skipped);
a throwaway container of the pinned image on a fresh 0700 1000:1000 config dir answers
/ping, serves the v3 API with the key it generated and refuses a wrong one (401), accepts
/movies as a writable root folder; the runtime's client discovers the key from that
config.xml and reads the queue.
ombi keeps its Servarr connections in its own database, so the mesh has no
file to write them into. A run-once step reads the three bindings and pair
credentials and writes host, port, TLS, base path and key into ombi through
ombi's own API - only when they differ, and nothing else ombi keeps.
Until the operator accepts an app's key for this pair the mesh delivers a
value it minted, which no Servarr app accepts. The step tries the key against
the app first and, refused, writes nothing and fails naming the secret accept
that fixes it, so the old working key in ombi is never replaced by a dead one.
Declared last so its failing gates nothing else of ombi (ADR 0136), and
restart-on its six inputs so it runs again when a provider moves (ADR 0099).
A consumer on another machine (ombi first; jackett, bazarr and home-assistant
later) reached these by container name on HAL's shared network, which the mesh
does not have. Each app now provides <app>-api at mesh scope and serves the
software port, so the mesh tells a consumer where it is and redirects the
port to where the machine published it.
Named per app, not one servarr-api: a requirement is matched by name and
answered by exactly one provider per node, so a consumer cannot require one
name from three providers - and ombi's code is written against each app's
own API version (ADR 0027).
No grants and no provisioner: a Servarr instance has one API key, which the
mesh cannot mint. The operator accepts it as the pair credential for each
consumer (ADR 0092).
A host-network sidecar reaches its service over the machine's loopback, and
the mesh publishes that service on a machine port it assigns (ADR 0038) —
so dialling the software's port reaches whatever else holds it. On ace,
searxng's sidecar dialled 127.0.0.1:8080 and got unifi's inform port. The
same shape in bazarr, bookshelf, lidarr, nzbget, qbittorrent, radarr and
sonarr; each now asks with ${port:N} (hq 088). Found in review of ace's
module preparation.
ombi's definition named /services/ombi/config (a HAL machine path) in three
places and pinned an image older than the one ace runs. Ombi migrates its
own SQLite schema, so a take onto the older pin (v4.53.10-ls267) would start
it on a database the newer build (ls269) already touched.
- config is a pathless placed directory, mounted as ${dir:config}
- a state directory placed at the assignment root carries route.json
- image pinned to the digest ace runs today (v4.53.10-ls269)
- the sidecar reaches ombi on the machine port the mesh assigns
(${port:3579}) rather than assuming 3579 is free
- the sidecar no longer mounts ombi's data directory: MESH_OMBI_CONFIG_DIR
is read by no code, and the mount exposed the databases for nothing
Verified: catalogue tests (MESH_CATALOGUE set, 6 pass, none skipped); the
pinned image starts as PUID 1000 in a 0700 dir and answers /api/v1/Status
200; data owned 1001:2000 (ace's media ids) under a 1000:1000 dir is
re-owned by the image's init and serves 200; a minted ApiKey is refused
(401) - the api-key secret must be accepted from ombi's own settings.
The route binding still named /var/lib/searxng-module, the directory the
previous commit placed elsewhere — the host would have written it into a
directory nothing declares. Same shape as gitea and nextcloud.
The module ran searxng on the image's built-in settings, which serve html
only — so the module's own search tool (format=json) was refused by the
software it fronts. And there was no way to configure it per machine: the
only file settings reach was the sidecar's.
settings.yml is now the module's one mergeable file (JSON is YAML): generic
defaults in the manifest (json format on, limiter and image proxy off,
valkey wired), and whatever differs per machine — base_url, method,
autocomplete, suspended times — set as the assignment's settings. The
secret key is filled on the machine through ${secret:secret}, so the
secrets-in-environment exception and the env file go. Directories are
placed. Image pinned to 2026.9.20, what ace runs today (the old pin was
older, 2026.9.1).
The sidecar's config.json is no longer mergeable: settings merge into every
mergeable file of a module, and the sidecar would have received searxng's
keys. It only ever read an optional url, which its env already carries.
Verified on ace: the pinned image serves html and json from a read-only,
root-owned 0600 JSON settings.yml.
novox/hq ADR 0147, issue 129. Every internal HTTPS name fails verification
on every machine: the certificates are genuine and nothing on a machine has
ever been told what issued them. The proxy's fetch answers for the proxy and
for nothing else — a browser, git over HTTPS and every module calling another
by an internal name read the machine's own trust store.
The module requires internal-acme-ca, fetches the root over the mesh's own
network (no prior trust to have; that is what this establishes), installs it
among the machine's anchors and refreshes the extracted bundles. Being
unassigned stops the unit, and stopping it takes the anchor away and
refreshes them again.
Arch's layout is named out loud: a machine that keeps anchors elsewhere fails
visibly rather than writing a file nothing reads.
What was in the catalogue was the first thing I built, not the thing ADR 0146
describes. It dialled raw ports on machine addresses from one hosting form and
emitted nothing, so findings would have sat in a file on the machine — the exact
thing issue 145 is about. It was never registered, never assigned, and never ran.
0146 says names per hosting form, fetched over TLS with the certificate verified,
and machines discovered over the bus. That shares nothing with this but the word
checker, so it goes rather than being bent into shape. Recorded as work to be
analysed and built deliberately.
Connectivity is checked by hand in the meantime, against the services the mesh
already runs.
The mesh asserts three things are callable (ADR 0144) — what runs on the same
machine, another machine's service exposed to the private network, and another
machine's service exposed publicly — and has never checked any of them. The first
was broken for eleven hours while the mesh reported every machine healthy.
This runs on every machine, on the cadence the mesh already has, in its own
container: the same position every other module calls from. Not the host and not
the control plane, both of which reach these addresses by paths no ordinary caller
uses and would have passed throughout that outage.
**Its probe is its own endpoint, and that is the point.** Declared reachable over
the private network like any other service, so it is admitted by exactly the rule
that governs every internally-exposed service and fails when that rule is wrong.
The tempting target is a service every machine has, and those are the ones never
closed — ssh above all — which would have passed while the thing that actually
broke was a service exposed to the private network.
It resolves before it dials and says which failed, because a name that does not
resolve and a port that does not answer have different owners. One failure is not
a fault: a machine rebooting is ordinary, so a path is broken after consecutive
runs and the count travels with the result. It reports and repairs nothing.
novox/hq ADR 0145. Eight tests; the consecutive-failure logic proved by reverting
it once. Not yet registered or assigned.
75 endpoints across 50 modules, named from what each one is for rather than by a
rule: mail's seven protocol ports are smtp, imaps, submission and the rest; unifi's
nine are inform, stun, discovery, the two portal ports and syslog; minio's two are
s3 and console; the resolver's two are dns-udp and dns-tcp.
And 35 route contributions name the endpoint they serve instead of repeating its
port. A route and a listen both carried a port and nothing said they were the same
thing; now one of them does. gitea's path-level deny rule names neither, because it
is a rule about a name rather than an endpoint.
novox/hq ADR 0138. The words shipped a release ahead in mesh-controller #138 and
#139, and the control plane running today is built from that merge — checked before
this was written, because an unknown manifest key is refused and a catalogue using
one against an older control plane would stop resolving.
umami declared its port reachable from anywhere, reasoning that the collection
endpoint tracked browsers POST to must be public. That is true of the name and
not of the port: both its surfaces are served through the proxy by name, so the
port is how the proxy reaches it and nothing else (ADR 0045).
Measured, which is how this was found: with the port open to the internet, the
dashboard's login page was served over plain HTTP directly on the machine's port,
bypassing every rule the proxy applies by path. The route stays exactly as it was,
so the collection endpoint keeps working.
It declared udp/53 only. The daemon listens on tcp/53 as well, and a resolver is
asked over tcp whenever an answer will not fit in a datagram — so on every
converged machine that port is closed while the service reports itself healthy
and the manifest reads as though the resolver were fully declared.
The same fault as issue 136 in miniature: the declaration covers part of what the
service does, and the gap is silent because nothing compares the two.
The module said the service must be running and nothing about boot, so the
machine's own way back in was enabled only because something before the mesh
had enabled it. All four machines happen to be enabled today; none of them is
enabled because the mesh says so, and a machine adopted tomorrow would run ssh
until its first reboot.
Not `state: running` alone for the same reason the module exists: this is the
one daemon whose absence cannot be fixed remotely.
jail.local named ufw as the ban action. Two machines on this mesh have no ufw,
and fail2ban does not check: it starts, the jail reads the log, counts the
attempts, runs the ban command, gets 127 -- 'ufw: command not found' -- and
logs an error nobody reads. The service is active, the mesh reports the module
applied, and the machine is not protected. Proven by banning a documentation
address on such a machine today.
The replacement is this module's own dualchain action, already used by the
recidive jail on all four machines, so it is not a new dependency. It bans in
DOCKER-USER as well as INPUT, which ufw's action did not, and it bans all
ports, which ufw's action did.
The recidive jail bans whoever keeps coming back by reading fail2ban's own
log, and fail2ban checks every jail's log file while it configures itself --
before it has created that log. On a machine where the file is not there
already, no jail is found for recidive, configuration fails, and the whole
service refuses to start, taking the sshd jail with it. Two machines assigned
this module today came up failed for exactly that reason; the two where it
worked had a log from years of the service running.
Declared create-once: the mesh puts an empty file there when it is absent and
never touches it again, because what grows in it is fail2ban's, and the
logrotate file this module already ships is what keeps it small.
This also reverts the previous two commits' fail2ban.local. It declared a
logtarget that the package already sets to the same path on every machine
here -- pacman reports the config pristine -- so it fixed nothing and said
something untrue about why.
The file that says where fail2ban logs was not in restart-on, so a change to
it would sit on disk with the running service unaware of it -- the same shape
as any other jail file this module already restarts for.
The recidive jail reads /var/log/fail2ban.log and this module ships the
logrotate file for it, but nothing ever told fail2ban to write there. Where
the package default stands, fail2ban logs to the journal, the recidive jail
finds no log file, and the whole service refuses to start -- taking the sshd
jail with it. Two machines assigned this module today came up failed; the two
where it worked had /etc/fail2ban/fail2ban.conf edited by hand, which a
package upgrade would have undone.
Declared in fail2ban.local, because fail2ban.conf belongs to the package.
It asks what it missed on every start and the answer never arrived: the control plane replayed each
build it held as a module's event from a module called "control-plane", which does not exist, so its
own account refused the publish and the graph kept the gap. The control plane now states those under
the seat it holds (novox/hq ADR 0134, mesh-controller #129), so this consumes that too — one handler,
because what a build means for the graph is the same whether the build machine says it as it happens
or the mesh says what it already held.
It brought its schema up inside its runtime, on every start. That made a schema it could not reach a
crash loop rather than a stop, with the module graph keeping a gap and nothing saying so — which is
how a whole morning's builds went unrecorded. The mesh now prepares this module's state before it
starts this version and does not start it if that failed (novox/hq ADR 0135): the work moves to an
entrypoint the image names in MESH_PREPARE, beside the entrypoints it already names.
The reason it was at start — that a step blocking the apply would block the very apply bringing the
overlay up — stopped being true when a step's failure became its module's business rather than the
machine's (ADR 0136).
Its runtime announces that it has just started and may have missed builds — the event the control
plane follows to replay them — and its manifest did not declare it. A module's authority on the bus
is derived from what it declares, so the publish was refused and the runtime died on start, in a
loop, with the mesh's graph never catching up.
Every module built from a repository was rebuilt for a change to any of them: one merge in this
repository meant twenty-six builds, which is what exhausted a public registry's pull limit. The
forge lists the files a merge changed and the event carries them, from the watcher and from the
merge tool alike; a merge that changed more files than were asked for says so, and the mesh then
treats the whole repository as changed rather than guessing.
An old merge past the first page of the forge's listing surfaced as newer pull requests were
updated, and was announced as if it had just happened; the mesh then rebuilt everything built from
that repository, once per old merge. The moment the watching began is kept with the record, and
only a merge made since is announced.
A trigger that fires silently is indistinguishable from one that did not fire (novox/hq issue
131); the announcement is now one line an operator can read.
/user/repos lists what the token's user owns, which for the mesh's administrator is nothing — so
the forge module watched an empty list and never announced a merge. It reads the forge's whole
view through the search endpoint, every page.
The merge tool emitted at the instant it acted; a merge made in the forge's own
pages or over its API emitted nothing, and the mesh went on believing every module
current with its source (novox/hq 04-ISSUES/131). Merged pull requests are now
watched the way repositories are: what the forge holds, asked for on a tick,
announced once, with the merge commit and the clone URL a build needs. What has
been announced is kept beside the module's state, so a restart does not announce
the whole history again, and a first tick with no record announces nothing.
The mesh's broker certificate is the operator's, kept outside any module and
mounted read-only by whatever serves the bus; the controller declares that
access and now so does nats, the same way. A mount nothing declares is refused
at registration (ADR 0030), which is how this was found.