Commit Graph
100 Commits
Author SHA1 Message Date
jschoubben 5b70ffd78f mosquitto, influxdb: the provisioner sees its grants where the mesh writes them
The received grants file names each pair secret by its HOST path. A sidecar
that mounts the placed grants directory at another path inside its
container sees mesh.json and not the secrets beside it — mosquitto's
provisioner reported ENOENT for a file that was there, and nodered's mqtt
step failed against a login that was never created. Mounted at its own
path now, as postgres does.
2026-09-30 17:19:33 +02:00
jschoubben 0020fa17cd Merge pull request 'mesh-console: the mesh's tools on the machine a person sits at' (#181) from feat/the-console into main
Reviewed-on: #181
2026-09-30 14:47:22 +00:00
jschoubben 22c6032144 Merge branch 'main' into feat/the-console 2026-09-30 14:47:11 +00:00
jschoubben 769f0724ca mesh-console: the mesh's tools on the machine a person sits at
The tool runtime's own client, mesh serve, started by the mesh on the credential it sealed to the
machine (novox/hq ADR 0152, design 34): invokes every tool, listens from the machine only, holds no
state. Checked with module check before it was ever registered (hq issue 148).
2026-09-30 16:20:06 +02:00
jschoubben 50ae89e718 influxdb: its admin password and operator token are its own secrets, so an existing instance's can be accepted
They came from `requires: secret`, minted by the vault — right for a fresh
setup (the image's INIT_* variables read them once), wrong for an instance
that already exists: setup is skipped, the minted values match nothing,
and the provisioner holds a token the server never issued (issue 100).
InfluxDB will not take a chosen token value, so the operator token must be
accepted from the instance (`secret accept ace influxdb admin-token`); the
password can be either. As own secrets both are minted for a fresh install
exactly as before, and accepted where the data already knows them.
Found migrating ace's influxdb.
2026-09-30 16:18:59 +02:00
jschoubben a2b9a9a411 Merge pull request 'dnsmasq: pass the DNSSEC bit down from the validating upstreams' (#179) from fix/resolver-passes-the-dnssec-bit into main 2026-09-30 13:17:38 +00:00
jschoubben 9523105df4 dnsmasq: pass the DNSSEC bit down from the validating upstreams
A program that checks its resolver validates — Mailu's admin does, at
start — could not use the mesh's resolver, and the one it ships instead
knows no mesh name (novox/hq issue 171). proxy-dnssec copies the AD bit
from 1.1.1.1 and 8.8.8.8, both of which validate.
2026-09-30 15:17:34 +02:00
jschoubben 4449f44cf1 Merge pull request 'mailu: admin asks the machine's resolver, not Mailu's own' (#178) from fix/mailu-admin-asks-the-machines-resolver into main 2026-09-30 13:13:29 +00:00
jschoubben 47f6e7d78b mailu: admin asks the machine's resolver, not Mailu's own
admin is the one container here that reaches something by a mesh name:
the database, bound in database.env. A mesh name is answered by the
machine's resolver, which the runtime hands every container that does
not name its own (novox/hq ADR 0148); Mailu's unbound knows no mesh
name, and since the mesh stopped copying names into containers admin
could not find its database. The others keep unbound — rspamd needs a
validating resolver for DNSBL lookups and asks for no mesh name.
2026-09-30 15:13:25 +02:00
jschoubben 7ab522ba31 Merge pull request 'dnsmasq: answer a container's query, which arrives on the runtime's bridge' (#176) from fix/110-the-resolver-answers-a-container into main 2026-09-30 12:31:30 +00:00
jschoubben 204bfbbaf9 dnsmasq: answer a container's query, which arrives on the runtime's bridge
dnsmasq admits a query by the interface it arrives on when told interface=,
and by the address it is sent to when told listen-address=. A container's
query is sent to the machine's private address but arrives on docker0, so
interface=mesh0 dropped it silently on every machine (novox/hq issue 110).
Name the address, not the interface.
2026-09-30 14:31:26 +02:00
jschoubben e0c09f46b6 Merge pull request 'dnsmasq: the runtime is reloaded when its dns file changes, and may then be restarted safely' (#175) from fix/110-the-runtime-reads-its-dns into main 2026-09-30 12:22:10 +00:00
jschoubben e0faf012be dnsmasq: the runtime is reloaded when its dns file changes, and may then be restarted safely
novox/hq 04-ISSUES/110. The module writes the runtime's `dns` key and
deliberately ordered no restart, because a restart stops every container.
It also ordered no reload, and the key holds only for containers created
after the runtime next starts. On two of four machines the runtime
predated the file — one since August — so every container there was
handed a public resolver and no mesh name resolved, while everything read
as fine. The third machine works only because its runtime happened to
restart later.

The file now also sets live-restore, which the runtime reads on a reload,
and the module declares the runtime reloaded when the file changes (ADR
0102: reload, don't restart). A reload still does not make `dns` take
effect; what it does is make the one restart that key needs keep every
container running. That restart stays the operator's, once per machine,
and is harmless from the second time on.

The mesh restarts nothing here. Undeclared, the runtime's unit goes back
to the state it was found in (ADR 0118).
2026-09-30 14:22:03 +02:00
jschoubben 4ec2ae1f7f postgres: the provisioner dials the port the mesh gave the store
MESH_PROVISION_POSTGRES named 127.0.0.1:5432 literally; the seat twin
(${seat:mesh-store:5432}) corrected it only on the machine holding the
mesh-store seat. On any other machine — ace, where the module provides
postgres-database without the seat — the twin is empty and the sidecar
would have dialled whatever else holds 5432 (HAL's postgres). ${port:5432}
is the machine port on every node; on novox the twin still says the same
6852, so nothing moves there.
2026-09-30 14:01:53 +02:00
jschoubben 9d716ed875 jackett: provide its Torznab API as jackett-api
sonarr, radarr, lidarr and bookshelf reached jackett as http://jackett:9117 (a HAL
container name) or https://indexers.zurag.be (its public route), typed into each app by
hand. The mesh has neither: an app now requires jackett-api and its downloads step writes
the bound address into the app.

Serves scheme, port and url-base; `at` and the machine port come from the binding. Mesh
scope, like sonarr-api: an indexer proxy shares no files with its consumers.

The pair credential is jackett's one API key. The mesh cannot mint it, so the operator
accepts it per consumer pair (ADR 0092), as #156 does for sonarr-api; a consumer's step
refuses a minted value and names the accept.
2026-09-30 13:05:36 +02:00
jschoubben 3c7aafdc21 nodered: its MQTT broker comes from the mesh
Node-RED's one broker node pointed at zurag.be:1884, where nothing listens. nodered now requires
mqtt-topic (asking for every topic: flows follow the devices' own) and a run-once `mqtt` step —
declared last, restarted when the binding, credential or settings change — points the mesh's broker
nodes at the bound broker through Node-RED's admin API with the module's api-token: the node the
step makes itself when none is named, or the ones an assignment names in `mqtt.brokers`. Only host,
port, TLS and the login change; the broker is asked first whether it takes the login; the deploy is
against the revision read ("nodes", so only that node restarts) and a digest makes a rerun a no-op.
A broker node nobody named is never touched. settings.js keeps `mqtt` and `topics` out of Node-RED.
2026-09-30 13:01:14 +02:00
jschoubben 1080f45012 mosquitto: a consumer's grant is the topics it asks for, and its binding says the port
mqtt-topic served nothing: with two listens the mesh could not say which port a consumer dials, so a
consumer had to type 1883 into its config. It now serves the MQTT listener's port (the machine's,
once assigned) and the scheme, so `${bound:mqtt-topic:port}` fills.

The provisioner confined every consumer to `<as>/#`, which leaves nothing for the consumers the
broker exists for: Home Assistant discovers under homeassistant/# and tasmota/discovery/#, and
Node-RED's flows follow the devices' own topics. A consumer now contributes `topics` (MQTT topic
filters) to its mqtt-topic requirement and is granted exactly those; with none, its own subtree as
before. Settings merge into contributions, so an operator narrows a grant per assignment. The role
is brought to exactly the wanted ACLs (stale ones removed), `holds` checks the ACLs too, and an
invalid list is refused, never quietly narrowed. Only the role named for the consumer is touched:
a client carried from the predecessor's password file keeps its own.
2026-09-30 13:01:14 +02:00
jschoubben 323ef9ec7e influxdb: provide influxdb-api, one mesh-made v1 credential per consumer
grafana's data source and Node-RED's influxdb nodes reached ace's
InfluxDB by a LAN IP or a public name nobody routes, with a credential
somebody made by hand. Now a consumer requires influxdb-api and is told
where it is, which org and default bucket it serves, and signs in with
the password the mesh minted for the pair.

The credential is a v1-compatibility authorization, made per grant by
the new provisioner: InfluxDB 2.x generates API tokens itself and
ignores one the caller sends, so a v2 token could only be accepted by
hand per pair; a v1 authorization takes a caller-chosen password (8-72
characters, the mesh mints 40) and reads/writes every bucket as a
database of its name over InfluxQL and line protocol. A consumer
contributes `access` (read, write, read-write) and, for writing, the
buckets; a missing bucket is made and never deleted. Only
authorizations named mesh_* and marked [mesh] are ever changed or
removed; anything else of that name is refused and left alone.

The org and default bucket are served facts the assignment's settings
set, reaching both the consumers and the provisioner's config.json.
2026-09-30 12:55:38 +02:00
jschoubben 8f459c7023 letta: placed state, a route, accepted keys, and a runtime that can log in
The module named /var/lib/letta, a layout no definition may carry (ADR
0112); state is now a placed directory holding the bindings and secrets.

letta had no route, while the server it replaces is reached by its public
name (a workflow calls it there). It now contributes one for its `web`
endpoint; reach is the assignment's.

The server needs an OpenAI key for agents on OpenAI models, and nothing
gave it one: `openai-api-key` is an own-secret, accepted from the
operator (a key someone chose, not one the mesh can mint). The
server-password is accepted the same way where a server already has
clients. Both, and the database password inside LETTA_PG_URI, stay in the
server's environment: letta 0.6.x reads settings from the environment
only, and its startup.sh starts an embedded PostgreSQL unless
LETTA_PG_URI is set - the declared reason now says so.

The runtime's tools never authenticated: the client sent only a Bearer
token, and 0.6.x's --secure mode checks X-BARE-PASSWORD ("password <it>")
and answers 401 otherwise. The client now sends both. Its password comes
from the runtime config file (the key client.ts reads first) instead of an
env-file, so the runtime container no longer carries a secret in its
environment.

Image: the same 0.6.8 image, now pinned by the index digest ace runs
rather than its amd64 manifest.

Verified: catalogue tests with MESH_CATALOGUE set; tsc -p tsconfig.json in
the mesh-tools build image. Throwaway containers: a fresh 0.6.8 with its
embedded PG and two blocks made through the API; stopped, copied, dumped
from the copy; restored (schema letta + vector pre-made by the superuser,
--no-owner --role, search_path set on the database as the original had
it) into a grant-shaped database on the postgres module's pgvector image
(PG17); started in this shape: alembic finds nothing to do, both blocks
are there, a wrong password is refused with 401, and the patched client
lists agents. Test containers and data removed.
2026-09-30 12:22:12 +02:00
jschoubben ac8556c590 baserow: placed directories, the database password from a file, and the build ace runs
The module named /var/lib/baserow and /services/baserow/data, a layout no
definition may carry (ADR 0112). State and data are now placed directories;
bindings and the grant's secret live in the placed state.

The all-in-one image's entrypoint honours DATABASE_PASSWORD_FILE (file_env
in /baserow.sh), so the grant's password is mounted rather than put in an
env-file, and "secrets-in-environment" is gone (ADR 0086). SECRET_KEY is no
longer minted: the image keeps it, and its JWT signing key, in the data
directory (.secret, .jwt_signing_key) and imports them on start, so a moved
data directory carries the keys its sessions and tokens were made with.
DISABLE_EMBEDDED_PSQL makes a missing grant fail loudly instead of starting
an empty embedded database.

BASEROW_PUBLIC_URL was http://localhost. Baserow answers only the host of
that URL - any other Host is looked up as a published builder site and gets
404, /api/_health/ included - so it is now https://${bound:route:name}
(depends on mesh-controller #149).

The runtime's tools could never have worked: its config was "{}", and the
client's Host override was silently dropped by Node's fetch, so calls by
container name would 404 even with credentials. The client now uses
node:http (which sends the Host it is given, with a Content-Length -
Baserow reads a chunked body as empty) and re-authenticates once when a
cached JWT is refused (access tokens last minutes, the runtime weeks). The
password is the accepted `admin` secret; the email is an assignment
setting merged into the same file, the host is the route's name.

Image pinned to the develop-latest build ace runs today (Baserow 2.3.4,
built 2026-09-18). The old pin (built 2026-09-04) is older than ace's data.

Verified: catalogue tests with MESH_CATALOGUE set; tsc -p tsconfig.json in
the mesh-tools build image. In throwaway containers of the pinned image: a
fresh embedded-PG instance with a user, workspace and 5-row table; stopped,
copied, dumped from the copy (start-only-db); restored with --no-owner
--role into a grant-shaped database on the pgvector image the postgres
module pins (PG17); started with this shape (root 0600 password file,
embedded PSQL disabled, copied data dir without postgres/): health 200,
the user logs in, the 5 rows are there, SECRET_KEY and the JWT key are
imported from the data dir. The patched client lists applications and rows
through the container name with the public Host, and recovers from a
refused token. Test containers and data removed.
2026-09-30 12:16:04 +02:00
jschoubben 62cecc8a2b supabase: the self-hosted stack as one module, its secrets rendered into files
ace runs Supabase under HAL as upstream's 13-container compose: a 2.2 GB
database (1.8 GB of it the dormant `novox` schema, 5.8 M rows in its largest
table), Kong at supabase.zurag.be, the pooler on 5433/6543. This is that stack
as a catalogue module, same images (the digests ace runs), same container
names so an assignment holds the running ones and a take replaces them.

The database stays inside the module. Supabase is a Postgres distribution: its
own image with pgsodium, pg_graphql, pg_net, vault and timescale preloaded, a
superuser (supabase_admin), a dozen reserved roles and a second database
(_supabase). A postgres-database grant - one database, one unprivileged role -
cannot hold it, so ace's data directory moves as a copy, not a dump/restore.

What HAL did by shell and environment the mesh now renders as files:
- kong.yml carries the anon/service keys and the dashboard login (owned by
  kong's uid 100), instead of an entrypoint that eval'd the environment;
- GoTrue reads a dotenv file (auth -c), PostgREST a config file, Vector its yml
  with the Logflare key in it, the database POSTGRES_PASSWORD_FILE and a
  jwt.sql rendered with the secret (owned by postgres, uid 105). None of these
  five containers has a secret in its environment.
- realtime, storage, meta, functions, analytics, studio and supavisor read
  their credentials from the environment only; each declares
  secrets-in-environment with the reason (ADR 0086).
- Upstreams are container names (supabase-db, supabase-kong, ...) instead of
  compose service names, which the mesh does not have.
- SITE_URL / API_EXTERNAL_URL / SUPABASE_PUBLIC_URL are
  https://${bound:route:name} (mesh-controller #149); on ace HAL rendered them
  as "https://supabase." - broken today.
- Vector reads the docker socket, as upstream does, but now includes only this
  module's containers instead of every container's logs on the machine.

Three things compose did that a declaration cannot, done as steps: a run-once
seed copies the image's /etc/postgresql-custom into the placed config
directory with cp -n (a named volume did that implicitly; never overwrites the
pgsodium root key), and two run-once gates wait for the database and for
Logflare, which compose expressed as depends_on: service_healthy.

The pooler bootstrap (pooler.exs) takes the tenant id and pool sizes from
settings.json, the module's one merge:json file, so ace keeps its tenant
"zurag"; and it repoints an existing tenant whose database host is not
supabase-db - HAL created ace's with host "db", which no longer resolves.

Secrets (vault, requires "secret"): postgres, jwt, anon-key,
service-role-key, dashboard-user, dashboard, logflare, pooler-vault,
key-base-a + key-base-b (concatenated: Phoenix wants 64+ bytes, a minted
secret is 40), openai. Three cannot be minted on any machine: anon-key and
service-role-key are JWTs signed with jwt, and pooler-vault must be exactly
32 bytes (AES-256-GCM, found in the bed). They are accepted. On ace every
secret the data already knows is accepted (all but key-base-a/b).

Not carried: Kong's 8443 and Logflare's 4000 on all interfaces (nothing
outside the module uses them); realtime's DB_ENC_KEY stays upstream's constant
(realtime deletes and re-seeds that tenant from its environment every start,
and the key must be exactly 16 bytes).

Verified: catalogue tests with MESH_CATALOGUE pointed here on mesh-controller
main and #149 (on main the render is refused for "name", never written
empty). The #149 resolution with stub providers, turned into a throwaway
stack of all 13 pinned digests with dummy secrets and the rendered files at
their owners and modes: the database initialised through the rendered scripts
(jwt setting applied, _analytics/_supavisor created, roles' password from
POSTGRES_PASSWORD_FILE); through Kong: REST 200 with the anon key and 401
without, auth health and settings 200, storage buckets 200, GraphQL 200, pg-meta
200, an edge function 200, Studio 401 without and 200 with the dashboard login,
realtime tenant health 200; the pooler in session and transaction mode as
postgres.zurag; a tenant set to host "db" was repointed to supabase-db by the
bootstrap and connections worked.
2026-09-30 12:12:27 +02:00
jschoubben 157fab5dc8 matrix: Conduit and Element, told their own name by the route
ace runs a Conduit homeserver (matrix.zurag.be, 5.4 GB of RocksDB, federating)
and Element Web under HAL, configured by environment with the domain
templated in, and Element's config.json carrying matrix.zurag.be literally.

A homeserver's server_name is its permanent identity - every user id, room id
and signature in the database carries it - and it is the name the module is
served under. So it comes from ${bound:route:name-homeserver} (mesh-controller
#149), rendered into a conduit.toml the container reads through CONDUIT_CONFIG,
and into Element's config.json (base_url, default_server_name, the room
directory). Without #149 the render is refused ("name-homeserver"), never
written empty. Two routes, one per endpoint: homeserver (label matrix, 6167)
and element (label element, 80).

Federation needs no 8448: Conduit answers /.well-known/matrix/server with
<name>:443, so peers federate through the route. HAL published 8448 on all
interfaces, but the router never forwarded it; checked from outside, the
well-known, federation version and client versions all answer on 443.

Registration defaults to off. HAL ran with CONDUIT_ALLOW_REGISTRATION=true,
which on ace means anyone on the internet can create an account with the
dummy flow (seen: /register offers m.login.dummy) - on 0.10.13, whose
successor 0.10.14 fixes an account-takeover by any local user. Existing
accounts are unaffected; the operator decides whether to reopen it.

Element's config.json is the one merge:json file (it tolerates `endpoints`),
so a machine can add keys. HAL's map_style_url is not carried: it embedded
a map-tile API key, which belongs in an assignment if wanted.

Images are the digests ace runs (Conduit 0.10.13, Element 1.12.28).

Verified: catalogue tests with MESH_CATALOGUE pointed here on mesh-controller
main and #149; a resolution on #149 renders both names (matrix.zurag.be,
element.zurag.be) into both files and both contributions; throwaway
containers of both pinned digests with the rendered files (root 0644, :ro):
client versions 200, well-known says matrix.zurag.be:443, register refused
M_FORBIDDEN, Element 200 serving the rendered config.json.
2026-09-30 11:59:56 +02:00
jschoubben 1247b8c27e redis: place its directories and run the build in use
The manifest stated /services/redis/data and /var/lib/redis-module - novox's
old layout, paths no definition may carry (ADR 0112). State is now the
assignment's root, grants and data are placed, and the config file, the
secret file, receives and grants all name them as ${dir:...}. Paths inside
the sidecar are its own view and are unchanged.

The data directory and config are owned 999:1000: the image's redis user is
uid 999 in gid 1000 (checked in both builds), which is who owns ace's data
today; 999:999 named a group the image does not use.

Image pinned to the 7.4.11-alpine build ace runs (2026-09-17); the old pin was
the same version, built in August. Older-than-running is never the pin.

Nothing is assigned it anywhere today, so no machine changes.

Verified: catalogue tests pass with MESH_CATALOGUE on this tree; the
declaration composes for ace with every path under /var/lib/redis. The pinned
image ran as a throwaway with a 0600 999:1000 config and a 0700 data dir:
unauthenticated PING is refused (NOAUTH), authenticated SET/GET works,
appendonly is on, the server runs as redis.
2026-09-30 11:57:46 +02:00
jschoubben 0fce3ebf5d mssql: place its directories and name the software's port, not a machine's
The manifest stated /var/lib/mssql, its grants and its SA file by path, and
declared it listens on 4848 - the port one machine's predecessor published,
which is an assignment's fact (ADR 0112, 0138). ace is moving its own
SQL Server (80 GB of work databases) onto the mesh, so the module has to be
the same on every machine.

- state is the assignment's root (place "."), grants is placed, and every
  reference (sa.env, the env-file, the SA and grants mounts, receives,
  grants, own-secrets) names them as ${dir:...}.
- the database endpoint listens on 1433, the port SQL Server uses; a machine
  that must keep an older number pins it in its assignment.
- the image pin is unchanged: it is the digest ace runs today (CU27,
  16.0.4295), the same as novox.

novox is untouched: rendered with novox's own setting ({"ports":{"1433":4848}})
through the controller's Declaration and Rules, every resource - paths,
container names, volumes, env-file, the 4848:1433 mapping, owners, modes - is
byte-identical to what main renders; the only difference is the firewall
rule's comment text (still port 4848, from the mesh).

Verified: catalogue tests pass with MESH_CATALOGUE on this tree. The pinned
image ran as a throwaway on a 0700 10001:0 data dir with a root-owned 0600
env-file (dummy SA), answered sqlcmd as sa; a scratch database stopped,
copied with cp -a, checksummed and started on the copy kept its rows and
CHECKSUM_AGG.
2026-09-30 11:57:45 +02:00
jschoubben c1a65e2354 nodered: the sidecar dials the port it was given
The sidecar runs on the host network and dialled 127.0.0.1:1880, the
software's port; the mesh publishes nodered on a machine port it assigns,
so the tools reached whatever else holds 1880, or nothing (hq 088).
2026-09-29 23:50:11 +02:00
jschoubben fe0ed3b74e influxdb: place its directories, hand secrets over as files, name its UI
The manifest named /services/influxdb and /var/lib/influxdb-module — one
machine's paths — and passed the admin password and token through the
environment. ace is moving its 2022 instance onto the mesh, so the module
has to be what it is on any machine.

- data, config and state are placed directories; the data keeps 1000:1000,
  the image's influxdb user, which is who owns ace's data today.
- the init secrets reach the image through its own
  DOCKER_INFLUXDB_INIT_{PASSWORD,ADMIN_TOKEN}_FILE; the vault's files are
  mounted read-only. secrets-in-environment is gone.
- the sidecar reads its token from the same file (MESH_INFLUXDB_TOKEN_FILE,
  added to client.ts) and reaches the server at its assigned machine port
  (${port:8086}) instead of assuming 8086. The unused config-dir mount,
  which held the CLI's copy of the admin token, is dropped.
- the api endpoint contributes a route: the web UI is how people use it,
  and reach is the assignment's to say.

Verified: catalogue tests pass with MESH_CATALOGUE pointed at this tree.
The pinned 2.9.1 image, run on a scratch copy of ace's 2.4.0 data, opens
it, runs its metadata migrations (backing up the pre-upgrade bolt/sqlite)
and hashes the two stored tokens; /health passes. A fresh setup through
the _FILE variables, with dummy secrets as root-owned 0600 files, accepts
the token (200 on /api/v2/buckets) and the password (204 on /signin).
client.ts typechecks strict and reads the token file, tolerating the
endpoints key in its config.
2026-09-29 23:42:18 +02:00
jschoubben 75eee9d4a0 jackett: its config dir is placed, and its tools find their own key
The manifest named /services/jackett/config and /var/lib/mesh/jackett/config.json — host
paths ADR 0112 takes out of definitions. The config dir is now a pathless directory
(${dir:config}) and the runtime's config and route binding live in a placed state dir, as
searxng does.

The image is pinned to v0.24.2627-ls34, the digest ace runs today; the old pin
(v0.24.2517-ls16) was older than the running version.

The runtime reached jackett at a fixed 127.0.0.1:9117; it now uses ${port:9117}, the
machine port the mesh actually assigned.

The tools never loaded: the client needed an API key nobody set. Like sonarr/radarr read
config.xml, it now reads APIKey from Jackett's own ServerConfig.json (the config dir is
already mounted read-only), so no secret goes into an assignment. jackett_indexers called
/api/v2.0/indexers, which is the web UI's endpoint and answers an API key with a redirect;
it now reads the Torznab t=indexers feed, and treats Torznab's 200-with-<error> as a
failure.

Verified: catalogue key tests (MESH_CATALOGUE set, not skipped); a throwaway container of
the pinned image on a fresh 0700 1000:1000 config dir serves its UI; the client discovers
the key from the generated ServerConfig.json, lists 617 indexers through the Torznab feed,
searches via /results, and a wrong key is refused; client.ts typechecks under --strict.
2026-09-29 23:41:21 +02:00
jschoubben 0c91e08bad nodered: its settings are files the mesh writes, and its editor is locked
The catalogue ran the image's defaults: no adminAuth, so a routed Node-RED
editor (which runs arbitrary code) was open to anyone who reached it, and
the module's own tools had no token to present to an install that was locked.

- settings.js (fixed, 0600, uid 1000) carries adminAuth: user admin checked
  against the admin secret -- a minted password, or the bcrypt hash an
  existing install held (accepted), so current logins keep working -- and a
  static bearer token (api-token) the sidecar presents. It loads settings.json
  beside it, the one mergeable file; endpoints is dropped there, and an
  optional timeZone sets process.env.TZ (assignments cannot set env).
- The sidecar's runtime config is no longer merged; it carries the token.
- Directories are placed (state, data), the route binds into state.
- Image pinned to 5.0.7 (a649dd71), what ace runs; the old pin was 5.0.6.
- deployFlows asks for API v2: v1 answers 204 with no body, which the client
  tried to parse as JSON.

Verified: catalogue tests pass against this tree. A throwaway 5.0.7 container
started with the generated files: anonymous /flows 401, bearer api-token 200,
bad token 401, password grant 200/403 with a minted password and with a
bcrypt-hash-accepted one; endpoints and timeZone do not reach /settings;
timeZone Europe/Brussels overrides TZ=Etc/UTC; v1 deploy 204, v2 deploy
answers {rev}.
2026-09-29 23:41:18 +02:00
jschoubben 37c212d5b4 unifi: placed data, mesh-assigned ports, an https route, and its password as a file
The module stated /services/unifi/data and fixed machine ports (8443:8443 and
eight more) — one installation's layout and numbers, which a definition may not
carry (ADR 0038, 0112). Data is now a placed directory (${dir:data}, 1000:1000,
0700), the container publishes its own ports and the mesh assigns the machine
side; an assignment pins them where devices already know them. The L2 endpoint
names the port the software uses (1900), not the one a machine published it on.

The sidecar dialled https://127.0.0.1:8443, true only while the machine port
equals the container's; it now asks for ${port:8443}. Its controller password
was a setting (plaintext in the mesh DB); it is now an own-secret written into
the one mergeable file (ADR 0086). The username stays a setting.

The web UI is contributed as a route to the "web" endpoint over https with
insecure upstream (the controller's own self-signed tls), as mailu's web-tls —
what HAL's hand-written traefik file for unifi does today.

Image pinned to the manifest list ace runs (8.0.24-ls221); the old pin was its
amd64 child, so the image is unchanged.

Verified: catalogue tests pass with MESH_CATALOGUE set; a throwaway container
of the pinned image on a fresh 1000:1000/0700 data dir answers /status (8.0.24,
up) and /inform; the sidecar client built from a config.json carrying site,
password, username and an endpoints key reaches it and is refused only on the
dummy credentials.
2026-09-29 23:39:50 +02:00
jschoubben 718fb12ef7 icecast: its passwords are a file the mesh writes, not the image's environment
The image seds ICECAST_*_PASSWORD from the environment into /etc/icecast.xml;
ADR 0086 wants secrets as files. icecast starts as root, reads its config, then
drops to uid 100, so a root-owned 0600 icecast.xml rendered with ${secret:...}
and mounted read-only works and the entrypoint's seds never fire (no env set).
The "secrets-in-environment" exemption and server.env are gone.

Also: directories are placed (state, logs owned 100:101 so the image's VOLUME
/var/log/icecast is not an anonymous volume per container, as HAL learned);
the server and sidecar share a module network, so the sidecar reaches
http://icecast:8000 instead of assuming machine port 8000 on the host; the
stream endpoint is routed (label "icecast"), as HAL served it via traefik.
Secrets remain mesh-vault grants (requires secret), now under ${dir:state}.

Verified: catalogue tests with MESH_CATALOGUE pointed at this tree; a
throwaway container of the pinned digest (the one ace runs) with the rendered
file (dummy secrets, root 0600, :ro): runs as icecast, status-json 200,
admin 401 without / 200 with the admin secret, a source PUT with the source
secret mounts, a listener receives it, a wrong source password gets 401, logs
land in the uid-100 directory.
2026-09-29 23:39:36 +02:00
jschoubben a32394ec22 mosquitto: its directories are placed, not stated, and it runs the build in use
The module stated /var/lib/mosquitto-module and /services/mosquitto/data —
novox's layout, a path no definition may carry (ADR 0112). State, grants and
data are now placed directories (${dir:state}, ${dir:grants}, ${dir:data}),
the admin secret lives beside the broker account under the mesh's own state,
and the receives/grants maps follow the grants directory. Paths inside the
sidecar are its own view and are unchanged.

Image pinned to the 2.1.2 build ace's predecessor runs (2026-09-17); the old
pin was the same version, built in June.

Found preparing ace, whose broker carries a password-file user (an IoT switch
and home-assistant). Carrying it is a data step, not a manifest one: the
migration repo has scripts/mosquitto-pwdfile-to-dynsec.py, which moves $7$
PBKDF2 entries into the dynsec store hash-for-hash (tested end to end).
2026-09-29 23:05:41 +02:00
jschoubben 63a255c5cb searxng: bind its route where its state now lives
The route binding still named /var/lib/searxng-module, the directory the
previous commit placed elsewhere — the host would have written it into a
directory nothing declares. Same shape as gitea and nextcloud.
2026-09-29 22:41:56 +02:00
jschoubben 7ad1fbd5c6 searxng: its settings are a file the mesh writes, not the image's defaults
The module ran searxng on the image's built-in settings, which serve html
only — so the module's own search tool (format=json) was refused by the
software it fronts. And there was no way to configure it per machine: the
only file settings reach was the sidecar's.

settings.yml is now the module's one mergeable file (JSON is YAML): generic
defaults in the manifest (json format on, limiter and image proxy off,
valkey wired), and whatever differs per machine — base_url, method,
autocomplete, suspended times — set as the assignment's settings. The
secret key is filled on the machine through ${secret:secret}, so the
secrets-in-environment exception and the env file go. Directories are
placed. Image pinned to 2026.9.20, what ace runs today (the old pin was
older, 2026.9.1).

The sidecar's config.json is no longer mergeable: settings merge into every
mergeable file of a module, and the sidecar would have received searxng's
keys. It only ever read an optional url, which its env already carries.

Verified on ace: the pinned image serves html and json from a read-only,
root-owned 0600 JSON settings.yml.
2026-09-29 22:33:38 +02:00
jschoubben 67f5f4cffd Merge pull request 'ca-trust: a machine trusts the mesh's authority because a module put its root there' (#142) from feat/ca-trust into main 2026-09-29 14:06:36 +00:00
jschoubben 8797335fbc ca-trust: a machine trusts the mesh's authority because a module put its root there
novox/hq ADR 0147, issue 129. Every internal HTTPS name fails verification
on every machine: the certificates are genuine and nothing on a machine has
ever been told what issued them. The proxy's fetch answers for the proxy and
for nothing else — a browser, git over HTTPS and every module calling another
by an internal name read the machine's own trust store.

The module requires internal-acme-ca, fetches the root over the mesh's own
network (no prior trust to have; that is what this establishes), installs it
among the machine's anchors and refreshes the extracted bundles. Being
unassigned stops the unit, and stopping it takes the anchor away and
refreshes them again.

Arch's layout is named out loud: a machine that keeps anchors elsewhere fails
visibly rather than writing a file nothing reads.
2026-09-29 15:07:40 +02:00
jschoubben ebf5ba2d4c Remove the network-checker module: it does not do what was decided
What was in the catalogue was the first thing I built, not the thing ADR 0146
describes. It dialled raw ports on machine addresses from one hosting form and
emitted nothing, so findings would have sat in a file on the machine — the exact
thing issue 145 is about. It was never registered, never assigned, and never ran.

0146 says names per hosting form, fetched over TLS with the certificate verified,
and machines discovered over the bus. That shares nothing with this but the word
checker, so it goes rather than being bent into shape. Recorded as work to be
analysed and built deliberately.

Connectivity is checked by hand in the meantime, against the services the mesh
already runs.
2026-09-29 14:43:25 +02:00
jschoubben 784a5a6514 A network-checker module: dial what the mesh claims, from where the callers are
The mesh asserts three things are callable (ADR 0144) — what runs on the same
machine, another machine's service exposed to the private network, and another
machine's service exposed publicly — and has never checked any of them. The first
was broken for eleven hours while the mesh reported every machine healthy.

This runs on every machine, on the cadence the mesh already has, in its own
container: the same position every other module calls from. Not the host and not
the control plane, both of which reach these addresses by paths no ordinary caller
uses and would have passed throughout that outage.

**Its probe is its own endpoint, and that is the point.** Declared reachable over
the private network like any other service, so it is admitted by exactly the rule
that governs every internally-exposed service and fails when that rule is wrong.
The tempting target is a service every machine has, and those are the ones never
closed — ssh above all — which would have passed while the thing that actually
broke was a service exposed to the private network.

It resolves before it dials and says which failed, because a name that does not
resolve and a port that does not answer have different owners. One failure is not
a fault: a machine rebooting is ordinary, so a path is broken after consecutive
runs and the count travels with the result. It reports and repairs nothing.

novox/hq ADR 0145. Eight tests; the consecutive-failure logic proved by reverting
it once. Not yet registered or assigned.
2026-09-29 13:44:16 +02:00
jschoubben f118344246 Every module names its endpoints, and every route names the one it serves
75 endpoints across 50 modules, named from what each one is for rather than by a
rule: mail's seven protocol ports are smtp, imaps, submission and the rest; unifi's
nine are inform, stun, discovery, the two portal ports and syslog; minio's two are
s3 and console; the resolver's two are dns-udp and dns-tcp.

And 35 route contributions name the endpoint they serve instead of repeating its
port. A route and a listen both carried a port and nothing said they were the same
thing; now one of them does. gitea's path-level deny rule names neither, because it
is a rule about a name rather than an endpoint.

novox/hq ADR 0138. The words shipped a release ahead in mesh-controller #138 and
#139, and the control plane running today is built from that merge — checked before
this was written, because an unknown manifest key is refused and a catalogue using
one against an older control plane would stop resolving.
2026-09-29 11:51:39 +02:00
jschoubben 9eb1265bc8 A routed module listens from the mesh, not from anywhere
umami declared its port reachable from anywhere, reasoning that the collection
endpoint tracked browsers POST to must be public. That is true of the name and
not of the port: both its surfaces are served through the proxy by name, so the
port is how the proxy reaches it and nothing else (ADR 0045).

Measured, which is how this was found: with the port open to the internet, the
dashboard's login page was served over plain HTTP directly on the machine's port,
bypassing every rule the proxy applies by path. The route stays exactly as it was,
so the collection endpoint keeps working.
2026-09-29 02:54:10 +02:00
jschoubben acedc5d9d9 The resolver declares both protocols it answers on
It declared udp/53 only. The daemon listens on tcp/53 as well, and a resolver is
asked over tcp whenever an answer will not fit in a datagram — so on every
converged machine that port is closed while the service reports itself healthy
and the manifest reads as though the resolver were fully declared.

The same fault as issue 136 in miniature: the declaration covers part of what the
service does, and the gap is silent because nothing compares the two.
2026-09-28 23:53:03 +02:00
jschoubben e145e2236c sshd: the daemon it owns starts at boot
The module said the service must be running and nothing about boot, so the
machine's own way back in was enabled only because something before the mesh
had enabled it. All four machines happen to be enabled today; none of them is
enabled because the mesh says so, and a machine adopted tomorrow would run ssh
until its first reboot.

Not `state: running` alone for the same reason the module exists: this is the
one daemon whose absence cannot be fixed remotely.
2026-09-28 21:43:27 +02:00
jschoubben 026421fd6e fail2ban: ban through an action every machine has
jail.local named ufw as the ban action. Two machines on this mesh have no ufw,
and fail2ban does not check: it starts, the jail reads the log, counts the
attempts, runs the ban command, gets 127 -- 'ufw: command not found' -- and
logs an error nobody reads. The service is active, the mesh reports the module
applied, and the machine is not protected. Proven by banning a documentation
address on such a machine today.

The replacement is this module's own dualchain action, already used by the
recidive jail on all four machines, so it is not a new dependency. It bans in
DOCKER-USER as well as INPUT, which ufw's action did not, and it bans all
ports, which ufw's action did.
2026-09-28 20:48:24 +02:00
jschoubben 7c18cdbd39 fail2ban: declare the log its own recidive jail reads
The recidive jail bans whoever keeps coming back by reading fail2ban's own
log, and fail2ban checks every jail's log file while it configures itself --
before it has created that log. On a machine where the file is not there
already, no jail is found for recidive, configuration fails, and the whole
service refuses to start, taking the sshd jail with it. Two machines assigned
this module today came up failed for exactly that reason; the two where it
worked had a log from years of the service running.

Declared create-once: the mesh puts an empty file there when it is absent and
never touches it again, because what grows in it is fail2ban's, and the
logrotate file this module already ships is what keeps it small.

This also reverts the previous two commits' fail2ban.local. It declared a
logtarget that the package already sets to the same path on every machine
here -- pacman reports the config pristine -- so it fixed nothing and said
something untrue about why.
2026-09-28 20:45:32 +02:00
jschoubben c5af8635c8 fail2ban: restart when the log declaration changes
The file that says where fail2ban logs was not in restart-on, so a change to
it would sit on disk with the running service unaware of it -- the same shape
as any other jail file this module already restarts for.
2026-09-28 20:42:42 +02:00
jschoubben f8ca36aacf fail2ban: declare where it logs, so the recidive jail has a file to read
The recidive jail reads /var/log/fail2ban.log and this module ships the
logrotate file for it, but nothing ever told fail2ban to write there. Where
the package default stands, fail2ban logs to the journal, the recidive jail
finds no log file, and the whole service refuses to start -- taking the sshd
jail with it. Two machines assigned this module today came up failed; the two
where it worked had /etc/fail2ban/fail2ban.conf edited by hand, which a
package upgrade would have undone.

Declared in fail2ban.local, because fail2ban.conf belongs to the package.
2026-09-28 20:41:09 +02:00
jschoubben 016ddb2b3a The catalogue hears what it missed
It asks what it missed on every start and the answer never arrived: the control plane replayed each
build it held as a module's event from a module called "control-plane", which does not exist, so its
own account refused the publish and the graph kept the gap. The control plane now states those under
the seat it holds (novox/hq ADR 0134, mesh-controller #129), so this consumes that too — one handler,
because what a build means for the graph is the same whether the build machine says it as it happens
or the mesh says what it already held.
2026-09-28 16:08:46 +02:00
jschoubben 4258f01614 The catalogue prepares its own schema instead of migrating at start
It brought its schema up inside its runtime, on every start. That made a schema it could not reach a
crash loop rather than a stop, with the module graph keeping a gap and nothing saying so — which is
how a whole morning's builds went unrecorded. The mesh now prepares this module's state before it
starts this version and does not start it if that failed (novox/hq ADR 0135): the work moves to an
entrypoint the image names in MESH_PREPARE, beside the entrypoints it already names.

The reason it was at start — that a step blocking the apply would block the very apply bringing the
overlay up — stopped being true when a step's failure became its module's business rather than the
machine's (ADR 0136).
2026-09-28 15:40:28 +02:00
jschoubben eff11b1d4d The catalogue declares the event it emits on starting
Its runtime announces that it has just started and may have missed builds — the event the control
plane follows to replay them — and its manifest did not declare it. A module's authority on the bus
is derived from what it declares, so the publish was refused and the runtime died on start, in a
loop, with the mesh's graph never catching up.
2026-09-28 09:49:03 +02:00
jschoubben 3b77dde666 A merge says which files it changed
Every module built from a repository was rebuilt for a change to any of them: one merge in this
repository meant twenty-six builds, which is what exhausted a public registry's pull limit. The
forge lists the files a merge changed and the event carries them, from the watcher and from the
merge tool alike; a merge that changed more files than were asked for says so, and the mesh then
treats the whole repository as changed rather than guessing.
2026-09-28 09:17:38 +02:00
jschoubben 2b8a668d06 A merge older than the watching is history, not news
An old merge past the first page of the forge's listing surfaced as newer pull requests were
updated, and was announced as if it had just happened; the mesh then rebuilt everything built from
that repository, once per old merge. The moment the watching began is kept with the record, and
only a merge made since is announced.
2026-09-28 05:08:20 +02:00
jschoubben ac5630bee2 A merge the forge announces is said in its log
A trigger that fires silently is indistinguishable from one that did not fire (novox/hq issue
131); the announcement is now one line an operator can read.
2026-09-28 04:44:17 +02:00
jschoubben d70cb18ea0 The forge watches every repository, not the administrator's own
/user/repos lists what the token's user owns, which for the mesh's administrator is nothing — so
the forge module watched an empty list and never announced a merge. It reads the forge's whole
view through the search endpoint, every page.
2026-09-28 04:31:20 +02:00
jschoubben eb62289f89 The forge announces every merge, whoever made it
The merge tool emitted at the instant it acted; a merge made in the forge's own
pages or over its API emitted nothing, and the mesh went on believing every module
current with its source (novox/hq 04-ISSUES/131). Merged pull requests are now
watched the way repositories are: what the forge holds, asked for on a tick,
announced once, with the merge commit and the clone URL a build needs. What has
been announced is kept beside the module's state, so a restart does not announce
the whole history again, and a first tick with no record announces nothing.
2026-09-28 02:54:48 +02:00
jschoubben f5969a2f9f Merge pull request 'nats declares the certificate directory it mounts' (#123) from feat/nats-serves-the-meshs-certificate into main 2026-09-27 22:31:20 +00:00
jschoubben 7f3d259cf5 nats: drop the access to a certificate directory it no longer mounts 2026-09-28 00:16:23 +02:00
jschoubben 721149eda1 nats declares the certificate directory it mounts
The mesh's broker certificate is the operator's, kept outside any module and
mounted read-only by whatever serves the bus; the controller declares that
access and now so does nats, the same way. A mount nothing declares is refused
at registration (ADR 0030), which is how this was found.
2026-09-28 00:15:59 +02:00
jschoubben 8e27bc1e36 Merge pull request 'nats serves the mesh's existing broker certificate' (#122) from feat/nats-serves-the-meshs-certificate into main 2026-09-27 22:07:29 +00:00
jschoubben f1212620e4 nats serves the mesh's existing broker certificate
The module mounted a TLS directory nothing fills, so the server could not start
on a mesh that was not raised by the genesis template. The mesh already has a
broker certificate every machine pins by fingerprint and the controller trusts;
serving the new bus with it means no pin changes when a machine moves and there
is no second certificate to be wrong about. No ca_file: the directory has none,
and a pinning client checks the leaf and nothing else.
2026-09-28 00:04:10 +02:00
jschoubben b1b18ae390 Merge pull request 'nats declares the upstream image it is built from' (#121) from fix/nats-declares-its-base into main 2026-09-27 21:40:50 +00:00
jschoubben 3f0a174392 nats declares the upstream image it is built from
The recipe started FROM the upstream server's digest directly, and the build machine
refuses that: every base is declared under build.on and copied into the mesh's own
store before a build, so a build never reaches out to a registry the mesh does not
run (novox/hq ADR 0097). Found the first time the module was built on a real mesh.

Same digest, now declared as NATS_BASE and arriving as a build argument; the
Dockerfile says where it comes from and why the digest is the index's.
2026-09-27 23:40:16 +02:00
jschoubben c0edebefb9 Merge pull request 'lavinmq, amqp-ping and amqp-email-forwarder leave the catalogue' (#120) from feat/amqp-leaves-the-catalogue into main 2026-09-27 21:30:19 +00:00
jschoubben b9605a6e3b lavinmq, amqp-ping and amqp-email-forwarder leave the catalogue
AMQP is not a provision (novox/hq ADR 0131, design 28 task 5.4). These were the
only three manifests that named it: the broker that provided it, a proof that a
grant worked end to end, and a forwarder reading mail off a queue. Removed, not
converted — a module that wants messaging wants the mesh's bus, reached through
the sdk and named by the mesh-broker seat, and either of the last two is re-done
against that if wanted, as a new module under the record.

The controller refuses a manifest naming amqp from its next release, so these
could not be re-registered anyway. Nothing else in the catalogue referenced them.
2026-09-27 23:27:24 +02:00
jschoubben dca84d3bb4 Merge pull request 'lavinmq holds mesh-broker again, now the check reads the store' (#119) from restore/broker-claim into main 2026-09-27 20:24:22 +00:00
jschoubben 9f9ce92d0f lavinmq holds mesh-broker again, now the check reads the store
The claim comes back for the third and last time. The store's row says the bus seat
answers for `amqp`, lavinmq provides `amqp`, and with mesh-controller#89 the check
that judges a claim reads that row instead of a copy compiled into the build
machine. So this is accepted for the reason it should have been all along.

Restores the holder the controller composes its own bus address through, which is
what ends tonight's crash loop. nats takes the seat over when the cutover is done
deliberately, not because the seat emptied itself.
2026-09-27 22:23:37 +02:00
jschoubben 6e5b2557ba Merge pull request 'Undo: claiming mesh-broker made lavinmq unbuildable' (#118) from revert/broker-seat-claim into main 2026-09-27 19:53:49 +00:00
jschoubben f97b7dd544 Undo: lavinmq cannot hold mesh-broker, and claiming it makes lavinmq unbuildable
Putting the claim back was wrong on its own terms. `mesh-broker` delivers
`mesh-bus`, and a seat that delivers a provision may only be held by a module that
provides it — so the claim is refused at registration:

  lavinmq claims mesh-broker, whose holder answers for "mesh-bus",
  and lavinmq does not provide "mesh-bus" at mesh scope

Which means the merged claim does not restore the holder, it stops lavinmq being
built at all. Removed again.

The seat being empty is still the live fault, and it has only one valid answer: the
holder must provide `mesh-bus`, and the module that does is nats. Recorded against
the rollout, because it moves a step that was optional into the critical path.
2026-09-27 21:27:55 +02:00
jschoubben 5f5798ec8a Merge pull request 'lavinmq keeps mesh-broker until something else can take it' (#117) from fix/broker-seat-must-stay-held into main 2026-09-27 19:17:45 +00:00
jschoubben 42550dbe43 lavinmq keeps mesh-broker until something else can take it
Taking the claim off made the seat unheld, and the controller dereferences that
seat to find its own bus (to-be 26, "the one exception is the controller itself").
Unheld, the composed address fell back to a default port nothing serves, and the
control plane crash-looped: "cannot reach the broker named in MESH_BROKER_AMQP:
dial tcp 127.0.0.1:5672". The broker itself never stopped — it is healthy on the
port the mesh actually assigned it.

lavinmq becoming an ordinary provider is right, and it is still a provider of amqp
here. What was wrong is the order: the seat has to pass from one holder to the next,
and it cannot be empty in between, because the thing that reads it is the thing that
would have to fix it.
2026-09-27 21:13:13 +02:00
jschoubben 51713dd631 Merge pull request 'The Go base has to be 1.26 for what compiles the controller's code' (#116) from fix/go-126-base into main 2026-09-27 19:00:38 +00:00
jschoubben 4cda964a43 The Go base has to be 1.26 for what compiles the controller's code
builder and route-proxy both build from the mesh-controller repository's context,
so its go.mod is theirs, and `nats.go v1.54.0` puts that at `go >= 1.26`. Pinned at
1.25.14 they cannot compile it: the build machine's own build failed with "go.mod
requires go >= 1.26.0 (running go 1.25.14)".

Each moves to the 1.26.8 digest of the flavour it already used — alpine for
builder, debian for route-proxy — so nothing changes but the compiler version.
2026-09-27 20:58:09 +02:00
jschoubben adb02da136 Merge pull request 'The nats module, and every manifest's event names made local' (#115) from feat/nats-genesis into main 2026-09-27 17:31:04 +00:00
jschoubben 06954b5a70 Merge main: the trunk's seat names, this branch's event names
Two lines of work renamed the same seats differently. The trunk named them for their
scope — node-scoped ones `node-*`, leaving `the-artifact-store`, `npm-package-registry`
and `git` as they were — and this branch had renamed ten of them to `mesh-*`. The trunk's
set is what the live controller loads and what the live seats were actually renamed to, so
a manifest claiming this branch's name is one the running mesh refuses. Three of them
needed reverting by hand: git had auto-merged this branch's names where the trunk had not
touched those lines, which is the quiet kind of merge result.

Event names are this branch's, because the trunk has not converted them and they are what
issue 127 was about.

Verdaccio goes with the trunk's removal of it. The template work on dnsmasq's roster fact
is the trunk's, sitting beside this branch's local event names in the same file — the one
hunk where both changes landed together.

75 manifests, all parsing, no claim outside the trunk's set and no event name left in the
old bus's form.
2026-09-27 18:25:57 +02:00
jschoubben 3d7d896014 Merge pull request 'fail2ban never bans a tunnel peer: ignoreip names the mesh range' (#113) from fix/fail2ban-ignores-the-mesh-range into main 2026-09-27 14:55:55 +00:00
jschoubben 278610c0c3 fail2ban never bans a tunnel peer: ignoreip names the mesh range
The jail.local [DEFAULT] gains ignoreip = 127.0.0.1/8 ::1 ${machine:mesh-range}
— localhost plus the mesh's own private range, named through the placeholder
rather than hardcoded (data is the mesh's, ADR 0112). Without it fail2ban could
ban the mesh's own nodes on 10.10.0.0/24; on novox that rule survived only in
memory from a now-deleted HAL file and would be lost on the next restart.
2026-09-27 16:55:34 +02:00
jschoubben b58a3b487d A build's outcome belongs to the role, not to the module holding it
ADR 0121. The builder declared `built` as its own event, so every consumer depended
on which module happens to be the build machine today. It is the build-machine
role's event now: the builder declares none of its own, and the catalogue listens
for `mesh-build-machine.built` rather than `builder.built`.

Nothing changes about what reaches the catalogue. What changes is that it survives
the build machine being a different module, which is the whole reason the mesh has a
word for a role.
2026-09-27 15:37:21 +02:00
jschoubben 7b06a7a408 Event names are local now, in the manifests and in the code
Every module named its events the way the old bus spelled a routing key —
`module.<module>.<verb>`. Design 29 says a module names an event locally and the
mesh works out where it lands, so all 37 were stale against a rule already
decided. On the new bus that derives into a namespace belonging to a module
called "module", so no cross-module subscription in the mesh matched anything:
nothing failed, nothing reacted (novox/hq 04-ISSUES/127).

36 manifests converted, and 43 files of module code with them. The code mattered
as much as the manifests: the runtime builds the subject from what `emit()` is
handed, so a converted manifest with unconverted code would have had the
permission and the subject disagree.

Three things the new check found on the way:

- `photos` emitted an event its manifest never declared, which the new bus refuses
  outright. Declared.
- `showcase` waited for an event nothing emits, so its demo could never be
  triggered — only `showcase` may publish under its own name. It emits both halves
  now.
- `distribution` declared an event named after a different module. It emits
  `image.pushed` under its own name. An event about a *role* belongs on the seat,
  where the name outlives whoever holds it, but the sdk has no way to publish on a
  seat yet, so that stays recorded rather than declared.

The audit logger's "everything" pattern is `**` rather than the old bus's `#`.
2026-09-27 14:42:28 +02:00
jschoubben f0a6ce8d4a Merge pull request 'Rename seat claims to mesh-*/node-*; retire verdaccio (ADR 0121)' (#112) from feat/system-seats-named-by-scope into main 2026-09-27 12:32:29 +00:00
jschoubben 6bedcd3f21 Rename seat claims to the mesh-*/node-* convention; retire verdaccio (ADR 0121)
Claims renamed to match the controller's seat set: node-dns-resolver (dnsmasq),
node-intrusion-prevention (fail2ban), node-packet-filter (nftables),
node-resolver-config (resolv-conf, resolved-split-dns), node-uplink
(networkmanager, systemd-networkd, dhcpcd), mesh-build-machine (builder, +mesh
scope), mesh-catalog (mesh-catalog). showcase now declares its own seat and
claims it. verdaccio removed — the mesh keeps distribution as its registry and
gitea already serves npm, so a second npm registry is redundant.
2026-09-27 14:30:56 +02:00
jschoubben a093c88c32 nats declares its own server settings, and where the mesh's users go
The split the controller now makes, from this side. The module's own
configuration — ports, TLS, JetStream — is a declared file resource, because those
are properties of this container and change when its image does. `bus-users` names
where the mesh writes every account and permission, in the same directory, and the
module's configuration includes it.

**Both files in one directory because they have to be.** An absolute include path
is resolved relative to the including file's directory: nats-server given
`include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf looks for
/etc/nats-server/etc/nats/accounts.conf and refuses to start. Verified against the
server, and recorded in the configuration itself where somebody moving a file will
read it.

**`verify: true` is gone, and it was refusing every connection in the mesh.** It
makes the server demand a client certificate; a host pins this server's exact
certificate and authenticates with the password the mesh minted, and presents none.
Found by building this image and connecting to it as a host would.

The entrypoint now waits for both files and watches the mesh's half: the module's
own does not change without a new declaration, and that recreates the container
anyway. Verified end to end against this image — the mesh's user list rewritten,
the module noticing and reloading the server itself with no signal from outside,
and the connection the mesh already had still working afterwards.
2026-09-27 02:50:35 +02:00
jschoubben f67f0ca9bc Merge pull request 'dnsmasq owns its resolver format: node-zones is a template (ADR 0120)' (#111) from feat/roster-facts-are-templates into main 2026-09-26 23:51:29 +00:00
jschoubben 968473219a dnsmasq owns its resolver format: node-zones is a template, not a controller formatter (ADR 0120)
The node-zones fact was a path; the local=/address= syntax lived in the
control plane. It is dnsmasq's configuration language, so it moves into
dnsmasq's manifest as a template over the roster. The mesh renders it; it
reads none of it. Output is unchanged.

Lands with mesh-controller's ADR 0120 change — the two are one schema step.
2026-09-27 01:33:28 +02:00
jschoubben ad219beee2 Merge pull request 'The uplink's managers are modules: networkmanager, systemd-networkd, dhcpcd (hq ADR 0117)' (#110) from feat/the-uplink-modules into main 2026-09-26 23:00:30 +00:00
jschoubben ae99204a8c Claim the renamed seats (novox/hq ADR 0118)
Ten manifests claim mesh-* names now. What they PROVIDE is unchanged: gitea
still provides git and npm-package-registry, and a consumer requires the
interface, not the seat.
2026-09-26 23:07:09 +02:00
jschoubben 226eab4c6f Merge pull request 'dnsmasq: the operator's own names have a home the mesh never rewrites' (#109) from feat/dnsmasq-has-a-home-for-operator-names into main 2026-09-26 20:43:53 +00:00
jschoubben bb8f2e76a9 dnsmasq: the operator's own names have a home the mesh never rewrites
A workstation's job includes names that are neither a mesh machine nor
a routed name (novox/hq 122): shanks carries 13 Mediahuis entries in
/etc/hosts, and mesh-wireguard replaces /etc/hosts whole when taken —
so without this they vanish, and the take gates the node. Two homes,
neither the mesh's to own: conf-dir=/etc/dnsmasq.d/,*.conf (drop-in
directives, HAL's dnsmasq-app used exactly this) and
addn-hosts=/etc/hosts.local (plain host lines). The mesh creates and
rewrites neither; a machine with none loses nothing. The migration
moves such names here BEFORE the /etc/hosts take, closing the window.
2026-09-26 22:43:29 +02:00
jschoubben 86882cacd4 nats provides mesh-bus (novox/hq ADR 0120) 2026-09-26 21:17:21 +02:00
jschoubben ea7f6796e8 lavinmq is a provider, not foundation
It claims no seat: mesh-broker is the NATS server's (novox/hq ADR 0119).
The amqp interface stays exactly as it is — a backing service a module may
require, like a database.
2026-09-26 21:08:40 +02:00
jschoubben 372450851f Merge pull request 'mssql: its data is placed — the last /services placement retires' (#108) from feat/mssql-data-is-placed into main 2026-09-26 18:36:16 +00:00
jschoubben f4e4e12c99 mssql: its data is placed — the last /services placement retires
The stated path was the adopted-data exception; with the take done and
the placement vocabulary live, the exception has no reason left. The
landing window renames the directory and recreates the container, since
a changed volume path does not do that by itself (hq 126).
2026-09-26 20:36:03 +02:00
jschoubben fd9be011c0 Merge pull request 'sshd: the operator's door is a module' (#107) from feat/sshd-module into main 2026-09-26 18:15:08 +00:00
jschoubben a85b0ee346 sshd: the operator's door is a module
The spec is the working system: HAL's 99-hal.conf, restated as
10-mesh.conf so lexical include order makes the mesh's answer the one
that wins while the predecessor's file is still on disk. Subsystem
stays the stock config's — first-set wins and it sits before the
Include. Port 22 from anywhere, said in listens with its reason: the
machines that need the door are exactly the ones not on the mesh yet,
and locking the operator out is the one failure a firewall must never
arrange.
2026-09-26 20:14:54 +02:00
jschoubben 34243c9e34 Merge pull request 'dnsmasq: the runtime's DNS is written into daemon.json, never over it' (#106) from fix/dnsmasq-writes-into-daemon-json into main 2026-09-26 18:12:20 +00:00
jschoubben 870a541072 dnsmasq: the runtime's DNS is written into daemon.json, never over it
The file is shared — the operator's insecure-registries for the mesh's
own store live there — and replacing it whole would break every pull
from that store the moment the module is taken (the 098 class, caught
in the pre-take diff this time). ADR 0102's verb is merge.
2026-09-26 20:12:06 +02:00
jschoubben 9b063a77b2 nats: the module, and an image that reloads in place
Step 1.1 and 1.2 of novox/hq ADR 0116. The server is a built artifact rather
than the upstream image directly, because it needs an entrypoint of its own:
the host can only recreate a container, and recreating the bus for every
permission change drops every connection and every in-flight ack. nats-server
reloads on SIGHUP by itself, so the config is mounted as a directory (not
digest-tracked, hq issue 103) and the entrypoint watches the one file.

Verified against the real server, not assumed: a user added to the config
connects, a revoked one is refused, both within one poll interval, with the
container's PID and restart count unchanged and "Reloaded: accounts" in its
log.

Two corrections found by checking rather than reading:
- the seat delivers nothing now (hq ADR 0117), and the controller's parser
  refused the manifest until it did — "nats claims mesh-broker, whose holder
  answers for amqp, and nats does not provide amqp"
- pinned to the multi-arch index digest; the first pin was the amd64
  manifest, which builds here and fails on any other architecture
2026-09-26 19:34:14 +02:00
jschoubben a59750fa28 Merge pull request 'mailu: the smtp provision serves the name its certificate answers to' (#105) from fix/smtp-serves-its-tls-name into main 2026-09-26 17:27:06 +00:00
jschoubben 81592a3b2c mailu: the smtp provision serves the name its certificate answers to
A consumer connecting by the binding's address meets a certificate for
mail.novox.be and refuses it — found live by the forwarder's cutover
proof, one send before production would have. The TLS name is mailu's
own fact (HOSTNAMES), so the binding carries it; consumers say
${bound:smtp:name} and verification holds.
2026-09-26 19:26:53 +02:00
jschoubben 705ceec1e7 Merge pull request 'Six modules name no /var/lib: the root is a place, the maps reference it' (#104) from feat/six-modules-name-no-var-lib into main 2026-09-26 16:21:00 +00:00
jschoubben afdd149ab7 Six modules name no /var/lib: the root is a place, the maps reference it
The state directories say place "." — the assignment's own root — and
every bind, secret, own-secret, receives and grants path references it
as ${dir:state}/…; grants directories that are their own resources are
placed by id. gitea's two coincidence strings from the first pass
(${dir:data}base.json — resolving correctly by pure concatenation) are
spelled honestly now. What still says /var/lib is inside containers —
the software's contract — or under /var/lib/mesh, the mesh's own
plumbing, which the requirements unification absorbs next. Every
resolved path is byte-identical to what runs; landing this is a no-op
on the node, and the converter checks its own boundaries this time.
2026-09-26 18:20:47 +02:00
jschoubben dde8b15483 Merge pull request 'mailu: seventeen data directories are placed, not stated' (#103) from feat/mailu-dirs-are-placed into main 2026-09-26 16:07:24 +00:00
jschoubben 8a046be198 mailu: seventeen data directories are placed, not stated
Each resolves to <root>/mailu/<id> — the maildir at
/var/lib/mailu/data-mail, certs at data-certs, and so on. Landing this
is a window, not an edit: seventeen renames on the node (the nested
data/ tree flattens to the ids), then the full stack recreated, because
a changed volume path does not recreate a container by itself (hq 126).
Ids are untouched on purpose — a renamed id orphans its held record,
and the mail spool is the wrong place to learn what a removal step does
with one.
2026-09-26 18:07:11 +02:00