The names region of /etc/hosts is written by mesh-wireguard and dnsmasq answers from it, but read
it once at start: after the mesh stopped publishing public names (hq ADR 0191) every machine's
hosts file was right and every resolver still answered mail.novox.be with a tunnel address. The
config comment that says so changes the file, which restarts dnsmasq once everywhere.
A package has one owning module, and sshd already owns openssh on every node — assigning
ssh-client was refused for both declaring it, and the refusal left every node unresolvable
until it was unassigned. Owning it was wrong besides: unassigning ssh-client would have
removed the package sshd serves from. The client binary ships in the package sshd holds.
Requires openssh; creates ~/.ssh (0700, owned by the operator account via
${machine:account}); writes every other node's Host block (HostName + User
<account>) into a marked region of ~/.ssh/config (home-scoped, into:block), so
`ssh <node>` reaches each peer as the right account and the operator's own
config is kept. Universal-tier: assigned wherever a person logs in; a node with
no account gets no config.
step-ca and public-acme both offered acme-ca on novox, and route-proxy's pin names a node, not
a module — so which one certified the public names depended on provider order. After the
controller restart on 2026-10-03 it came out as step-ca, and every public site served a
certificate no browser trusts. step-ca's own names are already certified through
internal-acme-ca; acme-ca is the public authority's alone.
The jail read gitea's container journal, which carries its sshd's lines,
but matched only the web login. 167 ssh attempts an hour from the
internet went unbanned. Two patterns, one per attempt: an unknown user,
and a user sshd refuses; tested against a day of the real log, 946
matches and none on an accepted login.
The home server banned the house's own router within an hour of the first public jail: the router
reflects local traffic, so every client in the building arrives as the gateway's address. Every
private range joins the mesh's own in the never-ban list.
fail2ban expands <HOST> to a named group, so two in one pattern is a duplicate group name and
the daemon refuses to start at all -- every jail on the machine, not just this one. Two patterns,
one <HOST> each: the certificate refused for an unserved name, and the request refused for one.
Caught live on the control node (hq ADR 0179).
The pattern ended at the line's end, which only the certificate refusal does; a request for
an unserved name carries trailing text and never matched. Caught against the live lines
before the jail counted anything (hq ADR 0179).
The module gains a runtime carrying only the fail2ban client with the daemon's socket shared in,
serving status/banned/ban/unban and its own fail2ban_settings. It declares jailing, so the
controller's composition lands in jail.d/mesh.conf and filter.d; mailu, route-proxy and gitea log to
the journal and declare a jail reading it by container name. The base is strict: three in a day for
a day, twice banned in two weeks for four; the mesh's range stays never banned.
Both sources sit in tools/, so tsc took tools/ as the root and wrote
dist/index.js, while the runtime loads dist/tools/index.js: the module
started and served nothing.
Five tools on the machine the lab runs on: check, run beds against
branches on the forge, a run's status, its log, and stop. A run checks
out every repository the lab builds, side by side, and runs the suite;
one at a time, answered at once with an id (novox/hq ADR 0172).
The seat's three verbs over the machine's own tools: the filter as enforced
(nftables and the legacy filter), the mesh's own table reloaded from its file,
and one rule set the mesh did not write removed by the name the host reports
it under (ADR 0168) — a predecessor's chain loses its jumps and goes, the
runtime's user chain is emptied back to its return, a table of the machine's
own goes whole; the mesh's tables, the runtime's chains, a built-in chain and
an active found firewall's chains are refused. Tested over the shapes two
machines of the first mesh reported live. The module's own tool stays.
Which DNS server the home network's DHCP hands out could be changed only
in the controller's own interface or by hand against its API (novox/hq
issue 198).
The addresses dnsmasq listens on beside the machine's are a setting, so a
machine that answers its own LAN can say so (novox/hq issue 198). Docker's
daemon.json no longer merges the module's settings: it needs none, and a
setting reaching it is a key dockerd refuses. The host still merges it
into the existing file.
A module's settings merge into every mergeable file it owns, so the
setting reached docker's daemon.json beside dnsmasq's config, where
dockerd would refuse it (novox/hq issue 198). Back to the fixed
loopback line until settings can be kept out of files they are not for.
Loopback by default, as before. A machine that answers its own LAN adds
its LAN address, and its DNS endpoints' reach opens the filter (novox/hq
issue 198). The mesh-wide default must be set before this lands.
Proven on a throwaway server: as the administrator a caller's $(SQLCMDPASSWORD) returned
the sa password, and a line beginning ':!!' ran a program in the tools container. The
statement now runs as mesh_mssql_reader (CONNECT ANY DATABASE, SELECT ALL USER SECURABLES),
with substitution off (-x), after the module's own text on the first line, and a line break
is refused. go-sqlcmd v1.10.0 is installed at a pinned digest: the image never had sqlcmd,
so every mssql tool failed with spawn sqlcmd ENOENT.
The verb wrapped the caller's text in BEGIN READ ONLY ... ROLLBACK as the superuser, so
'COMMIT; ...' left the transaction and, proven on a throwaway server, COPY TO PROGRAM ran a
shell command on the database host. The statement now runs as mesh_store_reader:
pg_read_all_data, no other grant, read-only transactions by role and session, its password
an own-secret the mesh mints. Without that password the call is refused. -q drops the
command tags that came back as rows keyed by BEGIN.
The same split lost them the other way round: the merge base held both changes, each branch had reset
the other's files, and the three-way merge kept neither. Both halves of hq ADR 0161 are on main again
with this.
The uplink branch was split from the vault's with the vault's files reset to a main that did not yet
hold #205; merging it afterwards took the older vault definition along (no claim, the refused event
names), and the vault could not be built. Restored to #205's state.
A module publishes under its own name only; secret.provisioned read as another module's event and the
builder refused the vault's definition today, so the seat claim could not be built. The three events lose
the prefix; nothing outside the vault listens for the old names.
The vault's claim makes a second provider of secret a second claimant, refused by name. The three
uplink definitions declare uplink-networkmanager, uplink-systemd-networkd and uplink-dhcpcd, which the
host reports for the manager it finds active, so the holder for a manager the machine does not run
is refused the way any missing capability is. Merges after the controller holds the seat and the host
reports the capability.
The store's databases and query are registered under mesh-store, so the runtime serves them on the
seat's subjects wherever postgres holds the seat and never lists them as postgres's own; the claim
names them, so the mesh can judge the holder without postgres listing the seat's verbs among its
tools. Scoped to what the store enables: creating a database stays postgres's tool.
bazarr, bookshelf, lidarr, nzbget, ombi, plex, qbittorrent, radarr,
sonarr and tautulli live in novox/mesh-media-catalog (#195 moved jackett
and left these behind). A build of this repository at a commit today
registered plex and nzbget from these copies, which carry no provides,
and ace's plan stopped resolving. Nothing here depends on the directories:
home-assistant consumes their provisions by name.
Named as the seat names them so the runtime finds them by name; the same calls as its own tools.
Its definition now lists its tools, which is what holding a seat with verbs demands at registration.
Merge before the controller declares the verbs on the mesh-store seat.
Both land in settings.js and the runtime's config file, and the containers that read them restart
on those files; a rotation is a new value and a restart. The broker credential says nothing yet: its
other party is the bus, and that rotation is the two-party form.
Ten tools the console lacked for the actions a review and a merge leave behind: close or reopen a
pull request whose work landed elsewhere, change its title or body, read its files, its diff and its
comments, reopen an issue, read one file at a ref, list branches, delete the branch a closed pull
request leaves. Each is the client's own call; `gitea_api` stays the escape hatch for the rest.
Tested against the fake forge through the compiled tools, the way the console calls them (13/13).
The container mounted /services/media literally — one installation's path
(ADR 0112). The access is now declared by id and mounted as ${access:media};
the assignment says where the library is (ace: /storage/media, hq 153).
The module named /var/lib/n8n, /services/n8n/n8n-data and n8n.novox.be -
paths and a domain no definition may carry (ADR 0112). State and data are
placed directories; the public name is ${bound:route:name} (depends on
mesh-controller #149), for N8N_HOST and WEBHOOK_URL alike.
The endpoint said 5682 while the container publishes 5678. 5682 was one
machine's host port; the endpoint is the software's port and the mesh
assigns the machine's (ADR 0038).
n8n had been run from an image in a registry that no longer exists: the
upstream image plus shadow, a `media` group (2000) with `node` in it, and
a global `uuid`. That recipe is now this module's Dockerfile, built on the
upstream 1.71.3 image named in build.on by digest, with uuid pinned to the
version the running image carries (14.0.1) - Code nodes require() it. The
media group is how the container writes into the shared media library, a
read-write `access` (ADR 0051), mounted where workflows expect it,
/media-library.
The workflows also use a redis (the Redis nodes of the chat workflows) and a
Selenium Chrome (the scraper), which the previous deployment ran beside n8n.
Both are containers on the module's own network, publishing nothing, pinned
to the digests in use; redis keeps its append-only file in a placed
directory.
The basic-auth secret is gone: N8N_BASIC_AUTH_* was removed in n8n 1.0 and
did nothing. The grant's password is a 0400 file owned by `node`, read
through DB_POSTGRESDB_PASSWORD_FILE, so nothing secret is in the
environment. The credentials' encryption key is n8n's own, in the data
directory (config), and moves with it - nothing to mint or accept.
Verified: catalogue tests with MESH_CATALOGUE set; the Dockerfile built
against the pinned base gives n8n 1.71.3, uid 1000 in group 2000, uuid
14.0.1 - the running image's shape. Throwaway containers: an instance on
PostgreSQL 15 with an owner, a workflow and an encrypted credential;
stopped, copied, dumped from the copy, restored (--no-owner --role, the
uuid-ossp extension pre-made by the superuser) into a grant-shaped
database on the postgres module's pgvector image (PG17); the new shape
(password from the file, data dir copied) serves /healthz, the owner logs
in, the workflow is listed, and the credential decrypts with the carried
key. The node user writes into a root:2000 0775 library through the media
group; redis and Selenium resolve by name on the module network and
Selenium reports ready. Test containers and data removed.
Ten definitions name each access by id; the path stays as the default an assignment may replace,
and the host side of every mount says ${access:<id>}. Resolved with no placement, every definition
names exactly the paths it named before (TestPlacedDirectoriesKeepTheirPaths, extended). On an
adopted machine the assignment now says `accesses: {<id>: <path>}` and the mount follows.
Needs the controller that knows an access id (mesh-controller #176); the running one refuses the
field at registration.
Consumers read `${bound:smtp:domain}` and `${bound:oidc-client:issuer}`, and both keys reached
them only because a module's settings were laid over everything it served. Issue 173 stops that: a
setting overrides a key a served fact declares and adds none. So the two providers declare the keys
their consumers read, as the operator's value (`${setting:…}`, ADR 0155), and the setting that
already carries each fills it. Nothing a consumer reads changes.
Merges first: under the controller that still merges settings over served facts this is the same
value, and the controller that stops merging (mesh-controller, feat/the-mesh-places-its-own-files)
needs these declared before it rolls out.
48 definitions stop naming /var/lib/mesh/<module>: the directory says `place: "mesh"`, the two
subdirectories beneath it (gitea's runtime state, anthropic-manager's output) state their path
beneath it, and every credential, binding, merged file and mount names it as ${dir:mesh-state}.
Resolved on the default root, every definition names exactly the paths it named before —
TestPlacedDirectoriesKeepTheirPaths in mesh-controller, run over both checkouts. Needs the
controller that knows the word (mesh-controller, same branch) one release ahead.
jackett leaves: it is registered from novox/mesh-media-catalog with sonarr,
radarr, lidarr, bazarr, nzbget, qbittorrent, bookshelf, plex, tautulli,
kometa and ombi (PRs 145-168 consolidated there). What those branches
changed outside the chain stays here: home-assistant's provisions
(sonarr-api, radarr-api, mqtt-topic — from #147) and searxng's sidecar
dialling the port it was given (#154).
Twenty-eight modules' data directories are placed: the root as place ".", a sub-directory named by
its id, and every host-side reference — binds, secrets, own secrets, grants, receives, file paths,
mounts, env-files — as ${dir:<id>}. Resolved on the default root every path is the one the manifest
named before, which the controller's TestPlacedDirectoriesKeepTheirPaths proves over both checkouts;
so no data moves and no machine sees a change. Five directories whose id is not their last segment
keep their path as a placement (novox/hq issue 119, ADR 0112, design 27).
The forge's own users are the mesh's to settle — making the builder's login
a site admin so private repos build (hq 229) — and the tools' token had no
write:admin. A token kept from before a scope was added lacks it, so the
client now treats the forge's 403 "required scope" like a 401: the source
re-mints by name with the whole list and retries once. The fake forge in the
tests learns /repos/search, which the client has used since 2026-09-28 and
which had left 9 of the 11 token tests failing on main.
listens said 4000 while the container publishes 8080; the old assignment's port setting hid it, and
the rename lost the setting, so the proxy dialled a port nothing answered (2026-09-30).
The rename changed the network resource's name and not the container's network, so the container
looked for a network that no longer exists and the site answered 502 (2026-09-30).
9090:9000 and 9443:9443 were the predecessor's machine numbers written into the
definition. The manifest now says 9000 and 9443 and the mesh assigns the
machine ports on each node (the portainer slice of #173, which is stale).
keycloak, minio and n8n are told their names from their route bindings; mailu takes its domain, site
name, website and proxy address as settings and its front's name from its route, and the provisioner
reads the domain from the merged config; builder and route-proxy package the controller from the git
seat; the applications built outside the mesh say so per container; matrix says which of the world's
servers it means; the site module is named website, and the why prose no longer names a name (novox/hq
ADR 0155, issues 122 and 134). module check passes over all 77.
#186 moved settings.js to /data for the image's health check but left
settings.json at /config; settings.js requires ./settings.json, so node-red
crashed at start. Both now mount under /data.
The image's /healthcheck.js requires /data/settings.js, so with the mesh's
settings mounted at /config the container ran fine but reported unhealthy
forever. Mount the same file at /data/settings.js and point --settings there.
The record is prose wrapped at a hundred columns; matched line by line, the first live search for a
sentence of ADR 0025 found nothing. A line is matched together with the next, emphasis marks are
ignored, and a hit still names the line it starts on.
The manifest declared its secrets as secrets.secret.<name> and required a
"secret" provision, so the controller saw no own secrets and every accept
was refused (the influxdb defect of #180). The directories functions and
storage shared their ids with the containers of the same name, which the
host refuses as two resources with one identity. The containers are now
edge-functions and storage-api; the placed directories keep their ids and
paths.
A module keeping a checkout of a repository of decisions, designs and issues from the git seat,
current on every announced merge and on a timer, answering records_search / records_read /
records_list / records_status / records_sync at the commit it read (novox/hq ADR 0025, ADR 0153).
The repository is a setting; it names no mesh.
The received grants file names each pair secret by its HOST path. A sidecar
that mounts the placed grants directory at another path inside its
container sees mesh.json and not the secrets beside it — mosquitto's
provisioner reported ENOENT for a file that was there, and nodered's mqtt
step failed against a login that was never created. Mounted at its own
path now, as postgres does.
The tool runtime's own client, mesh serve, started by the mesh on the credential it sealed to the
machine (novox/hq ADR 0152, design 34): invokes every tool, listens from the machine only, holds no
state. Checked with module check before it was ever registered (hq issue 148).
They came from `requires: secret`, minted by the vault — right for a fresh
setup (the image's INIT_* variables read them once), wrong for an instance
that already exists: setup is skipped, the minted values match nothing,
and the provisioner holds a token the server never issued (issue 100).
InfluxDB will not take a chosen token value, so the operator token must be
accepted from the instance (`secret accept ace influxdb admin-token`); the
password can be either. As own secrets both are minted for a fresh install
exactly as before, and accepted where the data already knows them.
Found migrating ace's influxdb.
A program that checks its resolver validates — Mailu's admin does, at
start — could not use the mesh's resolver, and the one it ships instead
knows no mesh name (novox/hq issue 171). proxy-dnssec copies the AD bit
from 1.1.1.1 and 8.8.8.8, both of which validate.
admin is the one container here that reaches something by a mesh name:
the database, bound in database.env. A mesh name is answered by the
machine's resolver, which the runtime hands every container that does
not name its own (novox/hq ADR 0148); Mailu's unbound knows no mesh
name, and since the mesh stopped copying names into containers admin
could not find its database. The others keep unbound — rspamd needs a
validating resolver for DNSBL lookups and asks for no mesh name.
The module's server container was called mesh-store and its data directory
/var/lib/mesh-store — the seat's name reused for the module's own resources,
a leftover from the first migration. On a machine whose postgres holds no
seat (ace, as a database provider only) that produced a container called
mesh-store holding nothing of the kind. The container is now named
postgres. The data directory keeps its path: a path change recreates a
running container on an empty directory (hq 126), and novox's store lives
there.
Rolling this out recreates novox's store container once (a restart on its
bind mount, no data moves). The two catalogue-test failures on this branch
(resolver_manifests_test) fail identically on main today.
dnsmasq admits a query by the interface it arrives on when told interface=,
and by the address it is sent to when told listen-address=. A container's
query is sent to the machine's private address but arrives on docker0, so
interface=mesh0 dropped it silently on every machine (novox/hq issue 110).
Name the address, not the interface.
novox/hq 04-ISSUES/110. The module writes the runtime's `dns` key and
deliberately ordered no restart, because a restart stops every container.
It also ordered no reload, and the key holds only for containers created
after the runtime next starts. On two of four machines the runtime
predated the file — one since August — so every container there was
handed a public resolver and no mesh name resolved, while everything read
as fine. The third machine works only because its runtime happened to
restart later.
The file now also sets live-restore, which the runtime reads on a reload,
and the module declares the runtime reloaded when the file changes (ADR
0102: reload, don't restart). A reload still does not make `dns` take
effect; what it does is make the one restart that key needs keep every
container running. That restart stays the operator's, once per machine,
and is harmless from the second time on.
The mesh restarts nothing here. Undeclared, the runtime's unit goes back
to the state it was found in (ADR 0118).
MESH_PROVISION_POSTGRES named 127.0.0.1:5432 literally; the seat twin
(${seat:mesh-store:5432}) corrected it only on the machine holding the
mesh-store seat. On any other machine — ace, where the module provides
postgres-database without the seat — the twin is empty and the sidecar
would have dialled whatever else holds 5432 (HAL's postgres). ${port:5432}
is the machine port on every node; on novox the twin still says the same
6852, so nothing moves there.
sonarr, radarr, lidarr and bookshelf reached jackett as http://jackett:9117 (a HAL
container name) or https://indexers.zurag.be (its public route), typed into each app by
hand. The mesh has neither: an app now requires jackett-api and its downloads step writes
the bound address into the app.
Serves scheme, port and url-base; `at` and the machine port come from the binding. Mesh
scope, like sonarr-api: an indexer proxy shares no files with its consumers.
The pair credential is jackett's one API key. The mesh cannot mint it, so the operator
accepts it per consumer pair (ADR 0092), as #156 does for sonarr-api; a consumer's step
refuses a minted value and names the accept.
ace's grafana reads InfluxDB through a data source somebody typed into
its database: a LAN address, a database InfluxDB 2 does not have, and a
password for a v1 user of an earlier instance. Nothing in the mesh knew
it existed, so migrating influxdb could only break it further.
grafana now requires influxdb-api, contributes read access, and the mesh
renders a provisioning file grafana reads at start: the address, port,
org's default bucket (as the InfluxQL database) and its own login from
the binding, the password by $__file from the pair credential the mesh
delivers, 0400 for grafana's uid 472. It is a data source of its own
name and uid, read-only in the UI and not the default, so the data
source a person made is never overwritten; a changed binding or a
rotated password restarts grafana, which re-reads the file.
Includes #152 (merged into this branch): influxdb provides influxdb-api.
Node-RED's one broker node pointed at zurag.be:1884, where nothing listens. nodered now requires
mqtt-topic (asking for every topic: flows follow the devices' own) and a run-once `mqtt` step —
declared last, restarted when the binding, credential or settings change — points the mesh's broker
nodes at the bound broker through Node-RED's admin API with the module's api-token: the node the
step makes itself when none is named, or the ones an assignment names in `mqtt.brokers`. Only host,
port, TLS and the login change; the broker is asked first whether it takes the login; the deploy is
against the revision read ("nodes", so only that node restarts) and a digest makes a rerun a no-op.
A broker node nobody named is never touched. settings.js keeps `mqtt` and `topics` out of Node-RED.
mqtt-topic served nothing: with two listens the mesh could not say which port a consumer dials, so a
consumer had to type 1883 into its config. It now serves the MQTT listener's port (the machine's,
once assigned) and the scheme, so `${bound:mqtt-topic:port}` fills.
The provisioner confined every consumer to `<as>/#`, which leaves nothing for the consumers the
broker exists for: Home Assistant discovers under homeassistant/# and tasmota/discovery/#, and
Node-RED's flows follow the devices' own topics. A consumer now contributes `topics` (MQTT topic
filters) to its mqtt-topic requirement and is granted exactly those; with none, its own subtree as
before. Settings merge into contributions, so an operator narrows a grant per assignment. The role
is brought to exactly the wanted ACLs (stale ones removed), `holds` checks the ACLs too, and an
invalid list is refused, never quietly narrowed. Only the role named for the consumer is touched:
a client carried from the predecessor's password file keeps its own.
grafana's data source requires influxdb-api, which influxdb provides only
on #152's branch; merged so this branch's catalogue has the provider of
everything grafana requires. #152 should merge first.
grafana's data source and Node-RED's influxdb nodes reached ace's
InfluxDB by a LAN IP or a public name nobody routes, with a credential
somebody made by hand. Now a consumer requires influxdb-api and is told
where it is, which org and default bucket it serves, and signs in with
the password the mesh minted for the pair.
The credential is a v1-compatibility authorization, made per grant by
the new provisioner: InfluxDB 2.x generates API tokens itself and
ignores one the caller sends, so a v2 token could only be accepted by
hand per pair; a v1 authorization takes a caller-chosen password (8-72
characters, the mesh mints 40) and reads/writes every bucket as a
database of its name over InfluxQL and line protocol. A consumer
contributes `access` (read, write, read-write) and, for writing, the
buckets; a missing bucket is made and never deleted. Only
authorizations named mesh_* and marked [mesh] are ever changed or
removed; anything else of that name is refused and left alone.
The org and default bucket are served facts the assignment's settings
set, reaching both the consumers and the provisioner's config.json.
The module named /var/lib/letta, a layout no definition may carry (ADR
0112); state is now a placed directory holding the bindings and secrets.
letta had no route, while the server it replaces is reached by its public
name (a workflow calls it there). It now contributes one for its `web`
endpoint; reach is the assignment's.
The server needs an OpenAI key for agents on OpenAI models, and nothing
gave it one: `openai-api-key` is an own-secret, accepted from the
operator (a key someone chose, not one the mesh can mint). The
server-password is accepted the same way where a server already has
clients. Both, and the database password inside LETTA_PG_URI, stay in the
server's environment: letta 0.6.x reads settings from the environment
only, and its startup.sh starts an embedded PostgreSQL unless
LETTA_PG_URI is set - the declared reason now says so.
The runtime's tools never authenticated: the client sent only a Bearer
token, and 0.6.x's --secure mode checks X-BARE-PASSWORD ("password <it>")
and answers 401 otherwise. The client now sends both. Its password comes
from the runtime config file (the key client.ts reads first) instead of an
env-file, so the runtime container no longer carries a secret in its
environment.
Image: the same 0.6.8 image, now pinned by the index digest ace runs
rather than its amd64 manifest.
Verified: catalogue tests with MESH_CATALOGUE set; tsc -p tsconfig.json in
the mesh-tools build image. Throwaway containers: a fresh 0.6.8 with its
embedded PG and two blocks made through the API; stopped, copied, dumped
from the copy; restored (schema letta + vector pre-made by the superuser,
--no-owner --role, search_path set on the database as the original had
it) into a grant-shaped database on the postgres module's pgvector image
(PG17); started in this shape: alembic finds nothing to do, both blocks
are there, a wrong password is refused with 401, and the patched client
lists agents. Test containers and data removed.
The module named /var/lib/baserow and /services/baserow/data, a layout no
definition may carry (ADR 0112). State and data are now placed directories;
bindings and the grant's secret live in the placed state.
The all-in-one image's entrypoint honours DATABASE_PASSWORD_FILE (file_env
in /baserow.sh), so the grant's password is mounted rather than put in an
env-file, and "secrets-in-environment" is gone (ADR 0086). SECRET_KEY is no
longer minted: the image keeps it, and its JWT signing key, in the data
directory (.secret, .jwt_signing_key) and imports them on start, so a moved
data directory carries the keys its sessions and tokens were made with.
DISABLE_EMBEDDED_PSQL makes a missing grant fail loudly instead of starting
an empty embedded database.
BASEROW_PUBLIC_URL was http://localhost. Baserow answers only the host of
that URL - any other Host is looked up as a published builder site and gets
404, /api/_health/ included - so it is now https://${bound:route:name}
(depends on mesh-controller #149).
The runtime's tools could never have worked: its config was "{}", and the
client's Host override was silently dropped by Node's fetch, so calls by
container name would 404 even with credentials. The client now uses
node:http (which sends the Host it is given, with a Content-Length -
Baserow reads a chunked body as empty) and re-authenticates once when a
cached JWT is refused (access tokens last minutes, the runtime weeks). The
password is the accepted `admin` secret; the email is an assignment
setting merged into the same file, the host is the route's name.
Image pinned to the develop-latest build ace runs today (Baserow 2.3.4,
built 2026-09-18). The old pin (built 2026-09-04) is older than ace's data.
Verified: catalogue tests with MESH_CATALOGUE set; tsc -p tsconfig.json in
the mesh-tools build image. In throwaway containers of the pinned image: a
fresh embedded-PG instance with a user, workspace and 5-row table; stopped,
copied, dumped from the copy (start-only-db); restored with --no-owner
--role into a grant-shaped database on the pgvector image the postgres
module pins (PG17); started with this shape (root 0600 password file,
embedded PSQL disabled, copied data dir without postgres/): health 200,
the user logs in, the 5 rows are there, SECRET_KEY and the JWT key are
imported from the data dir. The patched client lists applications and rows
through the container name with the public Host, and recovers from a
refused token. Test containers and data removed.
ace runs Supabase under HAL as upstream's 13-container compose: a 2.2 GB
database (1.8 GB of it the dormant `novox` schema, 5.8 M rows in its largest
table), Kong at supabase.zurag.be, the pooler on 5433/6543. This is that stack
as a catalogue module, same images (the digests ace runs), same container
names so an assignment holds the running ones and a take replaces them.
The database stays inside the module. Supabase is a Postgres distribution: its
own image with pgsodium, pg_graphql, pg_net, vault and timescale preloaded, a
superuser (supabase_admin), a dozen reserved roles and a second database
(_supabase). A postgres-database grant - one database, one unprivileged role -
cannot hold it, so ace's data directory moves as a copy, not a dump/restore.
What HAL did by shell and environment the mesh now renders as files:
- kong.yml carries the anon/service keys and the dashboard login (owned by
kong's uid 100), instead of an entrypoint that eval'd the environment;
- GoTrue reads a dotenv file (auth -c), PostgREST a config file, Vector its yml
with the Logflare key in it, the database POSTGRES_PASSWORD_FILE and a
jwt.sql rendered with the secret (owned by postgres, uid 105). None of these
five containers has a secret in its environment.
- realtime, storage, meta, functions, analytics, studio and supavisor read
their credentials from the environment only; each declares
secrets-in-environment with the reason (ADR 0086).
- Upstreams are container names (supabase-db, supabase-kong, ...) instead of
compose service names, which the mesh does not have.
- SITE_URL / API_EXTERNAL_URL / SUPABASE_PUBLIC_URL are
https://${bound:route:name} (mesh-controller #149); on ace HAL rendered them
as "https://supabase." - broken today.
- Vector reads the docker socket, as upstream does, but now includes only this
module's containers instead of every container's logs on the machine.
Three things compose did that a declaration cannot, done as steps: a run-once
seed copies the image's /etc/postgresql-custom into the placed config
directory with cp -n (a named volume did that implicitly; never overwrites the
pgsodium root key), and two run-once gates wait for the database and for
Logflare, which compose expressed as depends_on: service_healthy.
The pooler bootstrap (pooler.exs) takes the tenant id and pool sizes from
settings.json, the module's one merge:json file, so ace keeps its tenant
"zurag"; and it repoints an existing tenant whose database host is not
supabase-db - HAL created ace's with host "db", which no longer resolves.
Secrets (vault, requires "secret"): postgres, jwt, anon-key,
service-role-key, dashboard-user, dashboard, logflare, pooler-vault,
key-base-a + key-base-b (concatenated: Phoenix wants 64+ bytes, a minted
secret is 40), openai. Three cannot be minted on any machine: anon-key and
service-role-key are JWTs signed with jwt, and pooler-vault must be exactly
32 bytes (AES-256-GCM, found in the bed). They are accepted. On ace every
secret the data already knows is accepted (all but key-base-a/b).
Not carried: Kong's 8443 and Logflare's 4000 on all interfaces (nothing
outside the module uses them); realtime's DB_ENC_KEY stays upstream's constant
(realtime deletes and re-seeds that tenant from its environment every start,
and the key must be exactly 16 bytes).
Verified: catalogue tests with MESH_CATALOGUE pointed here on mesh-controller
main and #149 (on main the render is refused for "name", never written
empty). The #149 resolution with stub providers, turned into a throwaway
stack of all 13 pinned digests with dummy secrets and the rendered files at
their owners and modes: the database initialised through the rendered scripts
(jwt setting applied, _analytics/_supavisor created, roles' password from
POSTGRES_PASSWORD_FILE); through Kong: REST 200 with the anon key and 401
without, auth health and settings 200, storage buckets 200, GraphQL 200, pg-meta
200, an edge function 200, Studio 401 without and 200 with the dashboard login,
realtime tenant health 200; the pooler in session and transaction mode as
postgres.zurag; a tenant set to host "db" was repointed to supabase-db by the
bootstrap and connections worked.
ace runs a Conduit homeserver (matrix.zurag.be, 5.4 GB of RocksDB, federating)
and Element Web under HAL, configured by environment with the domain
templated in, and Element's config.json carrying matrix.zurag.be literally.
A homeserver's server_name is its permanent identity - every user id, room id
and signature in the database carries it - and it is the name the module is
served under. So it comes from ${bound:route:name-homeserver} (mesh-controller
#149), rendered into a conduit.toml the container reads through CONDUIT_CONFIG,
and into Element's config.json (base_url, default_server_name, the room
directory). Without #149 the render is refused ("name-homeserver"), never
written empty. Two routes, one per endpoint: homeserver (label matrix, 6167)
and element (label element, 80).
Federation needs no 8448: Conduit answers /.well-known/matrix/server with
<name>:443, so peers federate through the route. HAL published 8448 on all
interfaces, but the router never forwarded it; checked from outside, the
well-known, federation version and client versions all answer on 443.
Registration defaults to off. HAL ran with CONDUIT_ALLOW_REGISTRATION=true,
which on ace means anyone on the internet can create an account with the
dummy flow (seen: /register offers m.login.dummy) - on 0.10.13, whose
successor 0.10.14 fixes an account-takeover by any local user. Existing
accounts are unaffected; the operator decides whether to reopen it.
Element's config.json is the one merge:json file (it tolerates `endpoints`),
so a machine can add keys. HAL's map_style_url is not carried: it embedded
a map-tile API key, which belongs in an assignment if wanted.
Images are the digests ace runs (Conduit 0.10.13, Element 1.12.28).
Verified: catalogue tests with MESH_CATALOGUE pointed here on mesh-controller
main and #149; a resolution on #149 renders both names (matrix.zurag.be,
element.zurag.be) into both files and both contributions; throwaway
containers of both pinned digests with the rendered files (root 0644, :ro):
client versions 200, well-known says matrix.zurag.be:443, register refused
M_FORBIDDEN, Element 200 serving the rendered config.json.
The manifest stated /services/redis/data and /var/lib/redis-module - novox's
old layout, paths no definition may carry (ADR 0112). State is now the
assignment's root, grants and data are placed, and the config file, the
secret file, receives and grants all name them as ${dir:...}. Paths inside
the sidecar are its own view and are unchanged.
The data directory and config are owned 999:1000: the image's redis user is
uid 999 in gid 1000 (checked in both builds), which is who owns ace's data
today; 999:999 named a group the image does not use.
Image pinned to the 7.4.11-alpine build ace runs (2026-09-17); the old pin was
the same version, built in August. Older-than-running is never the pin.
Nothing is assigned it anywhere today, so no machine changes.
Verified: catalogue tests pass with MESH_CATALOGUE on this tree; the
declaration composes for ace with every path under /var/lib/redis. The pinned
image ran as a throwaway with a 0600 999:1000 config and a 0700 data dir:
unauthenticated PING is refused (NOAUTH), authenticated SET/GET works,
appendonly is on, the server runs as redis.
The manifest stated /var/lib/mssql, its grants and its SA file by path, and
declared it listens on 4848 - the port one machine's predecessor published,
which is an assignment's fact (ADR 0112, 0138). ace is moving its own
SQL Server (80 GB of work databases) onto the mesh, so the module has to be
the same on every machine.
- state is the assignment's root (place "."), grants is placed, and every
reference (sa.env, the env-file, the SA and grants mounts, receives,
grants, own-secrets) names them as ${dir:...}.
- the database endpoint listens on 1433, the port SQL Server uses; a machine
that must keep an older number pins it in its assignment.
- the image pin is unchanged: it is the digest ace runs today (CU27,
16.0.4295), the same as novox.
novox is untouched: rendered with novox's own setting ({"ports":{"1433":4848}})
through the controller's Declaration and Rules, every resource - paths,
container names, volumes, env-file, the 4848:1433 mapping, owners, modes - is
byte-identical to what main renders; the only difference is the firewall
rule's comment text (still port 4848, from the mesh).
Verified: catalogue tests pass with MESH_CATALOGUE on this tree. The pinned
image ran as a throwaway on a 0700 10001:0 data dir with a root-owned 0600
env-file (dummy SA), answered sqlcmd as sa; a scratch database stopped,
copied with cp -a, checksummed and started on the copy kept its rows and
CHECKSUM_AGG.
GF_SERVER_ROOT_URL=https://grafana.zurag.be, KC_HOSTNAME=https://keycloak.novox.be
and the served issuer's novox default were domains in definitions — wrong on
every other machine (ADR 0112). The names now come from ${bound:route:name}
(mesh-controller #149, hq 122): grafana's in oidc.env, keycloak's in a
hostname.env its server reads. The issuer includes the realm and stays the
assignment's, with no default: unset, a consumer asking for it is refused
and the provisioner says so, rather than both quietly using novox's URL.
Rendered through mesh-controller #149 from these manifests on a zurag.be
node: KC_HOSTNAME=https://keycloak.zurag.be, GF_SERVER_ROOT_URL=
https://grafana.zurag.be, OIDC URLs from the issuer setting. Needs #149
merged and rolled out first.
HAL's grafana logged in through a hand-made Keycloak client whose secret sat
in its .env. Requiring oidc-client gives it a client the mesh makes and keeps:
the id and URLs come from the binding, the secret arrives as a file grafana
reads itself (__FILE), and the callback it contributes is what keycloak
registers as its redirect.
GF_SERVER_ROOT_URL is still a literal: a module cannot yet learn the public
name the mesh composes for its own endpoint (hq issue 122), and without it
grafana sends a redirect Keycloak refuses.