Commit Graph
440 Commits
Author SHA1 Message Date
jschoubben d2f03736fa grafana: its InfluxDB data source comes from the influxdb-api provision
ace's grafana reads InfluxDB through a data source somebody typed into
its database: a LAN address, a database InfluxDB 2 does not have, and a
password for a v1 user of an earlier instance. Nothing in the mesh knew
it existed, so migrating influxdb could only break it further.

grafana now requires influxdb-api, contributes read access, and the mesh
renders a provisioning file grafana reads at start: the address, port,
org's default bucket (as the InfluxQL database) and its own login from
the binding, the password by $__file from the pair credential the mesh
delivers, 0400 for grafana's uid 472. It is a data source of its own
name and uid, read-only in the UI and not the default, so the data
source a person made is never overwritten; a changed binding or a
rotated password restarts grafana, which re-reads the file.

Includes #152 (merged into this branch): influxdb provides influxdb-api.
2026-09-30 13:02:11 +02:00
jschoubben c5e273e232 Merge feat/influxdb-for-ace (#152) into feat/oidc-client-provision
grafana's data source requires influxdb-api, which influxdb provides only
on #152's branch; merged so this branch's catalogue has the provider of
everything grafana requires. #152 should merge first.
2026-09-30 12:55:43 +02:00
jschoubben 323ef9ec7e influxdb: provide influxdb-api, one mesh-made v1 credential per consumer
grafana's data source and Node-RED's influxdb nodes reached ace's
InfluxDB by a LAN IP or a public name nobody routes, with a credential
somebody made by hand. Now a consumer requires influxdb-api and is told
where it is, which org and default bucket it serves, and signs in with
the password the mesh minted for the pair.

The credential is a v1-compatibility authorization, made per grant by
the new provisioner: InfluxDB 2.x generates API tokens itself and
ignores one the caller sends, so a v2 token could only be accepted by
hand per pair; a v1 authorization takes a caller-chosen password (8-72
characters, the mesh mints 40) and reads/writes every bucket as a
database of its name over InfluxQL and line protocol. A consumer
contributes `access` (read, write, read-write) and, for writing, the
buckets; a missing bucket is made and never deleted. Only
authorizations named mesh_* and marked [mesh] are ever changed or
removed; anything else of that name is refused and left alone.

The org and default bucket are served facts the assignment's settings
set, reaching both the consumers and the provisioner's config.json.
2026-09-30 12:55:38 +02:00
jschoubben 5a906b757d keycloak, grafana: their public names come from the mesh, not the manifest
GF_SERVER_ROOT_URL=https://grafana.zurag.be, KC_HOSTNAME=https://keycloak.novox.be
and the served issuer's novox default were domains in definitions — wrong on
every other machine (ADR 0112). The names now come from ${bound:route:name}
(mesh-controller #149, hq 122): grafana's in oidc.env, keycloak's in a
hostname.env its server reads. The issuer includes the realm and stays the
assignment's, with no default: unset, a consumer asking for it is refused
and the provisioner says so, rather than both quietly using novox's URL.

Rendered through mesh-controller #149 from these manifests on a zurag.be
node: KC_HOSTNAME=https://keycloak.zurag.be, GF_SERVER_ROOT_URL=
https://grafana.zurag.be, OIDC URLs from the issuer setting. Needs #149
merged and rolled out first.
2026-09-30 00:49:08 +02:00
jschoubben d8ee88e487 grafana: log in through keycloak's oidc-client provision
HAL's grafana logged in through a hand-made Keycloak client whose secret sat
in its .env. Requiring oidc-client gives it a client the mesh makes and keeps:
the id and URLs come from the binding, the secret arrives as a file grafana
reads itself (__FILE), and the callback it contributes is what keycloak
registers as its redirect.

GF_SERVER_ROOT_URL is still a literal: a module cannot yet learn the public
name the mesh composes for its own endpoint (hq issue 122), and without it
grafana sends a redirect Keycloak refuses.
2026-09-30 00:39:16 +02:00
jschoubben 54557b77bf keycloak: provide oidc-client, one mesh-made client per consumer
A module that logs people in through Keycloak had to be given a client by
hand, with its secret copied into the consumer's environment. As a provision
the mesh derives the client id (the consumer's identity, mesh_<node>_<module>)
and mints its secret, and delivers both ends: keycloak creates exactly that
confidential client, the consumer names it through ${bound:oidc-client:as}.

The consumer says where its browser comes back to (`callback`) and which
endpoint it is reached on (`label`/`endpoint`), so the redirect is built from
the same names the mesh composes for its route. keycloak serves the issuer and
the endpoint paths under it; the issuer is the one value an assignment sets,
and the realm is read out of it, so consumer and client cannot disagree.

Only what the mesh made is touched: its clients carry mesh.provisioned=true;
a client of the same id without the mark is refused, never adopted, updated
or deleted. The runtime now gets the admin password as a file, which its
tools also needed and never had.
2026-09-30 00:39:16 +02:00
jschoubben a5e21cb438 grafana: its directories are placed, its admin password is a file, and it runs the build in use
The module stated /var/lib/grafana-module and /services/grafana/data, a
layout no definition may carry (ADR 0112). State and data are now placed
directories; the admin secret lives beside the broker account under the
mesh's own state.

The admin password reached grafana through an env-file. Grafana honours
GF_SECURITY_ADMIN_PASSWORD__FILE, so it is now a 0400 file owned by the
image's user (472) and mounted, and "secrets-in-environment" is gone
(ADR 0086).

The runtime sidecar was given no credential at all - its config file was
"{}", so GrafanaClient.fromEnv threw and the tools and the alert watcher
did nothing. It now carries user/password from the same secret, and it
calls grafana on the machine port the mesh assigned (${port:3000}) rather
than a literal 3000.

Image pinned to the 13.2.2 build ace's predecessor runs; the old pin was
13.2.1, older than the data it would open.

Verified: catalogue tests with MESH_CATALOGUE pointing here; a throwaway
container of the pinned image with the file-mounted secret answers
/api/health and authenticates admin with the file's value (default
admin/admin refused); restarted over the same data with a different file
value, the original password still holds - so a migrated instance's
password must be accepted, not minted; data owned by another uid fails to
start, so a moved data directory must be chowned to 472.
2026-09-29 23:43:00 +02:00
jschoubben fe0ed3b74e influxdb: place its directories, hand secrets over as files, name its UI
The manifest named /services/influxdb and /var/lib/influxdb-module — one
machine's paths — and passed the admin password and token through the
environment. ace is moving its 2022 instance onto the mesh, so the module
has to be what it is on any machine.

- data, config and state are placed directories; the data keeps 1000:1000,
  the image's influxdb user, which is who owns ace's data today.
- the init secrets reach the image through its own
  DOCKER_INFLUXDB_INIT_{PASSWORD,ADMIN_TOKEN}_FILE; the vault's files are
  mounted read-only. secrets-in-environment is gone.
- the sidecar reads its token from the same file (MESH_INFLUXDB_TOKEN_FILE,
  added to client.ts) and reaches the server at its assigned machine port
  (${port:8086}) instead of assuming 8086. The unused config-dir mount,
  which held the CLI's copy of the admin token, is dropped.
- the api endpoint contributes a route: the web UI is how people use it,
  and reach is the assignment's to say.

Verified: catalogue tests pass with MESH_CATALOGUE pointed at this tree.
The pinned 2.9.1 image, run on a scratch copy of ace's 2.4.0 data, opens
it, runs its metadata migrations (backing up the pre-upgrade bolt/sqlite)
and hashes the two stored tokens; /health passes. A fresh setup through
the _FILE variables, with dummy secrets as root-owned 0600 files, accepts
the token (200 on /api/v2/buckets) and the password (204 on /signin).
client.ts typechecks strict and reads the token file, tolerating the
endpoints key in its config.
2026-09-29 23:42:18 +02:00
mesh-admin 8064e5da8f Merge pull request 'searxng: its settings are a file the mesh writes, not the image's defaults' (#143) from feat/searxng-settings-as-a-file into main 2026-09-29 20:45:13 +00:00
jschoubben 63a255c5cb searxng: bind its route where its state now lives
The route binding still named /var/lib/searxng-module, the directory the
previous commit placed elsewhere — the host would have written it into a
directory nothing declares. Same shape as gitea and nextcloud.
2026-09-29 22:41:56 +02:00
jschoubben 7ad1fbd5c6 searxng: its settings are a file the mesh writes, not the image's defaults
The module ran searxng on the image's built-in settings, which serve html
only — so the module's own search tool (format=json) was refused by the
software it fronts. And there was no way to configure it per machine: the
only file settings reach was the sidecar's.

settings.yml is now the module's one mergeable file (JSON is YAML): generic
defaults in the manifest (json format on, limiter and image proxy off,
valkey wired), and whatever differs per machine — base_url, method,
autocomplete, suspended times — set as the assignment's settings. The
secret key is filled on the machine through ${secret:secret}, so the
secrets-in-environment exception and the env file go. Directories are
placed. Image pinned to 2026.9.20, what ace runs today (the old pin was
older, 2026.9.1).

The sidecar's config.json is no longer mergeable: settings merge into every
mergeable file of a module, and the sidecar would have received searxng's
keys. It only ever read an optional url, which its env already carries.

Verified on ace: the pinned image serves html and json from a read-only,
root-owned 0600 JSON settings.yml.
2026-09-29 22:33:38 +02:00
jschoubben 67f5f4cffd Merge pull request 'ca-trust: a machine trusts the mesh's authority because a module put its root there' (#142) from feat/ca-trust into main 2026-09-29 14:06:36 +00:00
jschoubben 8797335fbc ca-trust: a machine trusts the mesh's authority because a module put its root there
novox/hq ADR 0147, issue 129. Every internal HTTPS name fails verification
on every machine: the certificates are genuine and nothing on a machine has
ever been told what issued them. The proxy's fetch answers for the proxy and
for nothing else — a browser, git over HTTPS and every module calling another
by an internal name read the machine's own trust store.

The module requires internal-acme-ca, fetches the root over the mesh's own
network (no prior trust to have; that is what this establishes), installs it
among the machine's anchors and refreshes the extracted bundles. Being
unassigned stops the unit, and stopping it takes the anchor away and
refreshes them again.

Arch's layout is named out loud: a machine that keeps anchors elsewhere fails
visibly rather than writing a file nothing reads.
2026-09-29 15:07:40 +02:00
mesh-admin 53dc108603 Merge pull request 'Remove the network-checker module: it does not do what was decided' (#141) from chore/remove-the-network-checker-module into main 2026-09-29 12:43:34 +00:00
jschoubben ebf5ba2d4c Remove the network-checker module: it does not do what was decided
What was in the catalogue was the first thing I built, not the thing ADR 0146
describes. It dialled raw ports on machine addresses from one hosting form and
emitted nothing, so findings would have sat in a file on the machine — the exact
thing issue 145 is about. It was never registered, never assigned, and never ran.

0146 says names per hosting form, fetched over TLS with the certificate verified,
and machines discovered over the bus. That shares nothing with this but the word
checker, so it goes rather than being bent into shape. Recorded as work to be
analysed and built deliberately.

Connectivity is checked by hand in the meantime, against the services the mesh
already runs.
2026-09-29 14:43:25 +02:00
mesh-admin bbac08a7d2 Merge pull request 'A network-checker module: dial what the mesh claims, from where the callers are' (#140) from feat/a-network-checker-module into main 2026-09-29 11:44:37 +00:00
jschoubben 784a5a6514 A network-checker module: dial what the mesh claims, from where the callers are
The mesh asserts three things are callable (ADR 0144) — what runs on the same
machine, another machine's service exposed to the private network, and another
machine's service exposed publicly — and has never checked any of them. The first
was broken for eleven hours while the mesh reported every machine healthy.

This runs on every machine, on the cadence the mesh already has, in its own
container: the same position every other module calls from. Not the host and not
the control plane, both of which reach these addresses by paths no ordinary caller
uses and would have passed throughout that outage.

**Its probe is its own endpoint, and that is the point.** Declared reachable over
the private network like any other service, so it is admitted by exactly the rule
that governs every internally-exposed service and fails when that rule is wrong.
The tempting target is a service every machine has, and those are the ones never
closed — ssh above all — which would have passed while the thing that actually
broke was a service exposed to the private network.

It resolves before it dials and says which failed, because a name that does not
resolve and a port that does not answer have different owners. One failure is not
a fault: a machine rebooting is ordinary, so a path is broken after consecutive
runs and the count travels with the result. It reports and repairs nothing.

novox/hq ADR 0145. Eight tests; the consecutive-failure logic proved by reverting
it once. Not yet registered or assigned.
2026-09-29 13:44:16 +02:00
mesh-admin 0c31499fb0 Merge pull request 'Every module names its endpoints, and every route names the one it serves' (#139) from feat/modules-name-their-endpoints into main 2026-09-29 09:51:56 +00:00
jschoubben f118344246 Every module names its endpoints, and every route names the one it serves
75 endpoints across 50 modules, named from what each one is for rather than by a
rule: mail's seven protocol ports are smtp, imaps, submission and the rest; unifi's
nine are inform, stun, discovery, the two portal ports and syslog; minio's two are
s3 and console; the resolver's two are dns-udp and dns-tcp.

And 35 route contributions name the endpoint they serve instead of repeating its
port. A route and a listen both carried a port and nothing said they were the same
thing; now one of them does. gitea's path-level deny rule names neither, because it
is a rule about a name rather than an endpoint.

novox/hq ADR 0138. The words shipped a release ahead in mesh-controller #138 and
#139, and the control plane running today is built from that merge — checked before
this was written, because an unknown manifest key is refused and a catalogue using
one against an older control plane would stop resolving.
2026-09-29 11:51:39 +02:00
mesh-admin 822df220ab Merge pull request 'A routed module listens from the mesh, not from anywhere' (#138) from fix/a-routed-module-listens-from-the-mesh into main 2026-09-29 00:58:39 +00:00
jschoubben 9eb1265bc8 A routed module listens from the mesh, not from anywhere
umami declared its port reachable from anywhere, reasoning that the collection
endpoint tracked browsers POST to must be public. That is true of the name and
not of the port: both its surfaces are served through the proxy by name, so the
port is how the proxy reaches it and nothing else (ADR 0045).

Measured, which is how this was found: with the port open to the internet, the
dashboard's login page was served over plain HTTP directly on the machine's port,
bypassing every rule the proxy applies by path. The route stays exactly as it was,
so the collection endpoint keeps working.
2026-09-29 02:54:10 +02:00
mesh-admin 41cfc70b53 Merge pull request 'The resolver declares both protocols it answers on' (#137) from fix/the-resolver-declares-both-protocols into main 2026-09-28 22:04:19 +00:00
jschoubben acedc5d9d9 The resolver declares both protocols it answers on
It declared udp/53 only. The daemon listens on tcp/53 as well, and a resolver is
asked over tcp whenever an answer will not fit in a datagram — so on every
converged machine that port is closed while the service reports itself healthy
and the manifest reads as though the resolver were fully declared.

The same fault as issue 136 in miniature: the declaration covers part of what the
service does, and the gap is silent because nothing compares the two.
2026-09-28 23:53:03 +02:00
mesh-admin 521a8dd1e2 Merge pull request 'sshd: the daemon it owns starts at boot' (#136) from fix/sshd-declares-the-daemon-it-owns into main 2026-09-28 19:43:29 +00:00
jschoubben e145e2236c sshd: the daemon it owns starts at boot
The module said the service must be running and nothing about boot, so the
machine's own way back in was enabled only because something before the mesh
had enabled it. All four machines happen to be enabled today; none of them is
enabled because the mesh says so, and a machine adopted tomorrow would run ssh
until its first reboot.

Not `state: running` alone for the same reason the module exists: this is the
one daemon whose absence cannot be fixed remotely.
2026-09-28 21:43:27 +02:00
mesh-admin 4d7e37e319 Merge pull request 'fail2ban bans through an action every machine has' (#135) from fix/fail2ban-bans-through-what-every-machine-has into main 2026-09-28 18:48:26 +00:00
jschoubben 026421fd6e fail2ban: ban through an action every machine has
jail.local named ufw as the ban action. Two machines on this mesh have no ufw,
and fail2ban does not check: it starts, the jail reads the log, counts the
attempts, runs the ban command, gets 127 -- 'ufw: command not found' -- and
logs an error nobody reads. The service is active, the mesh reports the module
applied, and the machine is not protected. Proven by banning a documentation
address on such a machine today.

The replacement is this module's own dualchain action, already used by the
recidive jail on all four machines, so it is not a new dependency. It bans in
DOCKER-USER as well as INPUT, which ufw's action did not, and it bans all
ports, which ufw's action did.
2026-09-28 20:48:24 +02:00
mesh-admin af89bb11ff Merge pull request 'fail2ban declares the log its own recidive jail reads' (#134) from fix/fail2ban-declares-the-log-its-own-jail-reads into main 2026-09-28 18:45:34 +00:00
jschoubben 7c18cdbd39 fail2ban: declare the log its own recidive jail reads
The recidive jail bans whoever keeps coming back by reading fail2ban's own
log, and fail2ban checks every jail's log file while it configures itself --
before it has created that log. On a machine where the file is not there
already, no jail is found for recidive, configuration fails, and the whole
service refuses to start, taking the sshd jail with it. Two machines assigned
this module today came up failed for exactly that reason; the two where it
worked had a log from years of the service running.

Declared create-once: the mesh puts an empty file there when it is absent and
never touches it again, because what grows in it is fail2ban's, and the
logrotate file this module already ships is what keeps it small.

This also reverts the previous two commits' fail2ban.local. It declared a
logtarget that the package already sets to the same path on every machine
here -- pacman reports the config pristine -- so it fixed nothing and said
something untrue about why.
2026-09-28 20:45:32 +02:00
mesh-admin 4fb16b2e6b Merge pull request 'fail2ban restarts when the log declaration changes' (#133) from fix/fail2ban-restarts-on-its-log-target into main 2026-09-28 18:42:44 +00:00
jschoubben c5af8635c8 fail2ban: restart when the log declaration changes
The file that says where fail2ban logs was not in restart-on, so a change to
it would sit on disk with the running service unaware of it -- the same shape
as any other jail file this module already restarts for.
2026-09-28 20:42:42 +02:00
mesh-admin 87366c5f36 Merge pull request 'fail2ban declares where it logs, so the recidive jail has a file to read' (#132) from fix/fail2ban-declares-where-it-logs into main 2026-09-28 18:41:16 +00:00
jschoubben f8ca36aacf fail2ban: declare where it logs, so the recidive jail has a file to read
The recidive jail reads /var/log/fail2ban.log and this module ships the
logrotate file for it, but nothing ever told fail2ban to write there. Where
the package default stands, fail2ban logs to the journal, the recidive jail
finds no log file, and the whole service refuses to start -- taking the sshd
jail with it. Two machines assigned this module today came up failed; the two
where it worked had /etc/fail2ban/fail2ban.conf edited by hand, which a
package upgrade would have undone.

Declared in fail2ban.local, because fail2ban.conf belongs to the package.
2026-09-28 20:41:09 +02:00
mesh-admin 812355bf31 Merge pull request 'The catalogue hears what it missed' (#131) from feat/the-catalogue-hears-what-it-missed into main 2026-09-28 14:08:48 +00:00
jschoubben 016ddb2b3a The catalogue hears what it missed
It asks what it missed on every start and the answer never arrived: the control plane replayed each
build it held as a module's event from a module called "control-plane", which does not exist, so its
own account refused the publish and the graph kept the gap. The control plane now states those under
the seat it holds (novox/hq ADR 0134, mesh-controller #129), so this consumes that too — one handler,
because what a build means for the graph is the same whether the build machine says it as it happens
or the mesh says what it already held.
2026-09-28 16:08:46 +02:00
mesh-admin ea17bf46d2 Merge pull request 'The catalogue prepares its own schema instead of migrating at start' (#130) from feat/the-catalogue-prepares-its-own-schema into main 2026-09-28 13:40:30 +00:00
jschoubben 4258f01614 The catalogue prepares its own schema instead of migrating at start
It brought its schema up inside its runtime, on every start. That made a schema it could not reach a
crash loop rather than a stop, with the module graph keeping a gap and nothing saying so — which is
how a whole morning's builds went unrecorded. The mesh now prepares this module's state before it
starts this version and does not start it if that failed (novox/hq ADR 0135): the work moves to an
entrypoint the image names in MESH_PREPARE, beside the entrypoints it already names.

The reason it was at start — that a step blocking the apply would block the very apply bringing the
overlay up — stopped being true when a step's failure became its module's business rather than the
machine's (ADR 0136).
2026-09-28 15:40:28 +02:00
mesh-admin 4d9b4fdfa6 Merge pull request 'The catalogue declares the event it emits on starting' (#129) from fix/the-catalogue-declares-the-event-it-emits into main 2026-09-28 07:49:06 +00:00
jschoubben eff11b1d4d The catalogue declares the event it emits on starting
Its runtime announces that it has just started and may have missed builds — the event the control
plane follows to replay them — and its manifest did not declare it. A module's authority on the bus
is derived from what it declares, so the publish was refused and the runtime died on start, in a
loop, with the mesh's graph never catching up.
2026-09-28 09:49:03 +02:00
mesh-admin 4ead13d4d4 Merge pull request 'A merge says which files it changed' (#128) from feat/a-merge-rebuilds-what-it-changed into main 2026-09-28 07:20:05 +00:00
jschoubben 3b77dde666 A merge says which files it changed
Every module built from a repository was rebuilt for a change to any of them: one merge in this
repository meant twenty-six builds, which is what exhausted a public registry's pull limit. The
forge lists the files a merge changed and the event carries them, from the watcher and from the
merge tool alike; a merge that changed more files than were asked for says so, and the mesh then
treats the whole repository as changed rather than guessing.
2026-09-28 09:17:38 +02:00
mesh-admin 5ea4961980 Merge pull request 'A merge older than the watching is history, not news' (#127) from fix/a-merge-older-than-the-watching-is-history into main 2026-09-28 03:08:22 +00:00
jschoubben 2b8a668d06 A merge older than the watching is history, not news
An old merge past the first page of the forge's listing surfaced as newer pull requests were
updated, and was announced as if it had just happened; the mesh then rebuilt everything built from
that repository, once per old merge. The moment the watching began is kept with the record, and
only a merge made since is announced.
2026-09-28 05:08:20 +02:00
mesh-admin 719fb1e025 Merge pull request 'A merge the forge announces is said in its log' (#126) from fix/a-merge-announced-is-said into main 2026-09-28 02:44:19 +00:00
jschoubben ac5630bee2 A merge the forge announces is said in its log
A trigger that fires silently is indistinguishable from one that did not fire (novox/hq issue
131); the announcement is now one line an operator can read.
2026-09-28 04:44:17 +02:00
mesh-admin 1c995fa9fc Merge pull request 'The forge watches every repository, not the administrator's own' (#125) from fix/the-forge-watches-every-repository into main 2026-09-28 02:31:25 +00:00
jschoubben d70cb18ea0 The forge watches every repository, not the administrator's own
/user/repos lists what the token's user owns, which for the mesh's administrator is nothing — so
the forge module watched an empty list and never announced a merge. It reads the forge's whole
view through the search endpoint, every page.
2026-09-28 04:31:20 +02:00
mesh-admin 1d71787896 Merge pull request 'The forge announces every merge, whoever made it' (#124) from feat/the-forge-announces-every-merge into main 2026-09-28 01:03:48 +00:00
jschoubben eb62289f89 The forge announces every merge, whoever made it
The merge tool emitted at the instant it acted; a merge made in the forge's own
pages or over its API emitted nothing, and the mesh went on believing every module
current with its source (novox/hq 04-ISSUES/131). Merged pull requests are now
watched the way repositories are: what the forge holds, asked for on a tick,
announced once, with the merge commit and the clone URL a build needs. What has
been announced is kept beside the module's state, so a restart does not announce
the whole history again, and a first tick with no record announces nothing.
2026-09-28 02:54:48 +02:00
jschoubben f5969a2f9f Merge pull request 'nats declares the certificate directory it mounts' (#123) from feat/nats-serves-the-meshs-certificate into main 2026-09-27 22:31:20 +00:00