Commit Graph
100 Commits
Author SHA1 Message Date
jschoubben 8a458aa5c4 Merge pull request 'The forge mints its own API token with the admin account the vault delivers' (#49) from fix/gitea-mints-its-token into main 2026-09-24 08:41:13 +00:00
jschoubben d036b43bfc Merge main into fix/gitea-mints-its-token to pick up the resolver conversion and artifact-store seat fix 2026-09-24 10:40:57 +02:00
jschoubben 5a94b78891 Merge pull request 'Convert the resolver from the module it replaces: forward, answer at 127.0.0.1, point the runtime at it' (#50) from convert/dnsmasq-from-hal into main 2026-09-23 23:13:08 +00:00
jschoubben 3d81bf41c6 Convert the resolver from the module it replaces: forward, answer at 127.0.0.1, point the runtime at it
Read against hal/modules/dnsmasq-app (hal dnsmasq-app conversion, hq 08-connectivity).
On the machines it runs, the predecessor's dnsmasq answers every name: the mesh's own
itself, the rest forwarded to 1.1.1.1 and 8.8.8.8, its module's defaults; resolv.conf
names it alone at 127.0.0.1, and the container runtime's dns is the machine's tunnel
address, so the host and every container resolve the world through it. The nox module
forwarded nothing, listened on 127.0.0.55 — a convention of its own beside the one every
machine already followed — and read a machines file that the module's `facts` already
asks the mesh for, so taking it would have left an adopted machine with a resolv.conf
pointing at an address nothing answered on, and no upstream for anything else.

Now the resolver keeps `no-resolv` (the documented loop — finding its own address in
resolv.conf and becoming its own upstream — stays impossible) and forwards to the same
two explicit upstreams; listens on mesh0 and 127.0.0.1, which systemd-resolved does not
hold; requires `mesh-addressing`, since its data is the mesh's addresses; and writes the
runtime's `dns` into daemon.json beside whatever the machine had (ADR 0102), at this
machine's own address — `${machine:address}`, new in the controller. The runtime is not
restarted for it: it reads the key at start, not on reload, and a restart stops every
container; on the machine this replaces the value is already there.

resolv-conf names the resolver alone, as the predecessor's file did; its placeholder second
line was a fallback nothing ever reached. resolved-split-dns follows the address. The
mesh's suffix as a local domain comes with the machines file, so a mesh name the resolver
does not know is refused here rather than asked upstream. mDNS is not carried: no module
does it and the design says mesh names are not multicast names.
2026-09-24 01:09:23 +02:00
jschoubben ed094031fd The forge mints its own API token with the admin account the vault delivers
The runtime beside the forge served 0 tools: GiteaClient.fromEnv required a token
(settings or MESH_GITEA_TOKEN), nobody had one to give — the mesh raised the forge —
and putting one in settings would store a secret in plaintext in the inventory. So
the fifteen tools registered nothing and the watcher logged "not watching".

What the mesh does deliver is the admin account: a login the manifest names and a
password the vault minted and the host unsealed into a file (ADR 0086). That is
enough to mint a token, so the module does (hq issue 100, the forge's tools):
POST /users/{admin}/tokens over basic auth, scoped to write:repository and
write:issue — the least the tools and the repo watcher need — kept at 0600 in the
module's own state (/var/lib/mesh/gitea/state, a new directory resource the runtime
mounts writable), read back on the next start, and minted afresh when the forge
answers 401 to it or the kept file is gone. A forge whose data came from the
predecessor has no mesh-admin: that is reported in plain words on every poll until
it clears, once per reason, not crash-looped. A configured token still wins and is
never minted over.

The mint happens on the first call, not at registration: a contributor is
synchronous, and a forge not yet answering must not keep the runtime from serving.
One source per kept file in a process, or the watcher and the tools would each
renew on a 401 and drop the other's token by name.
2026-09-23 23:48:36 +02:00
jschoubben 4d8be01c5f Merge pull request 'The artifact store's seat is one per mesh' (#48) from fix/the-artifact-store-is-one-per-mesh into main 2026-09-23 21:42:25 +00:00
jschoubben 833d8ec3ec The artifact store's seat is one per mesh
Review of the registry work found it node-scoped: a second `distribution` on another
machine resolved cleanly there, and only afterwards did the mesh notice `artifact-store`
offered by two nodes, with every consumer elsewhere refusing to choose. Worse, a node-scoped
requirement with one candidate installs that candidate, so anything that wanted the store
beside it would have raised a fresh, empty store on the wrong machine first.

There is one store in a mesh, which was already the effective rule; the claim now says it
where a second one is assigned, by name, instead of leaving consumers to discover it.
2026-09-23 23:40:54 +02:00
jschoubben 9f2c678355 Merge pull request 'Convert mssql from the module it replaces, not from a blank page' (#46) from convert/mssql-from-hal into main 2026-09-23 18:22:51 +00:00
jschoubben cae0d00a7d Convert mssql from the module it replaces, not from a blank page
Read against hal/modules/mssql: the predecessor publishes 4848:1433 and its connections
block hands every consumer localhost:4848; its data has lived in /services/mssql/data
since it was installed — 4.6 GB of it. The nox module had 1433, db-data, and a pin three
builds behind, so taking it would have moved the port every consumer was told about,
started on an empty directory, and downgraded the engine at the same time.

db-data was not a convention: only this module and mongodb used it, and both predecessors
say data. Now nothing moves at the cutover.
2026-09-23 20:22:34 +02:00
jschoubben c001e054fd Merge pull request 'Pin the analytics module to the version the machine it replaces is running' (#45) from fix/umami-matches-production into main 2026-09-23 00:48:13 +00:00
jschoubben b69edbb16f Pin the analytics module to the version the machine it replaces is running
The pin was a month old — 2026-08-20 against the 2026-09-17 image the container runs.
Third module in a row whose pin had aged into a downgrade; the pattern is hq issue 099.
2026-09-23 02:47:23 +02:00
jschoubben b0d48d5c54 Merge pull request 'Pin the package registry to the version the machine it replaces is running' (#44) from fix/verdaccio-matches-production into main 2026-09-23 00:37:31 +00:00
jschoubben a3eca8f4e6 Pin the package registry to the version the machine it replaces is running: 6.10.3, not 6.10.1
The rule the forge's cutover taught: a module that takes over a running service must not
carry an older image than the one running, or the cutover is a downgrade nobody asked for.
6.10.4 exists; moving to it is an upgrade and its own act.
2026-09-23 02:24:07 +02:00
jschoubben f11534102d Merge pull request 'Pin the forge to the version the machine it replaces is running' (#43) from fix/gitea-matches-production into main 2026-09-23 00:28:26 +02:00
jschoubben a29475617a Pin the forge to the version the machine it replaces is running: 1.27.3, not 1.22.6 2026-09-23 00:28:11 +02:00
jschoubben 176bbd6085 Merge pull request 'route-adapter: provide route by writing into the predecessor's proxy (hq ADR 0104)' (#42) from feat/route-adapter into main 2026-09-23 00:12:48 +02:00
jschoubben bd5b349a0d route-adapter: provide route by writing into the predecessor's proxy (hq ADR 0104)
A node being adopted cannot take a web module: every module reachable by
name requires route, the mesh's only provider of it binds the two public
ports, and the predecessor's proxy holds them and serves every public
name there. Stopping the predecessor to break the circle darkens every
name at once, with every certificate to re-obtain in the same window.

So this answers the same provision without binding anything. It provides
route and receives the same contributions file, and writes each
contribution as one route file where the predecessor's file provider
reads, naming the predecessor's own certificate resolver so no
certificate is asked for. It removes a file it wrote when its
contribution goes and never touches a file it did not write — the name
and a marker inside both have to say it is the mesh's.

A step, not a daemon: run-once, re-run by restart-on over the received
file and the settings. The predecessor's dynamic directory is a node
setting, because it is a fact about one machine.

Migration scaffolding with a stated end: assigned only on an adopted
node, deleted when the predecessor's proxy retires.
2026-09-23 00:11:21 +02:00
jschoubben 4f4a0750de Merge pull request 'The store keeps the vector extension the predecessor's database had' (#41) from fix/store-keeps-pgvector into main 2026-09-22 23:30:34 +02:00
jschoubben 5eb1baf5fe Keep the vector extension the predecessor's database had: the store runs the pgvector build of the same major 2026-09-22 23:30:21 +02:00
jschoubben d0fed5b982 Merge pull request 'The forge's own address follows the port the node gave it (hq issue 088)' (#40) from feat/forge-address into main 2026-09-22 23:29:47 +02:00
jschoubben cf30d4b3d0 The forge's sidecar dials the port the machine put the forge on (hq issue 088)
MESH_GITEA_URL named 3000 outright. The forge's container publishes 3000 without
fixing the machine side, so the mesh assigns it — and on a node given that port
as a setting it is the operator's number. The mapping that lets the forge go on
binding 3000 does nothing for a caller dialling the machine's loopback, so the
sidecar dialled a port nothing was listening on wherever the two differed.

${port:3000} is the mesh's answer to exactly that question, and it now resolves
in a container's environment as it always has in a file. The forge declares it
listens on 3000, so it may ask.

Still names 3000: the route contribution (hq issue 089) and the `listens` and
`ports` entries, which are the software's own number and belong there.
2026-09-22 22:36:34 +02:00
jschoubben 5706c00566 Merge pull request 'The builder's package binding is settable per node (hq issue 085)' (#39) from feat/packages-port into main 2026-09-22 21:59:57 +02:00
jschoubben b2e29871fd Protect the address the builder sends its registry password to
`at` was settable. A node setting could point the builder's package binding at
any host, and the builder sends its registry credential there as basic auth — so
a setting meant for a port was a way to hand the password to somebody else.

novox/hq 04-ISSUES/085
2026-09-22 21:56:51 +02:00
jschoubben d63f006cb8 The builder's package binding takes its port from the node, not from the manifest
The forge is raised by hand at genesis, before any module provides
`package-registry`, so the builder carries a binding instead of resolving one —
and the port in it was rewritten, as text, by the installer. A later
registration of this manifest from the catalogue put 3000 back, silently, and
pointed the builder and its registry credential at whatever holds that port.

Made settable instead: the port is a per-node setting the controller holds, and
the catalogue's number is only its default. `provision`, `from` and `as` are
protected — a setting here is about where the forge answers, never about who the
binding is with.

novox/hq 04-ISSUES/085, ADR 0100
2026-09-22 21:39:54 +02:00
jschoubben 6e88982498 Merge pull request 'Adoption mode: guards, and a filter unit that never flushes the ruleset (hq ADR 0100, 0103)' (#38) from feat/adoption-mode into main 2026-09-22 21:01:56 +02:00
jschoubben 90104d0818 Reload the filter on a rule change rather than restart it, so the node is never unfiltered in between (hq ADR 0102) 2026-09-22 19:47:43 +02:00
jschoubben 84012fab2e Make the stock nftables unit's stop delete only the mesh's table on nodes that still have it enabled (hq ADR 0100) 2026-09-22 18:06:15 +02:00
jschoubben 0c37d7389d Guard the store and management ports on adopted nodes, and load the filter through a unit that never flushes the ruleset (hq ADR 0100) 2026-09-22 17:32:32 +02:00
jschoubben f4097b3c57 Merge pull request 'n8n and baserow no longer take the shared cache (issue 081)' (#37) from multiple-fixes into main 2026-09-22 13:53:14 +02:00
jschoubben 431627c510 baserow: its data directory is 0755, as the image ships it — the cache it runs for itself does so as another user, who must reach its own directory beneath (issue 081) 2026-09-22 13:46:01 +02:00
jschoubben 93634041db baserow: its texts no longer name the shared cache it stopped taking 2026-09-22 12:34:58 +02:00
jschoubben 24cbac816d n8n and baserow no longer take the shared cache: neither can keep its keys and channels under the login it is granted — baserow runs its own, n8n in its shipped mode needs none (novox/hq issue 081) 2026-09-22 12:26:33 +02:00
jschoubben b2a48558ea Merge pull request 'redis: a consumer's ACL loses the dangerous category (issue 080)' (#36) from multiple-fixes into main 2026-09-22 02:19:14 +02:00
jschoubben c71bbd497f redis: INFO is allowed back — client libraries ask it at connect, and it reads nothing a consumer keeps 2026-09-22 02:03:23 +02:00
jschoubben 421f4d8577 redis: a consumer's ACL loses the dangerous category — a key pattern does not confine FLUSHALL (novox/hq issue 080) 2026-09-22 01:39:29 +02:00
jschoubben 54af5e8536 Merge pull request 'route-proxy: the trust gate names the binding it reads; the server names the gate (ADR 0099)' (#35) from multiple-fixes into main 2026-09-21 23:58:48 +02:00
jschoubben 6fb2003b8d route-proxy: the trust gate names the binding it reads and runs again when the authority moved; the server follows it (ADR 0099) 2026-09-21 23:25:06 +02:00
jschoubben 64d62978f6 Merge pull request 'Multiple fixes: five modules keep their secrets from the vault (ADR 0094), the authority makes its own root (076), the builder on the host network, route-proxy declares its bases' (#34) from multiple-fixes into main 2026-09-21 22:58:24 +02:00
jschoubben ac651c7ac2 route-proxy: the trust container names its artifact 2026-09-21 22:55:35 +02:00
jschoubben 1c964cd571 route-proxy: the trust gate retries with a timeout and checks for a certificate; its images are declared artifacts (ADRs 0096, 0097, 0098) 2026-09-21 22:53:29 +02:00
jschoubben 28b8feebfc The authority makes its own root at first start, and the proxy fetches it through a gate
step-ca's root certificate, its key and that key's password were own secrets — random
bytes the mesh minted, which no certificate is (novox/hq 04-ISSUES/076). The mesh
mints only the CA password now; step-ca makes its root at first start and serves it
at /roots.pem, which the manifest now names beside the ACME directory. The route
proxy fetches that root over the mesh network in a run-once step before it starts,
instead of being handed a served fact that could only be written before anything ran
(ADR 0098).
2026-09-21 22:26:28 +02:00
jschoubben 82256fcdb0 The builder runs on the machine's network: it copies images into the mesh's registry itself now
The copy between registries (ADR 0096) reaches the mesh's registry over HTTP from
inside the builder's container, where loopback on the default bridge is not the
machine; docker push never noticed because it went through the machine's daemon.
The builder already holds the runtime's socket, so the host network adds nothing
it did not have.
2026-09-21 22:24:43 +02:00
jschoubben facd41806d Five modules keep their own secrets from the vault, under local names; route-proxy declares its bases
gitea, umami, influxdb, icecast and mailu require a secret and keep each of theirs
under a local name (novox/hq ADR 0094); the broker account stays their own. The
route proxy's recipe starts FROM the bases its manifest declares (ADR 0097).
2026-09-21 22:16:10 +02:00
jschoubben 126969a829 Merge pull request 'The controller's manifest lives in the controller's repository, not here (hq issue 072)' (#33) from feat/one-controller-manifest into main 2026-09-21 19:23:18 +02:00
jschoubben a514d9827c The controller's manifest lives in the controller's repository, not here
ADR 0069 put it there; genesis now takes it from the build it runs (mesh-host,
novox/hq 04-ISSUES/072). This copy was read by nothing else and had already drifted.
2026-09-21 15:17:47 +02:00
jschoubben e0d34723f7 Merge pull request 'Six modules take their secrets from files; the rest say precisely why not (issue 041)' (#32) from feat/migration-blockers into main 2026-09-21 13:48:46 +02:00
jschoubben 1289510538 mongodb names its secrets' owner: the image drops to its own user before it reads the password file
The official entrypoint re-executes itself as mongodb (uid 999) and only then reads
MONGO_INITDB_ROOT_PASSWORD_FILE, so a root-owned 0600 file is 'Permission denied' at line 83 and
the server never starts. secrets-owner is the mechanism ADR 0086 gives for exactly this.
2026-09-21 13:43:41 +02:00
jschoubben 50b639ff52 The env-file references to files the conversion removed go with them 2026-09-21 12:43:41 +02:00
jschoubben 0b2f3bfbd3 grafana stays a declared exception until a bed exercises its admin password 2026-09-21 12:31:59 +02:00
jschoubben 1a28e5aec6 Six modules take their secrets from files; the rest say precisely why not
From the survey of every env-file secret (ADR 0086, issue 041): amqp-ping,
minio, mongodb and grafana use the _FILE twin their software honours;
mesh-catalog and model-usage read DATABASE_URL_FILE (a file the mesh
templates, mounted where only the runtime reads it); grafana's secret files
belong to its own account. Two dead deliveries removed: a line nothing read
in amqp-email-forwarder, and mailu's secret.env on four containers that
never read it. The 25 exceptions that remain carry the surveyed reason —
convertible and awaiting a bed, convertible through a generated config file,
the application's own code, or not convertible.
2026-09-21 12:29:25 +02:00
jschoubben db597bcb71 Merge pull request 'The controller reads its credentials from files; every other env-file secret says why' (#31) from feat/secret-not-in-environment into main 2026-09-21 11:59:08 +02:00
jschoubben 8647ea6671 The control plane's secret files belong to its own account (secrets-owner) 2026-09-21 10:19:00 +02:00
jschoubben 32dec5f0c2 The controller reads its credentials from files; every other env-file secret says why
ADR 0086. mesh-controller mounts its six own secrets and names them with
_FILE twins, so no credential of its own reaches its environment. The 35
containers that still read a secret through an env-file carry
secrets-in-environment with the reason; converting each where its software
accepts a path is the per-module work of issue 041.
2026-09-21 10:10:33 +02:00
jschoubben d03520f4ed Merge pull request 'Add mesh-vault; redis, postgres and lavinmq take their passwords from files' (#30) from feat/secrets-vault into main 2026-09-21 10:03:12 +02:00
jschoubben 82e513a360 Add the mesh-vault module; redis takes its password from it
mesh-vault provides `secret` (novox/hq ADR 0085, design 24). The value is
the pair credential the controller mints — the vault holds no copy, only a
ledger of who holds one, its fingerprint and every rotation, and two tools that
answer by fingerprint and never by value. Rotation is `rotate secret`,
unchanged machinery pointed at a secret with an owner (design 13). Named in the
mesh's own namespace, beside mesh-controller and mesh-catalog, because it is
the mesh's own code rather than wrapped software.

redis is the first consumer: its own password stops being an own-secret nothing
could rotate and becomes a `secret` it requires, read from the same file into
the same hole. The server now restarts on its config, or it would keep the
password it started with through every rotation (playbook 06).
2026-09-21 00:48:13 +02:00
jschoubben 8105ab6141 Merge pull request 'minio: pull image from quay.io (docker.io denies anonymous pulls)' (#29) from fix/minio-pull-from-quay into main 2026-09-20 22:03:50 +02:00
jschoubben 2ef7eb2a27 minio: pull image from quay.io (docker.io denies anonymous pulls)
Same digest, a registry that serves anonymous pulls. Unblocks raising minio in
the lab and the object-store cutover.
2026-09-20 22:01:20 +02:00
jschoubben 4ce9de6f3a Merge pull request 'Correct lavinmq Dockerfile's stale MESH_TOOL_MODULES comment (061 review)' (#28) from fix/lavinmq-dockerfile-comment into main 2026-09-20 13:42:03 +02:00
jschoubben 17243b72df Correct lavinmq Dockerfile's stale MESH_TOOL_MODULES comment (061 review)
The comment still said the provisioner is not listed and runs via a
container's args — but the 061 fix put it in MESH_TOOL_MODULES (serve
mode) and dropped the args. A future editor trusting the comment could
strip it again and silently reintroduce 061. Comment now matches the
code; only the run-once bootstrap runs via args.
2026-09-20 13:32:47 +02:00
jschoubben 43c9c973e2 Merge pull request 'The mesh builds its own catalogue (issue 060)' (#27) from feat/the-mesh-builds-its-catalogue into main 2026-09-20 12:53:04 +02:00
jschoubben 965c58fe44 Defer minio from the buildable set too (issue 060)
minio's runtime copies the `mc` client from minio/mc:latest — a
Docker Hub pull the mesh build environment cannot make (its docker
reaches the mesh registry, not public Hub), the same isolation that
blocks npm deps. apt-based installs (mongodb's mongosh, mosquitto)
build fine because the build has real internet for apt; only npm and
Docker Hub are redirected. Delivering an external binary or image layer
into a mesh build is the same open question as the npm deps — deferred
with them.
2026-09-18 02:51:09 +02:00
jschoubben b8d390cbae Defer model-usage and anthropic-manager from the buildable set (issue 060)
Both carry third-party runtime deps (pg; tweetnacl + sealedbox) that
their Dockerfile installs with npm — which 404s in a mesh build, whose
npm points at the mesh's own registry, not public npm. The workstation
build script got away with it by installing on a host with public npm.
Delivering a module's third-party deps into a mesh build is an open
question (how: publish to the mesh registry, or proxy); until it is
answered these two stay on the placeholder path they were already on.
The other 37 modules build from their own directory with no external
fetch.
2026-09-18 02:48:09 +02:00
jschoubben e362fb951c The rest of the catalogue becomes mesh-buildable (issue 060, batch 2)
The 21 media/home modules, the SaaS tool modules (cloudflare-dns,
confluence, gitlab, jira), model-usage, and the model-access trio get
the same Dockerfile + build section as batch 1. Scheduled-only modules
(anthropic-consumer, anthropic-manager, openai-consumer) deliberately
declare no MESH_TOOL_MODULES — every container of theirs names its
command. mosquitto's run-once bootstrap container builds from the same
artifact. Modules with third-party deps (model-usage: pg;
anthropic-manager: tweetnacl) install them beside their compiled code.

Also fixes cloudflare-dns's package.json, unparseable since its
description lost a closing quote.

Deliberately still without build sections: builder and mesh-controller
(the foundation builds them by its own path), distribution (provides
the artifact store — building it through itself is refused by design),
route-proxy (cross-repo build context, deferred), and the
upstream-image-only modules, which have no code to build.
2026-09-18 02:14:09 +02:00
jschoubben 591b26f712 Ten migration-critical modules become mesh-buildable (issue 060)
keycloak, mailu, minio, mongodb, mssql, nextcloud, portainer, redis,
umami and verdaccio get the Dockerfile + build section the eight
buildable modules already had; their runtime containers name the
artifact instead of a placeholder digest.

One convention, settled (060's open question, informed by 061): the
runtime container runs serve mode with every serve-time entrypoint in
MESH_TOOL_MODULES — tools serve, events flow, and a provider's
provisioner reconciles in the same process with the broker connected.
postgres, gitea and lavinmq are retrofitted from args-run provisioners,
which served no tools and emitted lifecycle events nowhere.

route-proxy is deferred: its build context is the mesh-controller
repository, a cross-repo shape the build section cannot yet express.
2026-09-18 02:02:21 +02:00
jschoubben 83a78b4fb3 Merge pull request 'Refs name the registry; lavinmq's runtime runs its provisioner (042/048/061)' (#26) from feat/refs-name-the-registry into main 2026-09-18 01:00:43 +02:00
jschoubben 1891b09c65 lavinmq's runtime container runs its provisioner
The runtime container named no command, so it ran the image default —
the tool host — and the provisioner entrypoint compiled beside it never
ran anywhere: no vhost was ever minted, while the grants sat applied in
its mounted directory. postgres already names its provisioner in args;
lavinmq now does the same. Surfaced by the built-store-cross-node bed,
run 9 — the first bed to reach the vhost assertion honestly.
2026-09-17 23:54:56 +02:00
jschoubben 2dab3069d2 The builder names the registry by the binding again (ADR 0082)
The one-line change e0c9219 parked "until there is a certificate" returns — with the
overlay recorded as the registry's transport security and every node's runtime told the
store speaks plain HTTP, a reference under the provider's internal name is one every
machine can pull. References minted at genesis stay loopback and are valid where they
matter, on the machine that made them.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:05:25 +02:00
jschoubben 728922a016 Merge pull request 'Foundation modules claim a mesh-scoped seat named after their server' (#25) from multi-node/foundation-seats into main 2026-09-17 02:15:59 +02:00
jschoubben fccc1e6552 The foundation modules claim a mesh-scoped seat named after their server
postgres claims mesh-store, lavinmq claims mesh-broker, and the controller's seat is
renamed the-controller -> mesh-controller so all three follow one convention. The resolver
refuses a second holder mesh-wide, so an adopted foundation module assigned to a second node
is refused rather than silently raising a second server. Closes hq issue 056 (ADR 0079).

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 02:13:26 +02:00
jschoubben a24362b74b Merge pull request 'The adopted broker declares its amqps bus port (5671) in listens' (#24) from fix/broker-declares-amqps-port into main 2026-09-17 01:44:47 +02:00
jschoubben 06fd436eca The adopted broker declares its amqps bus port (5671) in listens
mesh-broker serves the amqps bus on 5671 (the control plane and every module's
events) but the lavinmq module declared only 5672, so the firewall's forward
chain — where the broker's published ports are matched — never opened 5671, and
a consumer on another node could not reach the bus over the overlay. Part of
issue 055.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:01:54 +02:00
jschoubben f155493926 Merge pull request 'Foundation modules adopted, and the mesh-controller/foundation rename' (#23) from feat/foundation-and-rename into main 2026-09-16 23:22:47 +02:00
jschoubben 5e4dc3748e Phase 3.2: the lavinmq module adopts mesh-broker instead of raising its own
server is now the container the foundation raised — name mesh-broker, the same
pinned upstream lavinmq image, the same TLS args, ports and volumes — so the
applier adopts it in place. The second server, the lavinmq.ini bootstrap and the
module's own network are gone. The provisioner is host-networked to the broker's
loopback management (127.0.0.1:15672) and authenticates as lavinmq's default
guest, which the foundation broker runs with; the mesh bus stays on the / vhost,
a vhost-per-consumer beside it. One lavinmq now.

Issue 051 (WBS 3.2).

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:24:13 +02:00
jschoubben a63ef3d954 Phase 3.1: the postgres module adopts mesh-store instead of raising its own
server is now the container the foundation raised — same name (mesh-store),
same env (POSTGRES_PASSWORD/PGDATA), ports (127.0.0.1:5432:5432), volume
(mesh-store-data) and the same pinned upstream postgres image the foundation
runs — so the applier adopts it in place rather than raising a second postgres.
The provisioner is host-networked to reach the loopback store at 127.0.0.1:5432.
The module's own network and bind-mounted data dir are gone; there is one
postgres now, holding the controller's contexts and every module's database.

Issue 051 (WBS 3.1).

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:04:06 +02:00
jschoubben 41637befff Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben b520bd1825 The builder's workspace is a same-path bind, so sibling builds see the clone
A docker run -v from inside the builder resolves the path on the host: with the
workspace mounted at a different path inside than out, the SDK publish and any
bundle compile mounted an empty directory. Bind it at the same path both sides.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 16:02:26 +02:00
jschoubben 065ddd6d69 gitea provides the package registry; the builder gets its npm credential
gitea gains the package-registry provision: serves/receives/grants, an admin
own-secret, a postgres-shaped build, and a provisioner that creates a gitea user
per consumer with the mesh-minted password and seals nothing (hq ADR 0048). The
builder takes its registry credential as an own-secret rather than a resolved
provision, because gitea-as-module needs the base to build its provisioner and so
cannot resolve before the base — a cycle the own-secret avoids.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:27:26 +02:00
jschoubben abcba14edd Review: four manifests said something stale or nothing at all
dnsmasq still required resolver-data, a provision that died with the
mesh-resolver module — assigning it would refuse with "nothing provides
resolver-data". It asks for the node-zones fact now, at the same path its
config already reads, restarting on the fact's own id.

gitea and verdaccio both provide package-registry now — ADR 0075's provision,
which neither declared, so ADR 0014's "consumes from the private registry" had
no provider anywhere in the catalogue. Two providers, mesh-scoped: the resolver
refuses until one is assigned, and choosing is assigning, which is the designed
shape.

audit-logger runs a container and declared no capability, alone among the
containerised modules. A machine without a runtime would have been assigned it
and failed at apply rather than at assignment.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 22:03:33 +02:00
jschoubben 1cb33732f2 Move amqp-ping's source, so the mesh has something to notice 2026-09-15 21:03:43 +02:00
jschoubben 71bbc7dab0 Name the two modules after their software: nftables and distribution
A module's identity is the software it is (ADR 0040). Two were named after the
job instead, and the job already had a name.

firewall installs the nftables package and runs nftables.service. The seat it
claims is the-packet-filter, which is correctly named for the role. Calling the
module firewall named neither the software nor the provision, and promised that
any firewall could sit there — the false genericity the naming rule forbids.

registry runs Distribution, the OCI reference implementation, and provides
artifact-store. So registry was a third name for a thing that already had two,
which is how one word ended up meaning the module, the software and the concept
in the same paragraph.

The capability stays firewall, and correctly: a capability IS a functionality, so
a node having one and fail2ban requiring one are both right. Only the module
moves.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:51:00 +02:00
jschoubben bf1f67a485 listens names the port the software uses; serves is the mesh's to fill
Over-corrected: taking the port out of listens as well as serves made the module
declare it listens on nothing, and the parser said so.

ADR 0038 splits it. A module names the port its own software listens on, because
that is a fact about the software and it knows it. The mesh assigns the
machine-side number, because only the mesh knows what else is on the machine, and
it is the mesh that fills the assigned number into serves so a consumer is told
one number rather than three that agree by luck.

So what was wrong was writing a port into serves, not into listens.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 12:59:00 +02:00
jschoubben 9a41add136 showcase: a module that exercises everything a module can be
Written so the module system has something that proves itself rather than a claim
about what it supports, and guarded by a test in the catalogue's own suite so it
cannot quietly stop exercising things.

Nine of the host's eleven resource kinds, all three artifact kinds including the
one that compiles, all three ways a module's code can run, and all four things
that code can be: tools, an event consumer, a provisioner, and processes.

Two absences that are findings rather than gaps. `action` is refused to modules
outright — the link may not carry a command to run (ADR 0005), so a module that
needs something done ships a program that reconciles, which is what a run-once
process is. `service` puts an EXISTING unit into a state and installs none, which
is right for software shipping its own; code the mesh built has no unit until the
mesh writes one, and that is a process.

And it no longer picks its own port. ADR 0038 says a module cannot know what else
is on the machine it was assigned to, and names exactly the trap this fell into:
the number written three times — listens, serves, a container's ports — agreeing
only because one person wrote all three, with nothing checking. So it says what
it needs and the mesh assigns the number.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 12:58:30 +02:00
jschoubben c4e3ebee86 Move amqp-ping's source, so the mesh has something to notice 2026-09-15 01:42:57 +02:00
jschoubben d2dce34716 The catalogue asks on start, and registers a replay as history
A replayed build is registered exactly as any other and announced to nobody. A
module that moved months ago is not something anything should act on now:
emitting `upgraded` would have the control plane decide about a rollout, and
`rebuild-needed` would ask for builds of things already current.

Asked on every start rather than only the first, because a catalogue cannot tell
whether it has a gap — and the answer is idempotent, so asking when there is none
costs a message. Asked after subscribing, so a build arriving during the replay
is not lost between the two.

Closes novox/hq 04-ISSUES/050 with mesh-control.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:26:58 +02:00
jschoubben 4aa54fbbe0 Move amqp-ping's source, so the mesh has something to notice 2026-09-14 23:42:52 +02:00
jschoubben e0c92195d4 The builder names the registry by loopback until there is a certificate
Naming it from the binding was right and arrived too early. The moment the
machine had a name, the builder pushed to <node>.internal:5000 and the runtime
refused it: "http: server gave HTTP response to HTTPS client". The registry
serves plaintext, and anything that is not loopback is required to be HTTPS.

So there are two phases, and this is the first. Before the mesh has a certificate
authority of its own, loopback is the only trusted path that is honest — it is
trusted because it cannot leave the machine, not because anyone checked
anything. The mesh-reachable name belongs to the second phase, with TLS from the
mesh's own CA, and the binding expression returns then.

Not a revert of the reasoning: novox/hq issue 048 stays open and this is why. The
same one-line change lands again once a certificate module is running.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:38:05 +02:00
jschoubben af1e3afa49 amqp-ping declares a broker secret it never mounts
Its runtime is a tool host: it connects to the mesh's broker before it does
anything else. The module declares own-secrets.broker, and then its container
neither mounts that file nor names it, so the runtime started and said there was
no broker to reach, forever, in a restart loop.

lavinmq's runtime container does both, and is the shape this follows.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:20:48 +02:00
jschoubben c61c7f74f9 Convert lavinmq: it is not only a broker, so it has to be built
The mesh refused to place it — two of its three containers named
mesh-runtime-lavinmq@sha256:000…0, "a placeholder digest, which is never a real
image". That refusal was right, and the belief behind the placeholder was that
lavinmq needs no building because its broker is an upstream image.

The broker is upstream. The module is not the broker. It carries a run-once
bootstrap that writes the broker's configuration before it first starts, a
provisioner that grants each consumer its own vhost and user, a set of tools and
an event consumer — all of it this module's own TypeScript, and none of it
producible by naming somebody else's image.

So it gets what every module with code of its own gets: a Dockerfile standing on
the shared toolchain and runtime bases, a build block naming them, and containers
that name the artifact rather than a digest nothing can produce. Same recipe as
postgres, which is the converted module closest in shape — it has a provisioner
too.

Noted and deliberately not changed: postgres runs its provisioner from its
container's args, and lavinmq's equivalent container names none, so on this
manifest the provisioner is never started. That may be why, or may be a second
fault; it is left alone so the next run says which.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:05:10 +02:00
jschoubben 030509558c Name the registry where every machine can reach it, not where the builder stands
The builder's environment said MESH_REGISTRY=127.0.0.1:${bound:artifact-store:port}
— the port taken from the binding, the host pinned to loopback. So every artifact
the mesh builds was recorded under an address that means something only on the
machine holding the registry, and nothing else in the mesh could resolve it.

Loopback is correct for exactly one reader and the builder is not special: it
already requires artifact-store, and the binding states where the provider is on
the private network. It now uses both halves of what it was given.

Invisible with one machine, which is the only shape this had been proven in. The
registry module declares that every machine pulls from it and opens its port to
the mesh for that reason, so the reference it is handed has to be one a second
machine can use.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:00:48 +02:00
jschoubben 680b91546c Migrate audit-logger: it builds itself now
The first module moved onto the new build process. It named a placeholder digest
nothing could produce, so it only ever worked where somebody had pre-built its
image by hand. It names the two shared bases instead, and the mesh builds it.

Chosen first deliberately: it requires nothing, nothing requires it, and an
audit trail of every event on the mesh is the thing most worth having while
modules are being moved one at a time.
2026-09-14 12:58:30 +02:00
jschoubben 071afc1e2c Merge pull request 'Modules compile in the toolchain and ship on the runtime' (#22) from feat/build-and-run-are-two-images into main 2026-09-14 02:02:42 +02:00
jschoubben cf116a4932 Modules compile in the toolchain and ship on the runtime 2026-09-14 01:57:06 +02:00
jschoubben 5ea88b149b The three modules name their base rather than pinning a copy of it
Each named a digest produced inside a lab that no longer exists, so none of them
could be built anywhere else. They say which module they stand on now, and there
is deliberately no default — a build nobody told stops at the declaration rather
than at a reference that resolves to nothing.
2026-09-13 23:53:22 +02:00
jschoubben 720e3706b2 Merge pull request 'The catalogue holds the module graph, and the modules that make it build themselves' (#21) from feat/the-catalogue-module into main 2026-09-13 11:17:10 +02:00
jschoubben 729582cc55 Remove what was only there to move a commit
A line in postgres's recipe and a one-line file in amqp-ping, both added to
move a commit and watch the mesh notice. The proofs worked; neither was meant
to stay. MESH_MODULE is set from the sealed credential at run time anyway, so
baking it in was dead weight as well as noise.
2026-09-13 11:15:33 +02:00
jschoubben 21ad008879 Stale means the artifact moved, not the commit
A comment changed in a build recipe is a new commit and a byte-identical image.
Comparing commits called every module standing on it stale, so the mesh would
have rebuilt itself entirely to arrive back exactly where it started — and
listed each dependent once per commit that had produced the same image.
2026-09-13 02:48:21 +02:00
jschoubben 87243bc524 Every module stands on a base the mesh built
The base was a digest typed in by hand, for an image nothing in the mesh could
produce — so the graph held edges pointing at it with no version on the far end,
and the one change that reaches every module at once could never be noticed.
It is a module now, and these edges resolve.
2026-09-13 02:45:37 +02:00
jschoubben 1d0d9a3894 Name the module in its own runtime, and move the artifact with it 2026-09-13 01:58:48 +02:00
jschoubben 358c7d5a2d Move postgres's commit, to watch the mesh roll it out by itself 2026-09-13 01:57:00 +02:00
jschoubben d5300120d6 Move amqp-ping's commit, to see what the catalogue announces 2026-09-13 01:49:48 +02:00