Review of the registry hand-over. The store's seat was node-scoped, so a gate assigned to a
machine without the store pulled a second, empty store in beside it — behind the real
credentials and the public name, and offering `artifact-store` a second time so every
consumer elsewhere refused. `the-artifact-store` is one per mesh: a second store anywhere,
however it got there, is refused by name.
storage.delete.enabled was carried onto the store's own door, which the whole private
network reaches with no account (hq ADR 0082); anything on the overlay could have deleted a
manifest. Nothing needs it there — garbage collection was not carried. It stays on the gate
only, behind the registry's own auth, where tag retention runs.
hq ADR 0082/0104, the registry hand-over.
The predecessor serves the registry under a public name, behind htpasswd basic auth, with a
twenty-gigabyte body limit for layer pushes. The mesh's registry has no name, no lock and no
limit — by design inside the mesh, where the private network is the boundary and every node
pulls without an account (hq ADR 0082). Taking the name over must not change that.
A route on `distribution` itself would: contributing a route is requiring one, and the store
is raised at genesis on a node with no proxy. So the public door is `distribution-gate`, a
second registry process on the same volume, behind the registry's own htpasswd (the
predecessor's realm, the predecessor's file, carried in with `secret accept`), with the
route and its limit. It requires the store's storage as a node-scoped provision, so it can
only land beside the store. The store's own door is untouched — no auth, no htpasswd — which
is what keeps the builder's pushes and every node's pulls working.
Both processes read the predecessor's configuration where it changed behaviour: delete
enabled, which tag retention depends on; no per-process descriptor cache, which two
processes over one store cannot share; the CORS headers for the retired interface dropped.
route-adapter writes the limit as the predecessor's own buffering middleware, named after
the router, only when asked for — and skips a route whose limit it cannot read rather than
carrying what the module said not to.
hq ADR 0082/0104, the registry hand-over.
Read against hal/modules/mssql: the predecessor publishes 4848:1433 and its connections
block hands every consumer localhost:4848; its data has lived in /services/mssql/data
since it was installed — 4.6 GB of it. The nox module had 1433, db-data, and a pin three
builds behind, so taking it would have moved the port every consumer was told about,
started on an empty directory, and downgraded the engine at the same time.
db-data was not a convention: only this module and mongodb used it, and both predecessors
say data. Now nothing moves at the cutover.
The pin was a month old — 2026-08-20 against the 2026-09-17 image the container runs.
Third module in a row whose pin had aged into a downgrade; the pattern is hq issue 099.
The rule the forge's cutover taught: a module that takes over a running service must not
carry an older image than the one running, or the cutover is a downgrade nobody asked for.
6.10.4 exists; moving to it is an upgrade and its own act.
A node being adopted cannot take a web module: every module reachable by
name requires route, the mesh's only provider of it binds the two public
ports, and the predecessor's proxy holds them and serves every public
name there. Stopping the predecessor to break the circle darkens every
name at once, with every certificate to re-obtain in the same window.
So this answers the same provision without binding anything. It provides
route and receives the same contributions file, and writes each
contribution as one route file where the predecessor's file provider
reads, naming the predecessor's own certificate resolver so no
certificate is asked for. It removes a file it wrote when its
contribution goes and never touches a file it did not write — the name
and a marker inside both have to say it is the mesh's.
A step, not a daemon: run-once, re-run by restart-on over the received
file and the settings. The predecessor's dynamic directory is a node
setting, because it is a fact about one machine.
Migration scaffolding with a stated end: assigned only on an adopted
node, deleted when the predecessor's proxy retires.
MESH_GITEA_URL named 3000 outright. The forge's container publishes 3000 without
fixing the machine side, so the mesh assigns it — and on a node given that port
as a setting it is the operator's number. The mapping that lets the forge go on
binding 3000 does nothing for a caller dialling the machine's loopback, so the
sidecar dialled a port nothing was listening on wherever the two differed.
${port:3000} is the mesh's answer to exactly that question, and it now resolves
in a container's environment as it always has in a file. The forge declares it
listens on 3000, so it may ask.
Still names 3000: the route contribution (hq issue 089) and the `listens` and
`ports` entries, which are the software's own number and belong there.
`at` was settable. A node setting could point the builder's package binding at
any host, and the builder sends its registry credential there as basic auth — so
a setting meant for a port was a way to hand the password to somebody else.
novox/hq 04-ISSUES/085
The forge is raised by hand at genesis, before any module provides
`package-registry`, so the builder carries a binding instead of resolving one —
and the port in it was rewritten, as text, by the installer. A later
registration of this manifest from the catalogue put 3000 back, silently, and
pointed the builder and its registry credential at whatever holds that port.
Made settable instead: the port is a per-node setting the controller holds, and
the catalogue's number is only its default. `provision`, `from` and `as` are
protected — a setting here is about where the forge answers, never about who the
binding is with.
novox/hq 04-ISSUES/085, ADR 0100
step-ca's root certificate, its key and that key's password were own secrets — random
bytes the mesh minted, which no certificate is (novox/hq 04-ISSUES/076). The mesh
mints only the CA password now; step-ca makes its root at first start and serves it
at /roots.pem, which the manifest now names beside the ACME directory. The route
proxy fetches that root over the mesh network in a run-once step before it starts,
instead of being handed a served fact that could only be written before anything ran
(ADR 0098).
The copy between registries (ADR 0096) reaches the mesh's registry over HTTP from
inside the builder's container, where loopback on the default bridge is not the
machine; docker push never noticed because it went through the machine's daemon.
The builder already holds the runtime's socket, so the host network adds nothing
it did not have.
gitea, umami, influxdb, icecast and mailu require a secret and keep each of theirs
under a local name (novox/hq ADR 0094); the broker account stays their own. The
route proxy's recipe starts FROM the bases its manifest declares (ADR 0097).
ADR 0069 put it there; genesis now takes it from the build it runs (mesh-host,
novox/hq 04-ISSUES/072). This copy was read by nothing else and had already drifted.
The official entrypoint re-executes itself as mongodb (uid 999) and only then reads
MONGO_INITDB_ROOT_PASSWORD_FILE, so a root-owned 0600 file is 'Permission denied' at line 83 and
the server never starts. secrets-owner is the mechanism ADR 0086 gives for exactly this.
From the survey of every env-file secret (ADR 0086, issue 041): amqp-ping,
minio, mongodb and grafana use the _FILE twin their software honours;
mesh-catalog and model-usage read DATABASE_URL_FILE (a file the mesh
templates, mounted where only the runtime reads it); grafana's secret files
belong to its own account. Two dead deliveries removed: a line nothing read
in amqp-email-forwarder, and mailu's secret.env on four containers that
never read it. The 25 exceptions that remain carry the surveyed reason —
convertible and awaiting a bed, convertible through a generated config file,
the application's own code, or not convertible.
ADR 0086. mesh-controller mounts its six own secrets and names them with
_FILE twins, so no credential of its own reaches its environment. The 35
containers that still read a secret through an env-file carry
secrets-in-environment with the reason; converting each where its software
accepts a path is the per-module work of issue 041.
mesh-vault provides `secret` (novox/hq ADR 0085, design 24). The value is
the pair credential the controller mints — the vault holds no copy, only a
ledger of who holds one, its fingerprint and every rotation, and two tools that
answer by fingerprint and never by value. Rotation is `rotate secret`,
unchanged machinery pointed at a secret with an owner (design 13). Named in the
mesh's own namespace, beside mesh-controller and mesh-catalog, because it is
the mesh's own code rather than wrapped software.
redis is the first consumer: its own password stops being an own-secret nothing
could rotate and becomes a `secret` it requires, read from the same file into
the same hole. The server now restarts on its config, or it would keep the
password it started with through every rotation (playbook 06).
The comment still said the provisioner is not listed and runs via a
container's args — but the 061 fix put it in MESH_TOOL_MODULES (serve
mode) and dropped the args. A future editor trusting the comment could
strip it again and silently reintroduce 061. Comment now matches the
code; only the run-once bootstrap runs via args.
minio's runtime copies the `mc` client from minio/mc:latest — a
Docker Hub pull the mesh build environment cannot make (its docker
reaches the mesh registry, not public Hub), the same isolation that
blocks npm deps. apt-based installs (mongodb's mongosh, mosquitto)
build fine because the build has real internet for apt; only npm and
Docker Hub are redirected. Delivering an external binary or image layer
into a mesh build is the same open question as the npm deps — deferred
with them.
Both carry third-party runtime deps (pg; tweetnacl + sealedbox) that
their Dockerfile installs with npm — which 404s in a mesh build, whose
npm points at the mesh's own registry, not public npm. The workstation
build script got away with it by installing on a host with public npm.
Delivering a module's third-party deps into a mesh build is an open
question (how: publish to the mesh registry, or proxy); until it is
answered these two stay on the placeholder path they were already on.
The other 37 modules build from their own directory with no external
fetch.
The 21 media/home modules, the SaaS tool modules (cloudflare-dns,
confluence, gitlab, jira), model-usage, and the model-access trio get
the same Dockerfile + build section as batch 1. Scheduled-only modules
(anthropic-consumer, anthropic-manager, openai-consumer) deliberately
declare no MESH_TOOL_MODULES — every container of theirs names its
command. mosquitto's run-once bootstrap container builds from the same
artifact. Modules with third-party deps (model-usage: pg;
anthropic-manager: tweetnacl) install them beside their compiled code.
Also fixes cloudflare-dns's package.json, unparseable since its
description lost a closing quote.
Deliberately still without build sections: builder and mesh-controller
(the foundation builds them by its own path), distribution (provides
the artifact store — building it through itself is refused by design),
route-proxy (cross-repo build context, deferred), and the
upstream-image-only modules, which have no code to build.
keycloak, mailu, minio, mongodb, mssql, nextcloud, portainer, redis,
umami and verdaccio get the Dockerfile + build section the eight
buildable modules already had; their runtime containers name the
artifact instead of a placeholder digest.
One convention, settled (060's open question, informed by 061): the
runtime container runs serve mode with every serve-time entrypoint in
MESH_TOOL_MODULES — tools serve, events flow, and a provider's
provisioner reconciles in the same process with the broker connected.
postgres, gitea and lavinmq are retrofitted from args-run provisioners,
which served no tools and emitted lifecycle events nowhere.
route-proxy is deferred: its build context is the mesh-controller
repository, a cross-repo shape the build section cannot yet express.
The runtime container named no command, so it ran the image default —
the tool host — and the provisioner entrypoint compiled beside it never
ran anywhere: no vhost was ever minted, while the grants sat applied in
its mounted directory. postgres already names its provisioner in args;
lavinmq now does the same. Surfaced by the built-store-cross-node bed,
run 9 — the first bed to reach the vhost assertion honestly.
The one-line change e0c9219 parked "until there is a certificate" returns — with the
overlay recorded as the registry's transport security and every node's runtime told the
store speaks plain HTTP, a reference under the provider's internal name is one every
machine can pull. References minted at genesis stay loopback and are valid where they
matter, on the machine that made them.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
postgres claims mesh-store, lavinmq claims mesh-broker, and the controller's seat is
renamed the-controller -> mesh-controller so all three follow one convention. The resolver
refuses a second holder mesh-wide, so an adopted foundation module assigned to a second node
is refused rather than silently raising a second server. Closes hq issue 056 (ADR 0079).
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-broker serves the amqps bus on 5671 (the control plane and every module's
events) but the lavinmq module declared only 5672, so the firewall's forward
chain — where the broker's published ports are matched — never opened 5671, and
a consumer on another node could not reach the bus over the overlay. Part of
issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — name mesh-broker, the same
pinned upstream lavinmq image, the same TLS args, ports and volumes — so the
applier adopts it in place. The second server, the lavinmq.ini bootstrap and the
module's own network are gone. The provisioner is host-networked to the broker's
loopback management (127.0.0.1:15672) and authenticates as lavinmq's default
guest, which the foundation broker runs with; the mesh bus stays on the / vhost,
a vhost-per-consumer beside it. One lavinmq now.
Issue 051 (WBS 3.2).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — same name (mesh-store),
same env (POSTGRES_PASSWORD/PGDATA), ports (127.0.0.1:5432:5432), volume
(mesh-store-data) and the same pinned upstream postgres image the foundation
runs — so the applier adopts it in place rather than raising a second postgres.
The provisioner is host-networked to reach the loopback store at 127.0.0.1:5432.
The module's own network and bind-mounted data dir are gone; there is one
postgres now, holding the controller's contexts and every module's database.
Issue 051 (WBS 3.1).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A docker run -v from inside the builder resolves the path on the host: with the
workspace mounted at a different path inside than out, the SDK publish and any
bundle compile mounted an empty directory. Bind it at the same path both sides.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
gitea gains the package-registry provision: serves/receives/grants, an admin
own-secret, a postgres-shaped build, and a provisioner that creates a gitea user
per consumer with the mesh-minted password and seals nothing (hq ADR 0048). The
builder takes its registry credential as an own-secret rather than a resolved
provision, because gitea-as-module needs the base to build its provisioner and so
cannot resolve before the base — a cycle the own-secret avoids.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
dnsmasq still required resolver-data, a provision that died with the
mesh-resolver module — assigning it would refuse with "nothing provides
resolver-data". It asks for the node-zones fact now, at the same path its
config already reads, restarting on the fact's own id.
gitea and verdaccio both provide package-registry now — ADR 0075's provision,
which neither declared, so ADR 0014's "consumes from the private registry" had
no provider anywhere in the catalogue. Two providers, mesh-scoped: the resolver
refuses until one is assigned, and choosing is assigning, which is the designed
shape.
audit-logger runs a container and declared no capability, alone among the
containerised modules. A machine without a runtime would have been assigned it
and failed at apply rather than at assignment.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx