The comment still said the provisioner is not listed and runs via a
container's args — but the 061 fix put it in MESH_TOOL_MODULES (serve
mode) and dropped the args. A future editor trusting the comment could
strip it again and silently reintroduce 061. Comment now matches the
code; only the run-once bootstrap runs via args.
minio's runtime copies the `mc` client from minio/mc:latest — a
Docker Hub pull the mesh build environment cannot make (its docker
reaches the mesh registry, not public Hub), the same isolation that
blocks npm deps. apt-based installs (mongodb's mongosh, mosquitto)
build fine because the build has real internet for apt; only npm and
Docker Hub are redirected. Delivering an external binary or image layer
into a mesh build is the same open question as the npm deps — deferred
with them.
Both carry third-party runtime deps (pg; tweetnacl + sealedbox) that
their Dockerfile installs with npm — which 404s in a mesh build, whose
npm points at the mesh's own registry, not public npm. The workstation
build script got away with it by installing on a host with public npm.
Delivering a module's third-party deps into a mesh build is an open
question (how: publish to the mesh registry, or proxy); until it is
answered these two stay on the placeholder path they were already on.
The other 37 modules build from their own directory with no external
fetch.
The 21 media/home modules, the SaaS tool modules (cloudflare-dns,
confluence, gitlab, jira), model-usage, and the model-access trio get
the same Dockerfile + build section as batch 1. Scheduled-only modules
(anthropic-consumer, anthropic-manager, openai-consumer) deliberately
declare no MESH_TOOL_MODULES — every container of theirs names its
command. mosquitto's run-once bootstrap container builds from the same
artifact. Modules with third-party deps (model-usage: pg;
anthropic-manager: tweetnacl) install them beside their compiled code.
Also fixes cloudflare-dns's package.json, unparseable since its
description lost a closing quote.
Deliberately still without build sections: builder and mesh-controller
(the foundation builds them by its own path), distribution (provides
the artifact store — building it through itself is refused by design),
route-proxy (cross-repo build context, deferred), and the
upstream-image-only modules, which have no code to build.
keycloak, mailu, minio, mongodb, mssql, nextcloud, portainer, redis,
umami and verdaccio get the Dockerfile + build section the eight
buildable modules already had; their runtime containers name the
artifact instead of a placeholder digest.
One convention, settled (060's open question, informed by 061): the
runtime container runs serve mode with every serve-time entrypoint in
MESH_TOOL_MODULES — tools serve, events flow, and a provider's
provisioner reconciles in the same process with the broker connected.
postgres, gitea and lavinmq are retrofitted from args-run provisioners,
which served no tools and emitted lifecycle events nowhere.
route-proxy is deferred: its build context is the mesh-controller
repository, a cross-repo shape the build section cannot yet express.
The runtime container named no command, so it ran the image default —
the tool host — and the provisioner entrypoint compiled beside it never
ran anywhere: no vhost was ever minted, while the grants sat applied in
its mounted directory. postgres already names its provisioner in args;
lavinmq now does the same. Surfaced by the built-store-cross-node bed,
run 9 — the first bed to reach the vhost assertion honestly.
The one-line change e0c9219 parked "until there is a certificate" returns — with the
overlay recorded as the registry's transport security and every node's runtime told the
store speaks plain HTTP, a reference under the provider's internal name is one every
machine can pull. References minted at genesis stay loopback and are valid where they
matter, on the machine that made them.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
postgres claims mesh-store, lavinmq claims mesh-broker, and the controller's seat is
renamed the-controller -> mesh-controller so all three follow one convention. The resolver
refuses a second holder mesh-wide, so an adopted foundation module assigned to a second node
is refused rather than silently raising a second server. Closes hq issue 056 (ADR 0079).
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-broker serves the amqps bus on 5671 (the control plane and every module's
events) but the lavinmq module declared only 5672, so the firewall's forward
chain — where the broker's published ports are matched — never opened 5671, and
a consumer on another node could not reach the bus over the overlay. Part of
issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — name mesh-broker, the same
pinned upstream lavinmq image, the same TLS args, ports and volumes — so the
applier adopts it in place. The second server, the lavinmq.ini bootstrap and the
module's own network are gone. The provisioner is host-networked to the broker's
loopback management (127.0.0.1:15672) and authenticates as lavinmq's default
guest, which the foundation broker runs with; the mesh bus stays on the / vhost,
a vhost-per-consumer beside it. One lavinmq now.
Issue 051 (WBS 3.2).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — same name (mesh-store),
same env (POSTGRES_PASSWORD/PGDATA), ports (127.0.0.1:5432:5432), volume
(mesh-store-data) and the same pinned upstream postgres image the foundation
runs — so the applier adopts it in place rather than raising a second postgres.
The provisioner is host-networked to reach the loopback store at 127.0.0.1:5432.
The module's own network and bind-mounted data dir are gone; there is one
postgres now, holding the controller's contexts and every module's database.
Issue 051 (WBS 3.1).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A docker run -v from inside the builder resolves the path on the host: with the
workspace mounted at a different path inside than out, the SDK publish and any
bundle compile mounted an empty directory. Bind it at the same path both sides.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
gitea gains the package-registry provision: serves/receives/grants, an admin
own-secret, a postgres-shaped build, and a provisioner that creates a gitea user
per consumer with the mesh-minted password and seals nothing (hq ADR 0048). The
builder takes its registry credential as an own-secret rather than a resolved
provision, because gitea-as-module needs the base to build its provisioner and so
cannot resolve before the base — a cycle the own-secret avoids.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
dnsmasq still required resolver-data, a provision that died with the
mesh-resolver module — assigning it would refuse with "nothing provides
resolver-data". It asks for the node-zones fact now, at the same path its
config already reads, restarting on the fact's own id.
gitea and verdaccio both provide package-registry now — ADR 0075's provision,
which neither declared, so ADR 0014's "consumes from the private registry" had
no provider anywhere in the catalogue. Two providers, mesh-scoped: the resolver
refuses until one is assigned, and choosing is assigning, which is the designed
shape.
audit-logger runs a container and declared no capability, alone among the
containerised modules. A machine without a runtime would have been assigned it
and failed at apply rather than at assignment.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module's identity is the software it is (ADR 0040). Two were named after the
job instead, and the job already had a name.
firewall installs the nftables package and runs nftables.service. The seat it
claims is the-packet-filter, which is correctly named for the role. Calling the
module firewall named neither the software nor the provision, and promised that
any firewall could sit there — the false genericity the naming rule forbids.
registry runs Distribution, the OCI reference implementation, and provides
artifact-store. So registry was a third name for a thing that already had two,
which is how one word ended up meaning the module, the software and the concept
in the same paragraph.
The capability stays firewall, and correctly: a capability IS a functionality, so
a node having one and fail2ban requiring one are both right. Only the module
moves.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Over-corrected: taking the port out of listens as well as serves made the module
declare it listens on nothing, and the parser said so.
ADR 0038 splits it. A module names the port its own software listens on, because
that is a fact about the software and it knows it. The mesh assigns the
machine-side number, because only the mesh knows what else is on the machine, and
it is the mesh that fills the assigned number into serves so a consumer is told
one number rather than three that agree by luck.
So what was wrong was writing a port into serves, not into listens.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Written so the module system has something that proves itself rather than a claim
about what it supports, and guarded by a test in the catalogue's own suite so it
cannot quietly stop exercising things.
Nine of the host's eleven resource kinds, all three artifact kinds including the
one that compiles, all three ways a module's code can run, and all four things
that code can be: tools, an event consumer, a provisioner, and processes.
Two absences that are findings rather than gaps. `action` is refused to modules
outright — the link may not carry a command to run (ADR 0005), so a module that
needs something done ships a program that reconciles, which is what a run-once
process is. `service` puts an EXISTING unit into a state and installs none, which
is right for software shipping its own; code the mesh built has no unit until the
mesh writes one, and that is a process.
And it no longer picks its own port. ADR 0038 says a module cannot know what else
is on the machine it was assigned to, and names exactly the trap this fell into:
the number written three times — listens, serves, a container's ports — agreeing
only because one person wrote all three, with nothing checking. So it says what
it needs and the mesh assigns the number.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A replayed build is registered exactly as any other and announced to nobody. A
module that moved months ago is not something anything should act on now:
emitting `upgraded` would have the control plane decide about a rollout, and
`rebuild-needed` would ask for builds of things already current.
Asked on every start rather than only the first, because a catalogue cannot tell
whether it has a gap — and the answer is idempotent, so asking when there is none
costs a message. Asked after subscribing, so a build arriving during the replay
is not lost between the two.
Closes novox/hq 04-ISSUES/050 with mesh-control.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Naming it from the binding was right and arrived too early. The moment the
machine had a name, the builder pushed to <node>.internal:5000 and the runtime
refused it: "http: server gave HTTP response to HTTPS client". The registry
serves plaintext, and anything that is not loopback is required to be HTTPS.
So there are two phases, and this is the first. Before the mesh has a certificate
authority of its own, loopback is the only trusted path that is honest — it is
trusted because it cannot leave the machine, not because anyone checked
anything. The mesh-reachable name belongs to the second phase, with TLS from the
mesh's own CA, and the binding expression returns then.
Not a revert of the reasoning: novox/hq issue 048 stays open and this is why. The
same one-line change lands again once a certificate module is running.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Its runtime is a tool host: it connects to the mesh's broker before it does
anything else. The module declares own-secrets.broker, and then its container
neither mounts that file nor names it, so the runtime started and said there was
no broker to reach, forever, in a restart loop.
lavinmq's runtime container does both, and is the shape this follows.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh refused to place it — two of its three containers named
mesh-runtime-lavinmq@sha256:000…0, "a placeholder digest, which is never a real
image". That refusal was right, and the belief behind the placeholder was that
lavinmq needs no building because its broker is an upstream image.
The broker is upstream. The module is not the broker. It carries a run-once
bootstrap that writes the broker's configuration before it first starts, a
provisioner that grants each consumer its own vhost and user, a set of tools and
an event consumer — all of it this module's own TypeScript, and none of it
producible by naming somebody else's image.
So it gets what every module with code of its own gets: a Dockerfile standing on
the shared toolchain and runtime bases, a build block naming them, and containers
that name the artifact rather than a digest nothing can produce. Same recipe as
postgres, which is the converted module closest in shape — it has a provisioner
too.
Noted and deliberately not changed: postgres runs its provisioner from its
container's args, and lavinmq's equivalent container names none, so on this
manifest the provisioner is never started. That may be why, or may be a second
fault; it is left alone so the next run says which.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The builder's environment said MESH_REGISTRY=127.0.0.1:${bound:artifact-store:port}
— the port taken from the binding, the host pinned to loopback. So every artifact
the mesh builds was recorded under an address that means something only on the
machine holding the registry, and nothing else in the mesh could resolve it.
Loopback is correct for exactly one reader and the builder is not special: it
already requires artifact-store, and the binding states where the provider is on
the private network. It now uses both halves of what it was given.
Invisible with one machine, which is the only shape this had been proven in. The
registry module declares that every machine pulls from it and opens its port to
the mesh for that reason, so the reference it is handed has to be one a second
machine can use.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first module moved onto the new build process. It named a placeholder digest
nothing could produce, so it only ever worked where somebody had pre-built its
image by hand. It names the two shared bases instead, and the mesh builds it.
Chosen first deliberately: it requires nothing, nothing requires it, and an
audit trail of every event on the mesh is the thing most worth having while
modules are being moved one at a time.
Each named a digest produced inside a lab that no longer exists, so none of them
could be built anywhere else. They say which module they stand on now, and there
is deliberately no default — a build nobody told stops at the declaration rather
than at a reference that resolves to nothing.
A line in postgres's recipe and a one-line file in amqp-ping, both added to
move a commit and watch the mesh notice. The proofs worked; neither was meant
to stay. MESH_MODULE is set from the sealed credential at run time anyway, so
baking it in was dead weight as well as noise.
A comment changed in a build recipe is a new commit and a byte-identical image.
Comparing commits called every module standing on it stale, so the mesh would
have rebuilt itself entirely to arrive back exactly where it started — and
listed each dependent once per commit that had produced the same image.
The base was a digest typed in by hand, for an image nothing in the mesh could
produce — so the graph held edges pointing at it with no version on the far end,
and the one change that reaches every module at once could never be noticed.
It is a module now, and these edges resolve.
Build edges are discovered by building; requires and provides are stated by the
module about itself. Both belong in the graph and answer different questions —
and "what provides postgres-database" needed a sweep over every manifest, which
only something holding all of them can do.
The catalogue expected each edge to name a module and a commit. The builder
sends a pinned image reference — it cannot know which module produced it, that
is a fact about the graph. So every build that had been built on top of anything
was rejected, and only the modules built against nothing ever registered.
Versions now record what they published, and an edge resolves through that. An
edge to an artifact no module here produced is kept: it resolves by itself when
that module is registered, which is the ordinary case while a mesh fills in.
`run` imports an entrypoint without binding a broker — it exists for a step that
works offline and exits. Both the catalogue and amqp-ping subscribe on import,
so both died on the first on() with no broker bound.
The runtime base carries what every module needs, and a database driver is not
that. Installed into an empty directory because the module's package.json also
names the sdk, which lives in the base rather than on a registry.
Its provisioner container named an image nobody could produce — a zero digest
placeholder. It names an artifact instead, and the module says how to build it,
so the mesh can make the database provider the catalogue needs.
Its account is scoped from what it emits and consumes, and it declared neither —
which is why asking for a generic module account produced one that authenticated
and could do nothing, with the refusal surfacing a layer away as a permissions
error against a queue.
Declaring the announcement is not documentation here. It is what the permission
is derived from.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module's runtime image was assembled by a script copying the sdk and the tool
runtime out of neighbouring checkouts, so it could only be built on a workstation
that had them. That is why no module declared what it was made of and why
forty-seven point at a placeholder.
The tool runtime becomes an image a module's runtime is built FROM, published like
any other artifact. The module then builds from its own directory and that base —
one clone, which is what the builder can actually be asked for (novox/hq ADR 0069).
The dependency stops being a property of somebody's machine and becomes a build
edge, pinned to a digest the mesh's registry assigned.
The compiler is invoked by its real path rather than through node_modules/.bin:
those are symlinks to a launcher that requires its library relatively, and
resolving them while building the base leaves a launcher pointing at nothing.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
It links module-versions to each other and knows nothing about nodes; which
machine runs what stays the control plane's (novox/hq ADR 0070, 0072). Keeping
them apart is what lets the control plane carry on composing declarations while
this is down.
The builder announces what it built, this places it in the graph and announces
what that means, and the control plane hooks the meaning rather than the build
output. A rebuild producing the commit already current is registered and is not
an upgrade — announcing it would ripple outward forever through modules that did
not change.
Ordering is not computed. Modules stale and waiting on nothing that is itself
stale are announced as buildable; the rest stay stale and appear once whatever
they were waiting for is registered, so a chain and a diamond need no special
handling and nothing holds a plan.
Four tools over the graph: what this mesh holds, one module in full, what a
change to a module reaches, and what must be rebuilt and why. The edges are
derived from builds rather than declared, so they cannot drift from what the code
actually uses.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx