Collided with the LB container's own name. docker inspect minio resolved
to the network instead of the (not-yet-created) container, and mesh-host's
existence check crashed on the mismatched shape rather than reporting
absence -- a real mesh-host bug (fixed separately, mesh-host#25), but this
sidesteps it here without waiting on a host-level binary update.
files-api.novox.be (port 9000, the S3 data API) and files.novox.be (port
9001, the console) — same two names HAL routes today, via nginx's own
upstream split. Needed mesh-controller#55 (a module answering one
requirement several times) to exist first; it's merged and deployed.
The single standalone instance from the first pass didn't match HAL's actual
topology: HAL runs minio1-4, two drives each, behind an nginx load balancer
on 9000 (S3) and 9001 (console). This rewrite mirrors that exactly — same
node count, same erasure-coding command, same LB config — so the migration
is a real like-for-like move, not a simplification.
Only the two images that had to change did: the minio server (dead upstream,
already fixed in the prior commit) and nginx (1.19.2-alpine is long EOL;
repinned to current stable-alpine by digest). Data still lands on a fresh,
empty, mesh-owned path, never HAL's live drives.
The OIDC-wait entrypoint wrapper HAL used is dropped: it's a no-op when
MINIO_IDENTITY_OPENID_CONFIG_URL is unset (it always is here — no OIDC
integration was ever wired to minio itself), and this catalogue has no
container resource field for overriding a container's entrypoint anyway —
every converted module relies on the image's own entrypoint plus args,
which is exactly what the original single-node version already did.
minio/minio and minio/mc were pulled from all public registries on 2026-09-11;
pgsty's fork is the working replacement (hq issue 113). The runtime sidecar
had no Dockerfile and no build section at all — added, following the same
tsc-over-client/tools/provisioner shape every other converted module uses.
The data resource pointed straight at /services/minio/data/data1-1, one of
HAL's live 8-drive erasure-coded array — starting this module would have
written into production storage the mesh doesn't own. Moved to a fresh,
empty, mesh-owned directory; the actual data migration happens over the S3
API (rclone), not by sharing a disk path.
hq ADR 0107 / issue 115. HAL's own postgres and lavinmq both used a
directory bind (./db-data, ./data) for exactly this data -- the mesh's
adoption of them, three weeks ago, switched to a named Docker volume
instead, and searxng/distribution followed the same pattern since.
A named volume survives ordinary container recreation, same as a
directory bind -- that was never the problem. The problem is everything
else: docker rm -fv, docker volume rm, and docker system prune --volumes
all target it (one flag away from the docker rm -f this migration already
uses routinely); it is invisible to every tool this migration has used all
night (ls, find, grep across /var/lib, /services); and nothing outside
Docker's own volume machinery can back it up or notice it growing.
mesh-store carries the sharpest version: every database migrated tonight,
including keycloak's, live inside it.
Data already copied and verified on novox before this merges:
- mesh-registry-data -> /var/lib/mesh-registry (11G, diff -rq clean)
- mesh-broker-data -> /var/lib/mesh-broker (37M, cp -a)
- mesh-broker-tls -> /var/lib/mesh-broker-tls (12K, cp -a)
- mesh-store-data -> /var/lib/mesh-store (1.7G) -- mesh-store stopped
cleanly first, so the final copy is crash-consistent, not a live-file
copy of a running postgres; diff -rq clean after.
- searxng/valkey: not yet assigned anywhere, manifest-only fix, nothing
to copy.
Old named volumes left in place, not deleted, as the rollback path.
hq issue 091, measured 2026-09-22: 14 of 46 modules with containers fix the
machine side of a published port in their own manifest -- a fact about one
machine (which port HAL happened to publish it on) written into a
definition meant for any node. ADR 0038 already says the mesh assigns the
machine side and a module says only what it needs; the machinery already
does it (internal/inventory/ports.go's PortFor, declaration.go's
publishedOn rewrites a bare port automatically). These 14 just never
complied.
Fixed 11 of them -- stripped to the bare software port, letting assignment
take over: de-spiegel, gitea, hello-web (the demo module), mailu, mssql,
n8n, novox.be, only-office, photos, photos-eef, photos-filip.
Left alone, the two defensible kinds the issue names: postgres/lavinmq/
distribution (foundation, genesis-rewritten per ADR 0100 -- the number in
the manifest is a default, not a claim) and unifi (protocol/
device-discovery fixes the number; nothing else can find the controller).
gitea was a live one, not just a tidiness fix: its manifest said 2222:22,
but the actual adopted, running container is on 222. Harmless while held,
but the next 'take' would have recreated it on the wrong port and broken
SSH git access. Recorded 222 as novox's own setting for it (settings set
gitea -node novox) so that doesn't happen.
The first commit on this branch embedded ${seat:mesh-store:5432} directly
inside MESH_PROVISION_POSTGRES's URL. That placeholder resolves to an empty
string whenever mesh-store is on its own default port (novox/hq
internal/catalogue/seat_into.go: 'the mesh raised on the catalogue's own
ports never gives them a setting at all') -- which produces a malformed
connection string (host:/postgres) on exactly the common case, a fresh,
non-adopted mesh. It only worked here because novox's mesh-store happens to
be adopted at a non-default port.
mesh-controller's own manifest already has the right shape for this --
internal/envfile/port.go's NAME / NAME_PORT twin, composed by
mesh-controller's Placed(): the base value keeps its own port; a separate
_PORT variable carries the override, spliced in only when it says
something, and left alone -- not a fault -- when it's an unfilled
placeholder (a manifest ahead of the running controller).
Reverted the manifest to its original base value, added
MESH_PROVISION_POSTGRES_PORT as the twin, and taught client.ts (the only
consumer -- both tools/ and provisioner/ import it) the same precedence
envfile.Placed uses. Checked: no other file in the module reads
MESH_PROVISION_POSTGRES directly.
MESH_PROVISION_POSTGRES was a literal connection string naming port 5432 --
correct only when mesh-store happens to run on the mesh's own default. On
novox, mesh-store was adopted in place at HAL's original port (6852), and
the provisioner has been retrying-and-failing against 127.0.0.1:5432 ever
since, for every consumer including ones that already exist (gitea, umami,
mesh-catalog), not just a new grant.
Fixed with the same ${seat:mesh-store:5432} template mesh-controller's own
manifest already uses for the identical connection. No other module needed
this fix checked -- lavinmq's MESH_PROVISION_LAVINMQ already used
127.0.0.1:15672 unconditionally, but the broker's management port is fixed
by the module itself (127.0.0.1:15672 in lavinmq/module.json's own ports),
not by adoption, so it isn't the same bug.
Deployed #49 and the watcher immediately broke: GET /user/repos answered 403,
'required=[read:user]' — confirmed live against the running forge (1.27.3).
That route sits under gitea's user scope category despite listing
repositories, not repository as assumed.
Also gives the fake forge real scope enforcement on /user/repos, which is
why the original PR's test suite didn't catch this: it only checked the
token's value was valid, never that it carried the required scope.
Read against hal/modules/dnsmasq-app (hal dnsmasq-app conversion, hq 08-connectivity).
On the machines it runs, the predecessor's dnsmasq answers every name: the mesh's own
itself, the rest forwarded to 1.1.1.1 and 8.8.8.8, its module's defaults; resolv.conf
names it alone at 127.0.0.1, and the container runtime's dns is the machine's tunnel
address, so the host and every container resolve the world through it. The nox module
forwarded nothing, listened on 127.0.0.55 — a convention of its own beside the one every
machine already followed — and read a machines file that the module's `facts` already
asks the mesh for, so taking it would have left an adopted machine with a resolv.conf
pointing at an address nothing answered on, and no upstream for anything else.
Now the resolver keeps `no-resolv` (the documented loop — finding its own address in
resolv.conf and becoming its own upstream — stays impossible) and forwards to the same
two explicit upstreams; listens on mesh0 and 127.0.0.1, which systemd-resolved does not
hold; requires `mesh-addressing`, since its data is the mesh's addresses; and writes the
runtime's `dns` into daemon.json beside whatever the machine had (ADR 0102), at this
machine's own address — `${machine:address}`, new in the controller. The runtime is not
restarted for it: it reads the key at start, not on reload, and a restart stops every
container; on the machine this replaces the value is already there.
resolv-conf names the resolver alone, as the predecessor's file did; its placeholder second
line was a fallback nothing ever reached. resolved-split-dns follows the address. The
mesh's suffix as a local domain comes with the machines file, so a mesh name the resolver
does not know is refused here rather than asked upstream. mDNS is not carried: no module
does it and the design says mesh names are not multicast names.
The runtime beside the forge served 0 tools: GiteaClient.fromEnv required a token
(settings or MESH_GITEA_TOKEN), nobody had one to give — the mesh raised the forge —
and putting one in settings would store a secret in plaintext in the inventory. So
the fifteen tools registered nothing and the watcher logged "not watching".
What the mesh does deliver is the admin account: a login the manifest names and a
password the vault minted and the host unsealed into a file (ADR 0086). That is
enough to mint a token, so the module does (hq issue 100, the forge's tools):
POST /users/{admin}/tokens over basic auth, scoped to write:repository and
write:issue — the least the tools and the repo watcher need — kept at 0600 in the
module's own state (/var/lib/mesh/gitea/state, a new directory resource the runtime
mounts writable), read back on the next start, and minted afresh when the forge
answers 401 to it or the kept file is gone. A forge whose data came from the
predecessor has no mesh-admin: that is reported in plain words on every poll until
it clears, once per reason, not crash-looped. A configured token still wins and is
never minted over.
The mint happens on the first call, not at registration: a contributor is
synchronous, and a forge not yet answering must not keep the runtime from serving.
One source per kept file in a process, or the watcher and the tools would each
renew on a 401 and drop the other's token by name.
Review of the registry work found it node-scoped: a second `distribution` on another
machine resolved cleanly there, and only afterwards did the mesh notice `artifact-store`
offered by two nodes, with every consumer elsewhere refusing to choose. Worse, a node-scoped
requirement with one candidate installs that candidate, so anything that wanted the store
beside it would have raised a fresh, empty store on the wrong machine first.
There is one store in a mesh, which was already the effective rule; the claim now says it
where a second one is assigned, by name, instead of leaving consumers to discover it.
Read against hal/modules/mssql: the predecessor publishes 4848:1433 and its connections
block hands every consumer localhost:4848; its data has lived in /services/mssql/data
since it was installed — 4.6 GB of it. The nox module had 1433, db-data, and a pin three
builds behind, so taking it would have moved the port every consumer was told about,
started on an empty directory, and downgraded the engine at the same time.
db-data was not a convention: only this module and mongodb used it, and both predecessors
say data. Now nothing moves at the cutover.
The pin was a month old — 2026-08-20 against the 2026-09-17 image the container runs.
Third module in a row whose pin had aged into a downgrade; the pattern is hq issue 099.
The rule the forge's cutover taught: a module that takes over a running service must not
carry an older image than the one running, or the cutover is a downgrade nobody asked for.
6.10.4 exists; moving to it is an upgrade and its own act.
A node being adopted cannot take a web module: every module reachable by
name requires route, the mesh's only provider of it binds the two public
ports, and the predecessor's proxy holds them and serves every public
name there. Stopping the predecessor to break the circle darkens every
name at once, with every certificate to re-obtain in the same window.
So this answers the same provision without binding anything. It provides
route and receives the same contributions file, and writes each
contribution as one route file where the predecessor's file provider
reads, naming the predecessor's own certificate resolver so no
certificate is asked for. It removes a file it wrote when its
contribution goes and never touches a file it did not write — the name
and a marker inside both have to say it is the mesh's.
A step, not a daemon: run-once, re-run by restart-on over the received
file and the settings. The predecessor's dynamic directory is a node
setting, because it is a fact about one machine.
Migration scaffolding with a stated end: assigned only on an adopted
node, deleted when the predecessor's proxy retires.
MESH_GITEA_URL named 3000 outright. The forge's container publishes 3000 without
fixing the machine side, so the mesh assigns it — and on a node given that port
as a setting it is the operator's number. The mapping that lets the forge go on
binding 3000 does nothing for a caller dialling the machine's loopback, so the
sidecar dialled a port nothing was listening on wherever the two differed.
${port:3000} is the mesh's answer to exactly that question, and it now resolves
in a container's environment as it always has in a file. The forge declares it
listens on 3000, so it may ask.
Still names 3000: the route contribution (hq issue 089) and the `listens` and
`ports` entries, which are the software's own number and belong there.
`at` was settable. A node setting could point the builder's package binding at
any host, and the builder sends its registry credential there as basic auth — so
a setting meant for a port was a way to hand the password to somebody else.
novox/hq 04-ISSUES/085
The forge is raised by hand at genesis, before any module provides
`package-registry`, so the builder carries a binding instead of resolving one —
and the port in it was rewritten, as text, by the installer. A later
registration of this manifest from the catalogue put 3000 back, silently, and
pointed the builder and its registry credential at whatever holds that port.
Made settable instead: the port is a per-node setting the controller holds, and
the catalogue's number is only its default. `provision`, `from` and `as` are
protected — a setting here is about where the forge answers, never about who the
binding is with.
novox/hq 04-ISSUES/085, ADR 0100
step-ca's root certificate, its key and that key's password were own secrets — random
bytes the mesh minted, which no certificate is (novox/hq 04-ISSUES/076). The mesh
mints only the CA password now; step-ca makes its root at first start and serves it
at /roots.pem, which the manifest now names beside the ACME directory. The route
proxy fetches that root over the mesh network in a run-once step before it starts,
instead of being handed a served fact that could only be written before anything ran
(ADR 0098).
The copy between registries (ADR 0096) reaches the mesh's registry over HTTP from
inside the builder's container, where loopback on the default bridge is not the
machine; docker push never noticed because it went through the machine's daemon.
The builder already holds the runtime's socket, so the host network adds nothing
it did not have.
gitea, umami, influxdb, icecast and mailu require a secret and keep each of theirs
under a local name (novox/hq ADR 0094); the broker account stays their own. The
route proxy's recipe starts FROM the bases its manifest declares (ADR 0097).
ADR 0069 put it there; genesis now takes it from the build it runs (mesh-host,
novox/hq 04-ISSUES/072). This copy was read by nothing else and had already drifted.
The official entrypoint re-executes itself as mongodb (uid 999) and only then reads
MONGO_INITDB_ROOT_PASSWORD_FILE, so a root-owned 0600 file is 'Permission denied' at line 83 and
the server never starts. secrets-owner is the mechanism ADR 0086 gives for exactly this.
From the survey of every env-file secret (ADR 0086, issue 041): amqp-ping,
minio, mongodb and grafana use the _FILE twin their software honours;
mesh-catalog and model-usage read DATABASE_URL_FILE (a file the mesh
templates, mounted where only the runtime reads it); grafana's secret files
belong to its own account. Two dead deliveries removed: a line nothing read
in amqp-email-forwarder, and mailu's secret.env on four containers that
never read it. The 25 exceptions that remain carry the surveyed reason —
convertible and awaiting a bed, convertible through a generated config file,
the application's own code, or not convertible.
ADR 0086. mesh-controller mounts its six own secrets and names them with
_FILE twins, so no credential of its own reaches its environment. The 35
containers that still read a secret through an env-file carry
secrets-in-environment with the reason; converting each where its software
accepts a path is the per-module work of issue 041.
mesh-vault provides `secret` (novox/hq ADR 0085, design 24). The value is
the pair credential the controller mints — the vault holds no copy, only a
ledger of who holds one, its fingerprint and every rotation, and two tools that
answer by fingerprint and never by value. Rotation is `rotate secret`,
unchanged machinery pointed at a secret with an owner (design 13). Named in the
mesh's own namespace, beside mesh-controller and mesh-catalog, because it is
the mesh's own code rather than wrapped software.
redis is the first consumer: its own password stops being an own-secret nothing
could rotate and becomes a `secret` it requires, read from the same file into
the same hole. The server now restarts on its config, or it would keep the
password it started with through every rotation (playbook 06).
The comment still said the provisioner is not listed and runs via a
container's args — but the 061 fix put it in MESH_TOOL_MODULES (serve
mode) and dropped the args. A future editor trusting the comment could
strip it again and silently reintroduce 061. Comment now matches the
code; only the run-once bootstrap runs via args.
minio's runtime copies the `mc` client from minio/mc:latest — a
Docker Hub pull the mesh build environment cannot make (its docker
reaches the mesh registry, not public Hub), the same isolation that
blocks npm deps. apt-based installs (mongodb's mongosh, mosquitto)
build fine because the build has real internet for apt; only npm and
Docker Hub are redirected. Delivering an external binary or image layer
into a mesh build is the same open question as the npm deps — deferred
with them.
Both carry third-party runtime deps (pg; tweetnacl + sealedbox) that
their Dockerfile installs with npm — which 404s in a mesh build, whose
npm points at the mesh's own registry, not public npm. The workstation
build script got away with it by installing on a host with public npm.
Delivering a module's third-party deps into a mesh build is an open
question (how: publish to the mesh registry, or proxy); until it is
answered these two stay on the placeholder path they were already on.
The other 37 modules build from their own directory with no external
fetch.
The 21 media/home modules, the SaaS tool modules (cloudflare-dns,
confluence, gitlab, jira), model-usage, and the model-access trio get
the same Dockerfile + build section as batch 1. Scheduled-only modules
(anthropic-consumer, anthropic-manager, openai-consumer) deliberately
declare no MESH_TOOL_MODULES — every container of theirs names its
command. mosquitto's run-once bootstrap container builds from the same
artifact. Modules with third-party deps (model-usage: pg;
anthropic-manager: tweetnacl) install them beside their compiled code.
Also fixes cloudflare-dns's package.json, unparseable since its
description lost a closing quote.
Deliberately still without build sections: builder and mesh-controller
(the foundation builds them by its own path), distribution (provides
the artifact store — building it through itself is refused by design),
route-proxy (cross-repo build context, deferred), and the
upstream-image-only modules, which have no code to build.
keycloak, mailu, minio, mongodb, mssql, nextcloud, portainer, redis,
umami and verdaccio get the Dockerfile + build section the eight
buildable modules already had; their runtime containers name the
artifact instead of a placeholder digest.
One convention, settled (060's open question, informed by 061): the
runtime container runs serve mode with every serve-time entrypoint in
MESH_TOOL_MODULES — tools serve, events flow, and a provider's
provisioner reconciles in the same process with the broker connected.
postgres, gitea and lavinmq are retrofitted from args-run provisioners,
which served no tools and emitted lifecycle events nowhere.
route-proxy is deferred: its build context is the mesh-controller
repository, a cross-repo shape the build section cannot yet express.
The runtime container named no command, so it ran the image default —
the tool host — and the provisioner entrypoint compiled beside it never
ran anywhere: no vhost was ever minted, while the grants sat applied in
its mounted directory. postgres already names its provisioner in args;
lavinmq now does the same. Surfaced by the built-store-cross-node bed,
run 9 — the first bed to reach the vhost assertion honestly.
The one-line change e0c9219 parked "until there is a certificate" returns — with the
overlay recorded as the registry's transport security and every node's runtime told the
store speaks plain HTTP, a reference under the provider's internal name is one every
machine can pull. References minted at genesis stay loopback and are valid where they
matter, on the machine that made them.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
postgres claims mesh-store, lavinmq claims mesh-broker, and the controller's seat is
renamed the-controller -> mesh-controller so all three follow one convention. The resolver
refuses a second holder mesh-wide, so an adopted foundation module assigned to a second node
is refused rather than silently raising a second server. Closes hq issue 056 (ADR 0079).
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
mesh-broker serves the amqps bus on 5671 (the control plane and every module's
events) but the lavinmq module declared only 5672, so the firewall's forward
chain — where the broker's published ports are matched — never opened 5671, and
a consumer on another node could not reach the bus over the overlay. Part of
issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — name mesh-broker, the same
pinned upstream lavinmq image, the same TLS args, ports and volumes — so the
applier adopts it in place. The second server, the lavinmq.ini bootstrap and the
module's own network are gone. The provisioner is host-networked to the broker's
loopback management (127.0.0.1:15672) and authenticates as lavinmq's default
guest, which the foundation broker runs with; the mesh bus stays on the / vhost,
a vhost-per-consumer beside it. One lavinmq now.
Issue 051 (WBS 3.2).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
server is now the container the foundation raised — same name (mesh-store),
same env (POSTGRES_PASSWORD/PGDATA), ports (127.0.0.1:5432:5432), volume
(mesh-store-data) and the same pinned upstream postgres image the foundation
runs — so the applier adopts it in place rather than raising a second postgres.
The provisioner is host-networked to reach the loopback store at 127.0.0.1:5432.
The module's own network and bind-mounted data dir are gone; there is one
postgres now, holding the controller's contexts and every module's database.
Issue 051 (WBS 3.1).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A docker run -v from inside the builder resolves the path on the host: with the
workspace mounted at a different path inside than out, the SDK publish and any
bundle compile mounted an empty directory. Bind it at the same path both sides.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
gitea gains the package-registry provision: serves/receives/grants, an admin
own-secret, a postgres-shaped build, and a provisioner that creates a gitea user
per consumer with the mesh-minted password and seals nothing (hq ADR 0048). The
builder takes its registry credential as an own-secret rather than a resolved
provision, because gitea-as-module needs the base to build its provisioner and so
cannot resolve before the base — a cycle the own-secret avoids.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
dnsmasq still required resolver-data, a provision that died with the
mesh-resolver module — assigning it would refuse with "nothing provides
resolver-data". It asks for the node-zones fact now, at the same path its
config already reads, restarting on the fact's own id.
gitea and verdaccio both provide package-registry now — ADR 0075's provision,
which neither declared, so ADR 0014's "consumes from the private registry" had
no provider anywhere in the catalogue. Two providers, mesh-scoped: the resolver
refuses until one is assigned, and choosing is assigning, which is the designed
shape.
audit-logger runs a container and declared no capability, alone among the
containerised modules. A machine without a runtime would have been assigned it
and failed at apply rather than at assignment.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module's identity is the software it is (ADR 0040). Two were named after the
job instead, and the job already had a name.
firewall installs the nftables package and runs nftables.service. The seat it
claims is the-packet-filter, which is correctly named for the role. Calling the
module firewall named neither the software nor the provision, and promised that
any firewall could sit there — the false genericity the naming rule forbids.
registry runs Distribution, the OCI reference implementation, and provides
artifact-store. So registry was a third name for a thing that already had two,
which is how one word ended up meaning the module, the software and the concept
in the same paragraph.
The capability stays firewall, and correctly: a capability IS a functionality, so
a node having one and fail2ban requiring one are both right. Only the module
moves.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Over-corrected: taking the port out of listens as well as serves made the module
declare it listens on nothing, and the parser said so.
ADR 0038 splits it. A module names the port its own software listens on, because
that is a fact about the software and it knows it. The mesh assigns the
machine-side number, because only the mesh knows what else is on the machine, and
it is the mesh that fills the assigned number into serves so a consumer is told
one number rather than three that agree by luck.
So what was wrong was writing a port into serves, not into listens.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Written so the module system has something that proves itself rather than a claim
about what it supports, and guarded by a test in the catalogue's own suite so it
cannot quietly stop exercising things.
Nine of the host's eleven resource kinds, all three artifact kinds including the
one that compiles, all three ways a module's code can run, and all four things
that code can be: tools, an event consumer, a provisioner, and processes.
Two absences that are findings rather than gaps. `action` is refused to modules
outright — the link may not carry a command to run (ADR 0005), so a module that
needs something done ships a program that reconciles, which is what a run-once
process is. `service` puts an EXISTING unit into a state and installs none, which
is right for software shipping its own; code the mesh built has no unit until the
mesh writes one, and that is a process.
And it no longer picks its own port. ADR 0038 says a module cannot know what else
is on the machine it was assigned to, and names exactly the trap this fell into:
the number written three times — listens, serves, a container's ports — agreeing
only because one person wrote all three, with nothing checking. So it says what
it needs and the mesh assigns the number.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A replayed build is registered exactly as any other and announced to nobody. A
module that moved months ago is not something anything should act on now:
emitting `upgraded` would have the control plane decide about a rollout, and
`rebuild-needed` would ask for builds of things already current.
Asked on every start rather than only the first, because a catalogue cannot tell
whether it has a gap — and the answer is idempotent, so asking when there is none
costs a message. Asked after subscribing, so a build arriving during the replay
is not lost between the two.
Closes novox/hq 04-ISSUES/050 with mesh-control.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Naming it from the binding was right and arrived too early. The moment the
machine had a name, the builder pushed to <node>.internal:5000 and the runtime
refused it: "http: server gave HTTP response to HTTPS client". The registry
serves plaintext, and anything that is not loopback is required to be HTTPS.
So there are two phases, and this is the first. Before the mesh has a certificate
authority of its own, loopback is the only trusted path that is honest — it is
trusted because it cannot leave the machine, not because anyone checked
anything. The mesh-reachable name belongs to the second phase, with TLS from the
mesh's own CA, and the binding expression returns then.
Not a revert of the reasoning: novox/hq issue 048 stays open and this is why. The
same one-line change lands again once a certificate module is running.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx