Every module named its events the way the old bus spelled a routing key —
`module.<module>.<verb>`. Design 29 says a module names an event locally and the
mesh works out where it lands, so all 37 were stale against a rule already
decided. On the new bus that derives into a namespace belonging to a module
called "module", so no cross-module subscription in the mesh matched anything:
nothing failed, nothing reacted (novox/hq 04-ISSUES/127).
36 manifests converted, and 43 files of module code with them. The code mattered
as much as the manifests: the runtime builds the subject from what `emit()` is
handed, so a converted manifest with unconverted code would have had the
permission and the subject disagree.
Three things the new check found on the way:
- `photos` emitted an event its manifest never declared, which the new bus refuses
outright. Declared.
- `showcase` waited for an event nothing emits, so its demo could never be
triggered — only `showcase` may publish under its own name. It emits both halves
now.
- `distribution` declared an event named after a different module. It emits
`image.pushed` under its own name. An event about a *role* belongs on the seat,
where the name outlives whoever holds it, but the sdk has no way to publish on a
seat yet, so that stays recorded rather than declared.
The audit logger's "everything" pattern is `**` rather than the old bus's `#`.
The split the controller now makes, from this side. The module's own
configuration — ports, TLS, JetStream — is a declared file resource, because those
are properties of this container and change when its image does. `bus-users` names
where the mesh writes every account and permission, in the same directory, and the
module's configuration includes it.
**Both files in one directory because they have to be.** An absolute include path
is resolved relative to the including file's directory: nats-server given
`include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf looks for
/etc/nats-server/etc/nats/accounts.conf and refuses to start. Verified against the
server, and recorded in the configuration itself where somebody moving a file will
read it.
**`verify: true` is gone, and it was refusing every connection in the mesh.** It
makes the server demand a client certificate; a host pins this server's exact
certificate and authenticates with the password the mesh minted, and presents none.
Found by building this image and connecting to it as a host would.
The entrypoint now waits for both files and watches the mesh's half: the module's
own does not change without a new declaration, and that recreates the container
anyway. Verified end to end against this image — the mesh's user list rewritten,
the module noticing and reloading the server itself with no signal from outside,
and the connection the mesh already had still working afterwards.
Ten manifests claim mesh-* names now. What they PROVIDE is unchanged: gitea
still provides git and npm-package-registry, and a consumer requires the
interface, not the seat.
It claims no seat: mesh-broker is the NATS server's (novox/hq ADR 0119).
The amqp interface stays exactly as it is — a backing service a module may
require, like a database.
Step 1.1 and 1.2 of novox/hq ADR 0116. The server is a built artifact rather
than the upstream image directly, because it needs an entrypoint of its own:
the host can only recreate a container, and recreating the bus for every
permission change drops every connection and every in-flight ack. nats-server
reloads on SIGHUP by itself, so the config is mounted as a directory (not
digest-tracked, hq issue 103) and the entrypoint watches the one file.
Verified against the real server, not assumed: a user added to the config
connects, a revoked one is refused, both within one poll interval, with the
container's PID and restart count unchanged and "Reloaded: accounts" in its
log.
Two corrections found by checking rather than reading:
- the seat delivers nothing now (hq ADR 0117), and the controller's parser
refused the manifest until it did — "nats claims mesh-broker, whose holder
answers for amqp, and nats does not provide amqp"
- pinned to the multi-arch index digest; the first pin was the amd64
manifest, which builds here and fails on any other architecture
The state directories say place "." — the assignment's own root — and
every bind, secret, own-secret, receives and grants path references it
as ${dir:state}/…; grants directories that are their own resources are
placed by id. gitea's two coincidence strings from the first pass
(${dir:data}base.json — resolving correctly by pure concatenation) are
spelled honestly now. What still says /var/lib is inside containers —
the software's contract — or under /var/lib/mesh, the mesh's own
plumbing, which the requirements unification absorbs next. Every
resolved path is byte-identical to what runs; landing this is a no-op
on the node, and the converter checks its own boundaries this time.
Each resolves to <root>/mailu/<id> — the maildir at
/var/lib/mailu/data-mail, certs at data-certs, and so on. Landing this
is a window, not an edit: seventeen renames on the node (the nested
data/ tree flattens to the ids), then the full stack recreated, because
a changed volume path does not recreate a container by itself (hq 126).
Ids are untouched on purpose — a renamed id orphans its held record,
and the mail spool is the wrong place to learn what a removal step does
with one.
The module is nextcloud; its tree was /var/lib/nextcloud-module — a
historic spelling nothing depends on. The root moves to
/var/lib/nextcloud (a rename on the node, done in this change's
window), and html drops its path: the mesh resolves it to
<root>/nextcloud/html. Landing this requires the window: rename the
tree, push, recreate the container — a changed volume path does not
recreate one by itself (hq 126).
The mesh resolves both to <root>/gitea/<id> — where the 5.8G forge and
its grant files already sit, so the roll-out its upgrade policy makes
of this build changes no byte of the spec. The module root and the
mesh's plumbing stay stated.
The mesh resolves it to <root>/mongodb/data — where the granted
databases already sit. The provider's own state, grants and the mesh's
plumbing stay stated.
Seven data directories drop their paths; the mesh resolves each to
<root>/only-office/<id>, which is exactly where the data already sits —
a textual no-op on this node, and the first module speaking ADR 0112's
vocabulary. The module root and the mesh's own state stay stated.
The /services paths were the adopted-node pattern doing its job: take
replaced containers over the predecessor's data without moving a byte
(gitea set it — 'its data never moved'). With every cutover done the
exception has no reason left, and the operator called it: a nox
module's world is /var/lib/<module>, data included. Six modules
repathed; mssql keeps its /services path deliberately — it is still
held, HAL-run, and moves at its own take. Both trees are one
filesystem, so each move is a rename.
The provisioner derives a consumer's bucket from the login the mesh minted — 'derived from the
login, so teardown recomputes it with nothing to persist' — and never reads the bucket a manifest
contributed. Three modules contributed one anyway, and the value was decorative in two and wrong in
the third: photos told its container MINIO_BUCKET=photos, the predecessor's bucket, while its minted
key is scoped to mesh-novox-photos. Deployed as it stood, it would have authenticated and then been
denied on every object.
photos now names the bucket the mesh actually provisions, and the contributed bucket is gone from
all three: a value nothing reads, that reads as though it decides.
Verified against the live store before changing anything: the derived names are the populated ones —
mesh-novox-ncloud (77,886 objects, 174.9 GiB), mesh-novox-photos and mesh-novox-invoice. Nothing has
to move.
The controller now accesses /var/lib/mesh-broker-tls (mesh-controller
#54) and the push refused whole: lavinmq declared the directory as an
owned resource, and shared data is the operator's, owned by no module
(ADR 0051). lavinmq only ever reads the certs — genesis laid them down
— so it declares a read access like the controller does, and the
directory belongs to nobody.
The adapter skips what its one file shape cannot say. A body limit is the exception: the predecessor
has a buffering middleware and served its own registry name with exactly it, so this is written
rather than skipped, named after the router so the two halves cannot drift.
A limit that is not a whole positive number of bytes takes the route with it. Written without the
limit, the predecessor would carry what the module said not to carry and this module would report
success. Silence stays silence — no middleware, the predecessor's default.
The env block carried the key twice — mailu-webmail from #79's address
sweep, and a stray =webmail further down that survived it. Last write
wins in an env file, so the front resolved a name that answers nowhere
on the mesh's network and 502'd every logged-in webmail request. Latent
since the cutover: the SSO redirect the checks watched never touches
the upstream; the operator's first real login did.
Take one (#90) died on two real edge bugs, both fixed and pinned by
tests in mesh-controller (#66: autocert 404s unknown tokens itself;
#67: the internal authority 403s every public name before the token
lookup). The challenge path verified end to end reaching mailu's own
nginx before this flip.
The letsencrypt flavor served certbot's April-expired state to live IMAPS
users within minutes: autocert's HTTPHandler answers 404 itself for
tokens it does not hold and never consults the fallback for challenge
paths, so mailu's own client cannot answer through the path-scoped
route. cert flavor (valid to Nov 27) until route-proxy's handler
actually falls through.
PR #82 set TLS_FLAVOR=cert as the honest interim while the predecessor's
proxy owned /.well-known/acme-challenge outright. route-proxy took port
80 today and its handler passes unknown tokens through to routed paths
by design — the one line #82 promised, made now. The copied cert (valid
to Nov 27) stays on disk untouched; mailu's own certbot takes over from
here.
The manifest predated the working deployment on three axes: it declared a
data directory the running portainer never used (taking it would have
started empty), pinned an image digest the machine has moved past (issue
099), and contributed no route while portainer.novox.be rides a traefik
container label today. Now: the predecessor's portainer_data path, the
running image's digest, 9090:9000 kept as the predecessor's machine port
with the route contribution naming it, and 9443 kept for the runtime
sidecar's own TLS conversation.
A bare '80' tried to bind the node's port 80 — the edge's — instead of
auto-allocating. 9070 is the predecessor's number and the one the route
contribution already names.
The app reads MONGO_DB (default 'invoicing') for every operation and
uses the URL only to connect — listCollections ran against a database
the granted user cannot see. Same fault and same fix as photos' MONGO_DB,
found by the API's own logs at take.