The recidive jail bans whoever keeps coming back by reading fail2ban's own
log, and fail2ban checks every jail's log file while it configures itself --
before it has created that log. On a machine where the file is not there
already, no jail is found for recidive, configuration fails, and the whole
service refuses to start, taking the sshd jail with it. Two machines assigned
this module today came up failed for exactly that reason; the two where it
worked had a log from years of the service running.
Declared create-once: the mesh puts an empty file there when it is absent and
never touches it again, because what grows in it is fail2ban's, and the
logrotate file this module already ships is what keeps it small.
This also reverts the previous two commits' fail2ban.local. It declared a
logtarget that the package already sets to the same path on every machine
here -- pacman reports the config pristine -- so it fixed nothing and said
something untrue about why.
The file that says where fail2ban logs was not in restart-on, so a change to
it would sit on disk with the running service unaware of it -- the same shape
as any other jail file this module already restarts for.
The recidive jail reads /var/log/fail2ban.log and this module ships the
logrotate file for it, but nothing ever told fail2ban to write there. Where
the package default stands, fail2ban logs to the journal, the recidive jail
finds no log file, and the whole service refuses to start -- taking the sshd
jail with it. Two machines assigned this module today came up failed; the two
where it worked had /etc/fail2ban/fail2ban.conf edited by hand, which a
package upgrade would have undone.
Declared in fail2ban.local, because fail2ban.conf belongs to the package.
It asks what it missed on every start and the answer never arrived: the control plane replayed each
build it held as a module's event from a module called "control-plane", which does not exist, so its
own account refused the publish and the graph kept the gap. The control plane now states those under
the seat it holds (novox/hq ADR 0134, mesh-controller #129), so this consumes that too — one handler,
because what a build means for the graph is the same whether the build machine says it as it happens
or the mesh says what it already held.
It brought its schema up inside its runtime, on every start. That made a schema it could not reach a
crash loop rather than a stop, with the module graph keeping a gap and nothing saying so — which is
how a whole morning's builds went unrecorded. The mesh now prepares this module's state before it
starts this version and does not start it if that failed (novox/hq ADR 0135): the work moves to an
entrypoint the image names in MESH_PREPARE, beside the entrypoints it already names.
The reason it was at start — that a step blocking the apply would block the very apply bringing the
overlay up — stopped being true when a step's failure became its module's business rather than the
machine's (ADR 0136).
Its runtime announces that it has just started and may have missed builds — the event the control
plane follows to replay them — and its manifest did not declare it. A module's authority on the bus
is derived from what it declares, so the publish was refused and the runtime died on start, in a
loop, with the mesh's graph never catching up.
Every module built from a repository was rebuilt for a change to any of them: one merge in this
repository meant twenty-six builds, which is what exhausted a public registry's pull limit. The
forge lists the files a merge changed and the event carries them, from the watcher and from the
merge tool alike; a merge that changed more files than were asked for says so, and the mesh then
treats the whole repository as changed rather than guessing.
An old merge past the first page of the forge's listing surfaced as newer pull requests were
updated, and was announced as if it had just happened; the mesh then rebuilt everything built from
that repository, once per old merge. The moment the watching began is kept with the record, and
only a merge made since is announced.
A trigger that fires silently is indistinguishable from one that did not fire (novox/hq issue
131); the announcement is now one line an operator can read.
/user/repos lists what the token's user owns, which for the mesh's administrator is nothing — so
the forge module watched an empty list and never announced a merge. It reads the forge's whole
view through the search endpoint, every page.
The merge tool emitted at the instant it acted; a merge made in the forge's own
pages or over its API emitted nothing, and the mesh went on believing every module
current with its source (novox/hq 04-ISSUES/131). Merged pull requests are now
watched the way repositories are: what the forge holds, asked for on a tick,
announced once, with the merge commit and the clone URL a build needs. What has
been announced is kept beside the module's state, so a restart does not announce
the whole history again, and a first tick with no record announces nothing.
The mesh's broker certificate is the operator's, kept outside any module and
mounted read-only by whatever serves the bus; the controller declares that
access and now so does nats, the same way. A mount nothing declares is refused
at registration (ADR 0030), which is how this was found.
The module mounted a TLS directory nothing fills, so the server could not start
on a mesh that was not raised by the genesis template. The mesh already has a
broker certificate every machine pins by fingerprint and the controller trusts;
serving the new bus with it means no pin changes when a machine moves and there
is no second certificate to be wrong about. No ca_file: the directory has none,
and a pinning client checks the leaf and nothing else.
The recipe started FROM the upstream server's digest directly, and the build machine
refuses that: every base is declared under build.on and copied into the mesh's own
store before a build, so a build never reaches out to a registry the mesh does not
run (novox/hq ADR 0097). Found the first time the module was built on a real mesh.
Same digest, now declared as NATS_BASE and arriving as a build argument; the
Dockerfile says where it comes from and why the digest is the index's.
AMQP is not a provision (novox/hq ADR 0131, design 28 task 5.4). These were the
only three manifests that named it: the broker that provided it, a proof that a
grant worked end to end, and a forwarder reading mail off a queue. Removed, not
converted — a module that wants messaging wants the mesh's bus, reached through
the sdk and named by the mesh-broker seat, and either of the last two is re-done
against that if wanted, as a new module under the record.
The controller refuses a manifest naming amqp from its next release, so these
could not be re-registered anyway. Nothing else in the catalogue referenced them.
The claim comes back for the third and last time. The store's row says the bus seat
answers for `amqp`, lavinmq provides `amqp`, and with mesh-controller#89 the check
that judges a claim reads that row instead of a copy compiled into the build
machine. So this is accepted for the reason it should have been all along.
Restores the holder the controller composes its own bus address through, which is
what ends tonight's crash loop. nats takes the seat over when the cutover is done
deliberately, not because the seat emptied itself.
Putting the claim back was wrong on its own terms. `mesh-broker` delivers
`mesh-bus`, and a seat that delivers a provision may only be held by a module that
provides it — so the claim is refused at registration:
lavinmq claims mesh-broker, whose holder answers for "mesh-bus",
and lavinmq does not provide "mesh-bus" at mesh scope
Which means the merged claim does not restore the holder, it stops lavinmq being
built at all. Removed again.
The seat being empty is still the live fault, and it has only one valid answer: the
holder must provide `mesh-bus`, and the module that does is nats. Recorded against
the rollout, because it moves a step that was optional into the critical path.
Taking the claim off made the seat unheld, and the controller dereferences that
seat to find its own bus (to-be 26, "the one exception is the controller itself").
Unheld, the composed address fell back to a default port nothing serves, and the
control plane crash-looped: "cannot reach the broker named in MESH_BROKER_AMQP:
dial tcp 127.0.0.1:5672". The broker itself never stopped — it is healthy on the
port the mesh actually assigned it.
lavinmq becoming an ordinary provider is right, and it is still a provider of amqp
here. What was wrong is the order: the seat has to pass from one holder to the next,
and it cannot be empty in between, because the thing that reads it is the thing that
would have to fix it.
builder and route-proxy both build from the mesh-controller repository's context,
so its go.mod is theirs, and `nats.go v1.54.0` puts that at `go >= 1.26`. Pinned at
1.25.14 they cannot compile it: the build machine's own build failed with "go.mod
requires go >= 1.26.0 (running go 1.25.14)".
Each moves to the 1.26.8 digest of the flavour it already used — alpine for
builder, debian for route-proxy — so nothing changes but the compiler version.
Two lines of work renamed the same seats differently. The trunk named them for their
scope — node-scoped ones `node-*`, leaving `the-artifact-store`, `npm-package-registry`
and `git` as they were — and this branch had renamed ten of them to `mesh-*`. The trunk's
set is what the live controller loads and what the live seats were actually renamed to, so
a manifest claiming this branch's name is one the running mesh refuses. Three of them
needed reverting by hand: git had auto-merged this branch's names where the trunk had not
touched those lines, which is the quiet kind of merge result.
Event names are this branch's, because the trunk has not converted them and they are what
issue 127 was about.
Verdaccio goes with the trunk's removal of it. The template work on dnsmasq's roster fact
is the trunk's, sitting beside this branch's local event names in the same file — the one
hunk where both changes landed together.
75 manifests, all parsing, no claim outside the trunk's set and no event name left in the
old bus's form.
The jail.local [DEFAULT] gains ignoreip = 127.0.0.1/8 ::1 ${machine:mesh-range}
— localhost plus the mesh's own private range, named through the placeholder
rather than hardcoded (data is the mesh's, ADR 0112). Without it fail2ban could
ban the mesh's own nodes on 10.10.0.0/24; on novox that rule survived only in
memory from a now-deleted HAL file and would be lost on the next restart.
ADR 0121. The builder declared `built` as its own event, so every consumer depended
on which module happens to be the build machine today. It is the build-machine
role's event now: the builder declares none of its own, and the catalogue listens
for `mesh-build-machine.built` rather than `builder.built`.
Nothing changes about what reaches the catalogue. What changes is that it survives
the build machine being a different module, which is the whole reason the mesh has a
word for a role.
Every module named its events the way the old bus spelled a routing key —
`module.<module>.<verb>`. Design 29 says a module names an event locally and the
mesh works out where it lands, so all 37 were stale against a rule already
decided. On the new bus that derives into a namespace belonging to a module
called "module", so no cross-module subscription in the mesh matched anything:
nothing failed, nothing reacted (novox/hq 04-ISSUES/127).
36 manifests converted, and 43 files of module code with them. The code mattered
as much as the manifests: the runtime builds the subject from what `emit()` is
handed, so a converted manifest with unconverted code would have had the
permission and the subject disagree.
Three things the new check found on the way:
- `photos` emitted an event its manifest never declared, which the new bus refuses
outright. Declared.
- `showcase` waited for an event nothing emits, so its demo could never be
triggered — only `showcase` may publish under its own name. It emits both halves
now.
- `distribution` declared an event named after a different module. It emits
`image.pushed` under its own name. An event about a *role* belongs on the seat,
where the name outlives whoever holds it, but the sdk has no way to publish on a
seat yet, so that stays recorded rather than declared.
The audit logger's "everything" pattern is `**` rather than the old bus's `#`.
Claims renamed to match the controller's seat set: node-dns-resolver (dnsmasq),
node-intrusion-prevention (fail2ban), node-packet-filter (nftables),
node-resolver-config (resolv-conf, resolved-split-dns), node-uplink
(networkmanager, systemd-networkd, dhcpcd), mesh-build-machine (builder, +mesh
scope), mesh-catalog (mesh-catalog). showcase now declares its own seat and
claims it. verdaccio removed — the mesh keeps distribution as its registry and
gitea already serves npm, so a second npm registry is redundant.
The split the controller now makes, from this side. The module's own
configuration — ports, TLS, JetStream — is a declared file resource, because those
are properties of this container and change when its image does. `bus-users` names
where the mesh writes every account and permission, in the same directory, and the
module's configuration includes it.
**Both files in one directory because they have to be.** An absolute include path
is resolved relative to the including file's directory: nats-server given
`include /etc/nats/accounts.conf` from /etc/nats-server/nats.conf looks for
/etc/nats-server/etc/nats/accounts.conf and refuses to start. Verified against the
server, and recorded in the configuration itself where somebody moving a file will
read it.
**`verify: true` is gone, and it was refusing every connection in the mesh.** It
makes the server demand a client certificate; a host pins this server's exact
certificate and authenticates with the password the mesh minted, and presents none.
Found by building this image and connecting to it as a host would.
The entrypoint now waits for both files and watches the mesh's half: the module's
own does not change without a new declaration, and that recreates the container
anyway. Verified end to end against this image — the mesh's user list rewritten,
the module noticing and reloading the server itself with no signal from outside,
and the connection the mesh already had still working afterwards.
The node-zones fact was a path; the local=/address= syntax lived in the
control plane. It is dnsmasq's configuration language, so it moves into
dnsmasq's manifest as a template over the roster. The mesh renders it; it
reads none of it. Output is unchanged.
Lands with mesh-controller's ADR 0120 change — the two are one schema step.
networkmanager, systemd-networkd and dhcpcd each claim the-uplink and
declare only what keeps the machine's own network manager from
contradicting the mesh: the resolver file left to resolv-conf, mesh0
left alone. Never a link, profile or credential — the link is the
mesh's only channel to the machine, so NetworkManager and networkd are
reloaded on a change, never restarted, and dhcpcd (no reload; a restart
drops the address) takes its block at its next start.
Ten manifests claim mesh-* names now. What they PROVIDE is unchanged: gitea
still provides git and npm-package-registry, and a consumer requires the
interface, not the seat.
A workstation's job includes names that are neither a mesh machine nor
a routed name (novox/hq 122): shanks carries 13 Mediahuis entries in
/etc/hosts, and mesh-wireguard replaces /etc/hosts whole when taken —
so without this they vanish, and the take gates the node. Two homes,
neither the mesh's to own: conf-dir=/etc/dnsmasq.d/,*.conf (drop-in
directives, HAL's dnsmasq-app used exactly this) and
addn-hosts=/etc/hosts.local (plain host lines). The mesh creates and
rewrites neither; a machine with none loses nothing. The migration
moves such names here BEFORE the /etc/hosts take, closing the window.
It claims no seat: mesh-broker is the NATS server's (novox/hq ADR 0119).
The amqp interface stays exactly as it is — a backing service a module may
require, like a database.
The stated path was the adopted-data exception; with the take done and
the placement vocabulary live, the exception has no reason left. The
landing window renames the directory and recreates the container, since
a changed volume path does not do that by itself (hq 126).
The spec is the working system: HAL's 99-hal.conf, restated as
10-mesh.conf so lexical include order makes the mesh's answer the one
that wins while the predecessor's file is still on disk. Subsystem
stays the stock config's — first-set wins and it sits before the
Include. Port 22 from anywhere, said in listens with its reason: the
machines that need the door are exactly the ones not on the mesh yet,
and locking the operator out is the one failure a firewall must never
arrange.
The file is shared — the operator's insecure-registries for the mesh's
own store live there — and replacing it whole would break every pull
from that store the moment the module is taken (the 098 class, caught
in the pre-take diff this time). ADR 0102's verb is merge.
Step 1.1 and 1.2 of novox/hq ADR 0116. The server is a built artifact rather
than the upstream image directly, because it needs an entrypoint of its own:
the host can only recreate a container, and recreating the bus for every
permission change drops every connection and every in-flight ack. nats-server
reloads on SIGHUP by itself, so the config is mounted as a directory (not
digest-tracked, hq issue 103) and the entrypoint watches the one file.
Verified against the real server, not assumed: a user added to the config
connects, a revoked one is refused, both within one poll interval, with the
container's PID and restart count unchanged and "Reloaded: accounts" in its
log.
Two corrections found by checking rather than reading:
- the seat delivers nothing now (hq ADR 0117), and the controller's parser
refused the manifest until it did — "nats claims mesh-broker, whose holder
answers for amqp, and nats does not provide amqp"
- pinned to the multi-arch index digest; the first pin was the amd64
manifest, which builds here and fails on any other architecture
A consumer connecting by the binding's address meets a certificate for
mail.novox.be and refuses it — found live by the forwarder's cutover
proof, one send before production would have. The TLS name is mailu's
own fact (HOSTNAMES), so the binding carries it; consumers say
${bound:smtp:name} and verification holds.
The state directories say place "." — the assignment's own root — and
every bind, secret, own-secret, receives and grants path references it
as ${dir:state}/…; grants directories that are their own resources are
placed by id. gitea's two coincidence strings from the first pass
(${dir:data}base.json — resolving correctly by pure concatenation) are
spelled honestly now. What still says /var/lib is inside containers —
the software's contract — or under /var/lib/mesh, the mesh's own
plumbing, which the requirements unification absorbs next. Every
resolved path is byte-identical to what runs; landing this is a no-op
on the node, and the converter checks its own boundaries this time.
Each resolves to <root>/mailu/<id> — the maildir at
/var/lib/mailu/data-mail, certs at data-certs, and so on. Landing this
is a window, not an edit: seventeen renames on the node (the nested
data/ tree flattens to the ids), then the full stack recreated, because
a changed volume path does not recreate a container by itself (hq 126).
Ids are untouched on purpose — a renamed id orphans its held record,
and the mail spool is the wrong place to learn what a removal step does
with one.
The module is nextcloud; its tree was /var/lib/nextcloud-module — a
historic spelling nothing depends on. The root moves to
/var/lib/nextcloud (a rename on the node, done in this change's
window), and html drops its path: the mesh resolves it to
<root>/nextcloud/html. Landing this requires the window: rename the
tree, push, recreate the container — a changed volume path does not
recreate one by itself (hq 126).
The mesh resolves both to <root>/gitea/<id> — where the 5.8G forge and
its grant files already sit, so the roll-out its upgrade policy makes
of this build changes no byte of the spec. The module root and the
mesh's plumbing stay stated.
The mesh resolves it to <root>/mongodb/data — where the granted
databases already sit. The provider's own state, grants and the mesh's
plumbing stay stated.
Seven data directories drop their paths; the mesh resolves each to
<root>/only-office/<id>, which is exactly where the data already sits —
a textual no-op on this node, and the first module speaking ADR 0112's
vocabulary. The module root and the mesh's own state stay stated.
The /services paths were the adopted-node pattern doing its job: take
replaced containers over the predecessor's data without moving a byte
(gitea set it — 'its data never moved'). With every cutover done the
exception has no reason left, and the operator called it: a nox
module's world is /var/lib/<module>, data included. Six modules
repathed; mssql keeps its /services path deliberately — it is still
held, HAL-run, and moves at its own take. Both trees are one
filesystem, so each move is a rename.