It declared udp/53 only. The daemon listens on tcp/53 as well, and a resolver is
asked over tcp whenever an answer will not fit in a datagram — so on every
converged machine that port is closed while the service reports itself healthy
and the manifest reads as though the resolver were fully declared.
The same fault as issue 136 in miniature: the declaration covers part of what the
service does, and the gap is silent because nothing compares the two.
The module said the service must be running and nothing about boot, so the
machine's own way back in was enabled only because something before the mesh
had enabled it. All four machines happen to be enabled today; none of them is
enabled because the mesh says so, and a machine adopted tomorrow would run ssh
until its first reboot.
Not `state: running` alone for the same reason the module exists: this is the
one daemon whose absence cannot be fixed remotely.
jail.local named ufw as the ban action. Two machines on this mesh have no ufw,
and fail2ban does not check: it starts, the jail reads the log, counts the
attempts, runs the ban command, gets 127 -- 'ufw: command not found' -- and
logs an error nobody reads. The service is active, the mesh reports the module
applied, and the machine is not protected. Proven by banning a documentation
address on such a machine today.
The replacement is this module's own dualchain action, already used by the
recidive jail on all four machines, so it is not a new dependency. It bans in
DOCKER-USER as well as INPUT, which ufw's action did not, and it bans all
ports, which ufw's action did.
The recidive jail bans whoever keeps coming back by reading fail2ban's own
log, and fail2ban checks every jail's log file while it configures itself --
before it has created that log. On a machine where the file is not there
already, no jail is found for recidive, configuration fails, and the whole
service refuses to start, taking the sshd jail with it. Two machines assigned
this module today came up failed for exactly that reason; the two where it
worked had a log from years of the service running.
Declared create-once: the mesh puts an empty file there when it is absent and
never touches it again, because what grows in it is fail2ban's, and the
logrotate file this module already ships is what keeps it small.
This also reverts the previous two commits' fail2ban.local. It declared a
logtarget that the package already sets to the same path on every machine
here -- pacman reports the config pristine -- so it fixed nothing and said
something untrue about why.
The file that says where fail2ban logs was not in restart-on, so a change to
it would sit on disk with the running service unaware of it -- the same shape
as any other jail file this module already restarts for.
The recidive jail reads /var/log/fail2ban.log and this module ships the
logrotate file for it, but nothing ever told fail2ban to write there. Where
the package default stands, fail2ban logs to the journal, the recidive jail
finds no log file, and the whole service refuses to start -- taking the sshd
jail with it. Two machines assigned this module today came up failed; the two
where it worked had /etc/fail2ban/fail2ban.conf edited by hand, which a
package upgrade would have undone.
Declared in fail2ban.local, because fail2ban.conf belongs to the package.
It asks what it missed on every start and the answer never arrived: the control plane replayed each
build it held as a module's event from a module called "control-plane", which does not exist, so its
own account refused the publish and the graph kept the gap. The control plane now states those under
the seat it holds (novox/hq ADR 0134, mesh-controller #129), so this consumes that too — one handler,
because what a build means for the graph is the same whether the build machine says it as it happens
or the mesh says what it already held.
It brought its schema up inside its runtime, on every start. That made a schema it could not reach a
crash loop rather than a stop, with the module graph keeping a gap and nothing saying so — which is
how a whole morning's builds went unrecorded. The mesh now prepares this module's state before it
starts this version and does not start it if that failed (novox/hq ADR 0135): the work moves to an
entrypoint the image names in MESH_PREPARE, beside the entrypoints it already names.
The reason it was at start — that a step blocking the apply would block the very apply bringing the
overlay up — stopped being true when a step's failure became its module's business rather than the
machine's (ADR 0136).
Every module built from a repository was rebuilt for a change to any of them: one merge in this
repository meant twenty-six builds, which is what exhausted a public registry's pull limit. The
forge lists the files a merge changed and the event carries them, from the watcher and from the
merge tool alike; a merge that changed more files than were asked for says so, and the mesh then
treats the whole repository as changed rather than guessing.
2026-09-28 09:17:38 +02:00
11 changed files with 87 additions and 10 deletions
"why":"every name for this machine and what it runs \u2014 the mesh's own answered here, the rest forwarded",
"fixed":true
},
{
"port":53,
"protocol":"tcp",
"from":"mesh",
"why":"the same names over tcp, which a resolver answers on as well and is asked for whenever an answer will not fit in a datagram. Declared because the daemon serves it: a declaration that covers one of the two protocols its own service listens on leaves the other closed while everything reports success",
"content":"[INCLUDES]\n\nbefore = paths-arch.conf\n\n[DEFAULT]\n\n# Never act on the machine itself or on a tunnel peer: the mesh's private range is\n# ${machine:mesh-range}, named here rather than written as a value the module cannot\n# know (novox/hq ADR 0112). Without this, fail2ban could ban the mesh's own nodes.\nignoreip = 127.0.0.1/8 ::1 ${machine:mesh-range}\n\nbantime = 10m\nfindtime = 10m\nmaxretry = 5\n\nbanaction = ufw\nbanaction_allports = iptables-allports\n\n[sshd]\nenabled = true\nport = ssh\nlogpath = %(sshd_log)s\nbackend = %(sshd_backend)s\n"
"content":"[INCLUDES]\n\nbefore = paths-arch.conf\n\n[DEFAULT]\n\n# Never act on the machine itself or on a tunnel peer: the mesh's private range is\n# ${machine:mesh-range}, named here rather than written as a value the module cannot\n# know (novox/hq ADR 0112). Without this, fail2ban could ban the mesh's own nodes.\nignoreip = 127.0.0.1/8 ::1 ${machine:mesh-range}\n\nbantime = 10m\nfindtime = 10m\nmaxretry = 5\n\n# Ban through iptables, not through a firewall front-end the machine may not have. ufw is\n# installed on two of this mesh's machines and absent on the other two, and fail2ban finds out\n# only at ban time: the service reports healthy, the jail counts the attempt, the ban command\n# exits 127, and nothing is blocked. Proven on 2026-09-28 -- 'ufw: command not found' on a\n# machine the mesh reported as protected.\n#\n# The action below is this module's own, already used by the recidive jail on every machine\n# here, and it bans in DOCKER-USER as well as INPUT, so a container's published port is\n# covered too.\nbanaction = iptables-allports-dualchain\nbanaction_allports = iptables-allports-dualchain\n\n[sshd]\nenabled = true\nport = ssh\nlogpath = %(sshd_log)s\nbackend = %(sshd_backend)s\n"
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.