5.9 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| to-be | in-progress |
|
2026-10-02 |
|
31 — A module declares its fail2ban jail, and the mesh composes them per node
A node's intrusion filter should be composed from the modules it runs, the same way its firewall
is. The mesh already derives a node's nftables ruleset from every assigned module's listens and
guards (the Filtering mechanism). fail2ban is the same shape and is not modelled: a module that
runs an authenticating service — postgres, mssql, mailu — has a jail (a filter that reads its log
and a jail stanza that bans on it), and which jails a node's fail2ban runs should be exactly the
jails of the modules assigned to that node.
The predecessor did this with per-module files: postgres shipped postgres-auth.conf, mssql
shipped mssql-auth.conf, mailu shipped mailu.conf, and the node's fail2ban read whichever were
present. When HAL retired on novox those became dangling symlinks — fail2ban ran the jails only from
memory, and a restart would have dropped them. The base was salvaged (the fail2ban module now ships
sshd, recidive, and the ignoreip that spares the mesh's own range), but the service jails
are gone, because no nox module declares one yet.
The shape
- A module declares its jail in its manifest, naming no node and no path (ADR 0112): the filter
(the failregex, or a stock filter it uses) and the jail stanza (port, logpath, maxretry, bantime).
The
postgresmodule says what a postgres brute-force looks like and how to ban it; it does not say on which machine, because it does not know. - The mesh composes them per node. For each node, the jails of its assigned modules are gathered
and written into the fail2ban holder's
jail.d/(and filters intofilter.d/), exactly aslistens/guardsare gathered into the node's firewall. So a node running postgres gets the postgres jail; a node not running it does not. Thenode-intrusion-preventionholder receives them the way a provider receives its consumers' contributions. - The base stays the fail2ban module's:
sshd,recidive, and theignoreipnaming${machine:mesh-range}so a tunnel peer is never banned.
Why this, and not the module writing the file itself
A module could declare a file resource at /etc/fail2ban/jail.d/<x>.conf directly. Rejected: the
path is the fail2ban holder's to own (one module owns jail.d, as one module owns the firewall
table), the jail's logpath and defaults want the mesh's composition (the ignoreip, the ban action
the node uses), and two modules writing into one directory is the collision the seat/holder model
exists to prevent. The module declares what its jail is; the holder's composition decides how it
lands — the same split as listens (the module says the port; the mesh says the rule).
Why now
fail2ban on novox currently runs the service jails from memory only; the next restart drops them
(the ignoreip is safe on disk, so the mesh-partition risk is closed, but postgres/mssql/mailu
auth-banning would be lost). This is the mechanism that restores them properly, and it is needed as
each of those modules migrates to the other nodes — ace running postgres should get the postgres
jail, composed from the postgres module's manifest, without anyone editing a node.
References
- ADR 0112 — a module
names no node or path; its jail is declared the same way its
listensare - mesh-controller
internal/catalogue/adoption.go(Filtering— the firewall composition this mirrors),internal/catalogue/manifest.go(Listens/Guards, the fields a jail field sits beside) - mesh-catalog
modules/fail2ban(the base: sshd, recidive, ignoreip); the service modules (postgres,mssql,mailu) that will declare jails
Decided and built, 2026-10-02
ADR 0179
made this the rule and built it. A module declares jails — each a name, the failregex of a
failed attempt in its log, and the stanza's own keys — and the fail2ban module declares jailing:
the one file the stanzas compose into and the directory each filter lands in. The controller gathers
every assigned module's jails per node into those; the holder's daemon restarts on the composed file.
What made it workable was the log. A container's output went to a file of the runtime's own, under
a path that changes when the container is recreated, so no jail could read a container's service
however it logged. A container now declares logging: journald, the host runs it with the journal as
its driver, and a jail reads it with backend = systemd and a journalmatch on the container's
name — the same way the base's ssh jail has always read the ssh daemon. The first three doors: the
mail front end (every login failure on its proxying ports), the forge (a failed authentication
attempt) and the public proxy (a certificate or request for a name the mesh does not serve, which
the proxy now says in its log). The base is strict — three in a day for a day; twice banned in two
weeks for four — and the mesh's own range stays never banned.
The seat the module holds serves status, banned, ban and unban, from a runtime that carries
only the fail2ban client with the daemon's socket shared in; the jails are composed, the ban list is
the daemon's, and both are read through the console.
How it is checked: ADR 0179's table.