Files
hq/04-ISSUES/143-converging-does-not-retire-the-firewall-it-found/00-report.md
T
jschoubben ddb3980f09 Issue 143: correct the diagnosis — the step exists and did not fire
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.

What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.

So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
2026-09-29 02:33:10 +02:00

5.7 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
located 2026-09-29
mesh-host internal/apply/opening.go (retireFirewall)
mesh-host internal/apply/apply.go (the condition it is called under)

143 — Converging a machine does not retire the firewall it found, and says it does

What was observed

The control-node was converged on 2026-09-29, the first machine with a found firewall to be flipped — the two converged before it had none.

The preview said, and the flip repeated:

the found firewall (ufw) is disabled, never flushed: its configuration stays on disk
...
sent: the host loads the mesh's filter and disables the firewall it found

The mesh then reported the node converged, 372 resources applied, nothing failed. Afterwards, on the machine:

systemctl is-enabled ufw   -> enabled
systemctl is-active  ufw   -> active

Corrected 2026-09-29, an hour later, from reading the host rather than the declaration. The first account of this was wrong. It said the declaration carries no resource that would disable the found firewall, and that the sentence was printed by the command with nothing implementing it. The declaration indeed carries no such resource — but the mechanism was never meant to be one. It is a step in the host's own apply, retireFirewall, and it exists, is careful, and is strict: it refuses to retire anything until it has read back from the machine that the mesh's own table is loaded, it records the forward policies first so a half-done retirement can be retried, and it verifies ufw reports inactive afterwards.

What is established is narrower and stranger than "nothing implements it":

  • ufw was active and enabled two minutes after the flip, and the flip had reported the node converged with 372 resources applied and nothing failed.
  • The machine's own record now reads disabled_by_mesh: true — but it was written by a reconcile after an operator disabled ufw by hand, roughly fifty minutes later. A reconcile found ufw already inactive, asked it to be inactive, read that back, and recorded that the mesh had done it.
  • So the step did not take effect at the flip, and the machine's record now says it did.

The candidates are named rather than chosen, because the evidence does not separate them: the step is called only when the apply had no failures, and a skipped step is silent; the mesh's table is loaded by a service in the same apply, so whether it was loaded at the moment the step asked is an ordering question; and the host's own detail lines do not reach the journal, so what it decided is not recoverable after the fact.

Why it matters beyond this instance

It is a stated behaviour that does not happen, reported as success — the fault this repository exists to catch, and ADR 0100 states it as part of what the flip is: "loads the mesh's derived filter in place of its refusal-only table, and retires the found firewall by disabling it, never by flushing".

It could only be found on the first machine that had one. The two machines converged before this had no firewall to retire, so the step had never run, and nothing reported that it had not. That is the same shape as issue 136: a step that is silent when it does nothing.

The machine is left doubly filtered, which is not what either firewall describes. Every base chain at a hook runs and a drop in any is final, so the machine now enforces the intersection of the mesh's derived filter and a rule set left by the system being replaced. Nothing is broken by that today — measured from outside, mail, the proxy and git-over-ssh answer and the databases and admin interfaces are refused — but the machine's behaviour is described by neither of the two things claiming to describe it, and the stale set includes a rule for a broker that no longer exists.

And returning the node to adopted would be wrong in the other direction. ADR 0100 says that restores the found firewall by enabling it again; enabling something that was never disabled is harmless, but the mesh's belief about which firewall is in force has been wrong in both modes.

Open questions

  • Which side owns retiring it — a resource in the declaration, so it is applied and reported like everything else, or the flip as an act? A resource seems right: the flip is otherwise entirely expressed as one, and an act that only the command performs cannot be re-checked on a later reconcile.
  • What should a reconcile do if the found firewall is enabled again by hand, or by a package update? Convergence is a state, so presumably re-disable it and say so.
  • Should the preview say what it will do rather than what it does, until a step exists that does it? The wording was read as evidence twice in one session.
  • Is there a check that a sentence the mesh prints corresponds to something that happened? This is the second time in one session that a printed claim and the machine disagreed.
  • Why did the step not take effect? It is called only when the apply had no failures, and being skipped is silent. The mesh's table is loaded by a service in the same apply, so whether it was loaded when the step asked is an ordering question — and ADR 0100 makes loading it first a precondition rather than an expectation.
  • A step that records the mesh as having done what an operator did is worse than the omission. The record now says the mesh disabled ufw. Nothing distinguishes "we did this" from "we found it already so". Should it?
  • Why do the host's own detail lines not reach the journal? Everything it decided during the flip is unrecoverable, which is why this account has candidates instead of a cause.