A whole-mesh push refused outright the moment any single node failed to resolve:
every node was composed first, and one entry in `refusals` returned "nothing was
sent" before the send loop ran. So on the ADR 0056 lab, one module on the anchor
requiring a provision nobody had assigned a provider for — `nothing provides
"acme-ca", wanted by route-proxy` — stopped every OTHER machine from being sent
anything. Nothing converged anywhere, and the machines that went unconverged were
the ones with nothing wrong with them. The failure and the punishment were on
different machines.
This is the rule a92c11b established one level down, where an un-hostable module
stopped taking down the healthy modules beside it, applied one level up: the blast
radius of a fault is the thing that has it. A node whose declaration composes is
sent; a node whose does not is named, with its reason, and the push still ends
non-zero — skipping is not succeeding, and a command that exits cleanly having
missed a machine is a command that lies. The message now says how many machines
WERE sent, because "nothing was sent" was the claim that had become untrue.
sendTo keeps its all-or-nothing rule, and its comment now says that push
deliberately does not share it: a rotation reaching the consumer and refusing on
the provider leaves one end holding a credential the other has never heard of,
which is a real coupling between two named machines. A whole-mesh push has no such
coupling and never did.
Composing is split out of pushCommand so the rule can be tested without a broker
and a database.
novox/hq ADR 0056
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module a person assigns to a machine that cannot host it — its declared
capability has no detector there, as fail2ban does on a host with no firewall —
made Resolve refuse the entire node, so a whole-node push refused to send the
healthy modules beside it too. One module on the wrong machine took down every
other module on that node.
Assign already keeps such an assignment on purpose (it is what a person meant,
and acts.go says so), so the fix is on the resolve/push side: a directly-assigned
module the machine cannot host is left out of the closure and reported as
un-applied on the Resolution, rather than refusing the set. The healthy modules
still resolve, declare, and converge. A module that is *required* by something
running here and cannot be hosted still refuses — that set is genuinely
incoherent — so the distinction is who wanted it.
assign, plan and push now name the un-applied module and the missing capability,
via a shared WrongMachine message, so it is neither silently dropped nor fatal.
Reconciled two tests that encoded the old whole-node refusal for directly-assigned
un-hostable modules; added coverage for the healthy-modules-still-converge case
and the required-un-hostable-still-refuses distinction.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
`status` hung. It composes a declaration for every node to answer *is
this machine running what I would send it*, and composing one assigns
each module a machine port — so the question wrote to the database, and
wrote to the same rows as the machine it was asking about.
`port_assignment` is unique on (node, machine). Two transactions
inserting the same port do not race, they queue: the second waits on the
index until the first commits. A status polled every two seconds while a
node applies is two writers on those rows, and the poll stopped
returning rather than returning something wrong — which is the better
failure of the two, and still a failure.
The latent version of this was there before anything polled: two
compositions running at once could both allocate.
So allocation belongs to the send path alone. The mesh chooses a port
when it commits to sending one; every other caller reads what was
chosen. A module with nothing assigned has never been sent, which is
precisely what "waiting" means — the read needs no number to be right
about that, and inventing one would make the answer worse.
Named rather than passed as a bare bool: at three call sites, `true` and
`false` say nothing about which of these two things is meant.
Checked by the lab, which now polls status throughout an apply.
Working out what a machine should be reaches the identity context for its
certificate and the licence context for its model access. Both were opened —
and waited on — inside functions called for every node in a push. Two machines
hid it. Fifty would be fifty connect-and-wait cycles for data that does not
change while the push runs.
So a command holds what it has open, and passes it. Each context is opened on
first use rather than up front, because most commands need one and paying to
reach three would be the same waste from the other side.
The contexts stay separate, which is the point: this is one struct holding
three connections to three databases, not one connection to a shared one. No
context reaches another's store, and each still holds only its own credential
(novox/hq ADR 0008).
A pure move again — the gate is green before and after, and no test changed.
2,769 lines and 59 functions, holding command parsing, store opening,
resolution, the board, rotation, licences, builds and status rendering.
Nothing in it was wrong. It grew because appending was always the cheapest next
step, and no single edit was the one that should have been a new file.
That is exactly how novox/hq ADR 0001 records `hal/sdk` reaching 155 files and
34,636 lines — "containing code from every context", with each addition
avoiding a cycle and none of them the mistake. This is the same shape at 8% of
the size, which is why it is worth doing now rather than noting.
Eight files, along boundaries that already existed: what a machine is; the
private network; the catalogue; working out what one machine should be; sending
it; builds; the three questions; and reaching each context's store. main.go
keeps what a main is for — parsing arguments and dispatching.
A pure move. No behaviour changed, no test changed, and the gate is green
before and after — which is the only thing that makes a refactor this size
safe to do in one commit.