Review of the fix for hq issue 104 found three faults in it. A file applied
on an enrolled node — the mesh's own last declaration included — is applied
as the bundle is, so its resources are recorded as the machine's own and
what the mesh declared reads as undeclared: the plan removed the foundation.
`apply FILE` is for a machine the mesh has not spoken to, and is now refused
saying so whenever declared.json exists. The plan looked at what is held
before what the declaration says is taken, so the one cutover ADR 0100 says
must be previewed read as a hold; it now decides in holdOnAdopted's order,
models a step run inside a held container, and a test holds the plan's
sequence to the apply's outcomes. Genesis wrote the mode on every run, so a
re-run after `converge` left the state saying adopted while the kept,
signed declaration said converged, and the reconcile loop refused every five
minutes with no delivery coming to end it: genesis now writes the mode only
when none is recorded, and where the state and the verified kept declaration
disagree, the kept declaration wins and the repair is said.
Also: a file lock beside the state, taken by the link service, the host's
own commands and the installer alike, so a `reconcile` run by hand no
longer races the loop's save — chosen over refusing while a named service is
active, which would miss a `mesh-host run` started by hand; `--json
--dry-run` emits {plan} like an apply emits {plan, report}; the README's
duplicate flag line; and the bundle refusal is about the digest, not a claim
the carried bytes can never match what genesis applied.
An operator ran `mesh-host reconcile` on an adopted control-node with twelve
modules assigned. It applied the bundle the host carries — the genesis
declaration, foundation only, converged: recreated the store, failed on the
broker's held port, wrote the converged base filter and started its service,
and stopped at the first failing action. The filter closed the machine for
forty-five minutes. The host reported the node adopted in every report, the
declaration said converged, and nothing compared the two; nothing was printed
before acting (hq issue 104).
The host now records the node's mode — from every declaration the mesh sends,
and at genesis from what the operator said — and refuses, at the point of
application, a declaration that says the other mode, naming both and the act
that changes it. Only a declaration the link delivers, signed, changes the
mode: that is how `converge` and `adopt` arrive, so the flip still works and
nothing else can do it. Genesis marks the bundle consumed, with the digest of
what it applied, so `reconcile` holds a node the mesh has spoken to against
what the mesh last said and never the bundle, and refuses the carried bytes
when they are not what genesis applied. A file is refused when it is not what
the mesh last said: a declaration carries no sequence and no issued-at, so the
host cannot tell older from newer, and says so. Both commands print what they
would change — a hold, a removal, an action named as one — before touching
anything, and --dry-run is that list and nothing more.
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.
The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.
What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.
Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
Asked how the mesh would know if somebody edited their hosts file. It would
not. The file was rewritten within five minutes and the outcome said
"updated" -- which is exactly what the mesh changing its own mind looks like.
So the change vanished, nothing anywhere said why, and the obvious thing to do
is edit it again.
The host now records a digest of what it wrote, which is enough to tell the two
apart on the next pass:
the file matches the declaration unchanged
it matches what was last written updated -- the mesh changed its mind
it matches neither corrected -- somebody changed it here
The machine is put back either way, because holding it to what it was told is
the point. What changes is that it says so.
A digest rather than the content: the store is read on every reconcile and sits
beside the state on disk, and keeping every managed file twice would make it
grow with the size of the machine rather than with the number of resources.
Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.
Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.
A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.
And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.
The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.
With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.
Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.
State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.
Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.
Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.
Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.
It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.
It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.
The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.
Three faults found by running inside a raised machine rather than by reasoning:
A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.
Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.
And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.
30 new tests, 55 in total.