25 Commits
Author SHA1 Message Date
jochen 50776b8613 review: a retired configuration that comes back is retired again on its first original, and only wg-quick's own file is ever removed (hq ADR 0119)
Put back by hand, it was found afresh and its copy became the hold's original, so a machine that
kept restoring it kept growing copies and lost which one was first. The first original now stays
the record's, content that differs is kept once beside it, and the note says a rollback means
unassigning the private network. Nothing is removed unless it is <wireguard dir>/<iface>.conf,
not a path the mesh writes, and not a link, which would leave the key-bearing target behind.
2026-09-27 00:55:50 +02:00
jochen b462f461c6 A taken tunnel's found configuration is retired once the take is proven (hq ADR 0119)
Kept on disk it was the take's fallback; once the mesh's interface is up in its place and a peer
has handshaken with it, it is an unmaintained way back onto the network, held for ever. It is now
removed from where its unit reads it, its kept original verified first and left as it is, and the
hold ends. Until proven — no handshake, or wg not answering — it is kept and the report says why.
The retirement is recorded apart from holds, so later applies, an undeclare, and a reassignment
find it retired rather than missing, and nothing writes it back.
2026-09-27 00:47:57 +02:00
jochen 23a4436499 review: a unit whose file the mesh wrote is stopped when undeclared, and what was found survives a failed first apply (hq ADR 0118)
Records written before Found existed left the adoption guard and the converge filter loaded on
undeclare, then deleted their unit files from under them; a unit whose own file the mesh created
is now the mesh's, whatever its record says. Found is kept apart the moment it is read, so a
first apply that enabled and then failed is not read back as the machine's; boot is found the
first time the mesh sets it; a service once stateless, or moved to another unit, is found afresh
(the old unit given back). The unit is read after the reload that loads a file written in the
same apply, and removal reports what it actually did.
2026-09-27 00:41:02 +02:00
jochen 3112c881e4 Undeclaring gives a unit back the state it was found in, and removes a process the mesh made (hq ADR 0118, issue 130)
A service undeclared used to be stopped: unassigning the private network stopped the container
runtime, unassigning sshd would stop ssh, an uplink module would take the machine offline. The
host now records the unit's state when it first applies it and restores that on undeclare —
found running stays running; started by the mesh (the converge filter) is stopped again; nothing
is started on the way out; a pre-existing record leaves the unit alone.

An undeclared process had no removal at all and failed every apply on its node; its unit, timer
and bundle are now removed.
2026-09-27 00:21:25 +02:00
jochen 06aaac0820 A service may omit its state, so unassigning an uplink module never stops the machine's network manager (hq ADR 0117) 2026-09-27 00:11:58 +02:00
jochen fdc768c476 review: rebuild a file the mesh once wrote whole, keep links, give back a missing line end, and release a hold only after the write (hq issue 128) 2026-09-27 00:09:34 +02:00
jochen 1cb895346d Write into a marked block of a text file instead of over it, so a shared hosts file keeps every line that is not the mesh's (hq issue 128) 2026-09-26 23:51:36 +02:00
jschoubben 982b84310e Look at what a container mounts directly, accept a pre-upgrade label, and write the genesis secret without a newline
Review of the first cut found four things.

A directory mounted into a container is no longer looked inside, not even
for the files this host wrote there. The controller records every
provider's received and contributions file as a plain file under a mounted
directory, so folding those in would have recreated the route proxy — which
re-reads its routes live, by design — on every route change, and killed
every provisioner sidecar, which polls what it receives, mid-reconcile on
every grant. Whether a service reads a file under its directory once or
watches it is the service's; restart-on is how a module says "once", and it
stays the opt-in. Env-files and files mounted directly remain by content.

Genesis wrote the superuser secret as `value\n`; `secret accept` strips the
line ending by design, so the postgres module declared `value` — and with
a mounted file's content in the spec, phase three would have recreated the
store it meant to adopt in place, with the temporary control plane
connected to it. Genesis now writes the value alone. readCredentialFile
tolerated both endings already. Pinned with the bytes the genesis code
path writes, then the module's declaration of the same container: it must
reconcile.

A container carrying a label from before the host folded in what it reads
is accepted rather than recreated, when that label matches the spec as it
used to be computed: what it reads is recorded then, a change is caught
from that record from the next apply on, and the label is renewed at the
next genuine recreate. Recreating them all would have been a restart storm
across the mesh in declaration order, the store first. The trade-off is
stated in the code: a container already stale at upgrade time is not
caught, and could not have been either way.

The record of what a container read is looked up by its name when its
declared id has none — the bundle's `store` becomes `postgres.server` for
the same container — so a change on the day it is adopted still names the
file. The by-target lookup takes the most recently applied record, since
the bundle's record for the same target is never removed by the mesh's.

novox/hq 04-ISSUES/103
2026-09-23 23:40:27 +02:00
jschoubben c60228e719 Recreate a container when the content of a file it reads at creation changes
The host decided whether a container was still the one declared by a digest
of its declaration, and the declaration names an env-file's path and a
mount's path — never what is in them. So when the store was given a new
port, the host rewrote the forge's and the analytics service's environment
files, correctly, and left both containers running with the old port in
their environment: a container reads its env-file when it is CREATED, and
`docker restart` hands it the same environment again. Both looked healthy
until they answered 502.

What a running container takes in at creation is now part of its spec, by
content: every env-file, a file bind-mounted into it, and every file this
host wrote at or under a directory bind-mounted into it — the secrets,
bindings and configs under a module's state directories. The digest is the
one the store already records for a file the host wrote (`wrote`), read
from the state as it stands when the container is reached, so a file
rewritten earlier in the same apply is already the new one; a file the host
has no record of — an env-file a predecessor left, the superuser secret
genesis writes before any declaration names it — is read from disk, which
is what keeps adopting a running store in place a reconcile and not a
recreate.

Deliberately not part of it: what else is in a bind-mounted directory,
which is the service's own data and changes while it runs; a named volume;
a seed created once, which digests as the seed the host wrote and not as
what has grown in it; and a step — a run-once or scheduled container reads
its files when it runs and runs fresh each time. On an adopted node a held
container is held before any of this is looked at.

The host records what each container was created reading, per file, so
the recreate can say which file changed — "recreated: <file> changed" in
the report and, now with its detail, in the log. A container made before
this record existed is recreated once and says so.

novox/hq 04-ISSUES/103
2026-09-23 23:40:27 +02:00
jschoubben f08a8ea3f7 Refuse every file once the mesh has spoken, plan the cutover as one, and let the kept declaration repair the mode
Review of the fix for hq issue 104 found three faults in it. A file applied
on an enrolled node — the mesh's own last declaration included — is applied
as the bundle is, so its resources are recorded as the machine's own and
what the mesh declared reads as undeclared: the plan removed the foundation.
`apply FILE` is for a machine the mesh has not spoken to, and is now refused
saying so whenever declared.json exists. The plan looked at what is held
before what the declaration says is taken, so the one cutover ADR 0100 says
must be previewed read as a hold; it now decides in holdOnAdopted's order,
models a step run inside a held container, and a test holds the plan's
sequence to the apply's outcomes. Genesis wrote the mode on every run, so a
re-run after `converge` left the state saying adopted while the kept,
signed declaration said converged, and the reconcile loop refused every five
minutes with no delivery coming to end it: genesis now writes the mode only
when none is recorded, and where the state and the verified kept declaration
disagree, the kept declaration wins and the repair is said.

Also: a file lock beside the state, taken by the link service, the host's
own commands and the installer alike, so a `reconcile` run by hand no
longer races the loop's save — chosen over refusing while a named service is
active, which would miss a `mesh-host run` started by hand; `--json
--dry-run` emits {plan} like an apply emits {plan, report}; the README's
duplicate flag line; and the bundle refusal is about the digest, not a claim
the carried bytes can never match what genesis applied.
2026-09-23 23:35:49 +02:00
jschoubben 27c4b765b2 Refuse a declaration for the other mode, or older than the mesh's last, and say what an apply would change first
An operator ran `mesh-host reconcile` on an adopted control-node with twelve
modules assigned. It applied the bundle the host carries — the genesis
declaration, foundation only, converged: recreated the store, failed on the
broker's held port, wrote the converged base filter and started its service,
and stopped at the first failing action. The filter closed the machine for
forty-five minutes. The host reported the node adopted in every report, the
declaration said converged, and nothing compared the two; nothing was printed
before acting (hq issue 104).

The host now records the node's mode — from every declaration the mesh sends,
and at genesis from what the operator said — and refuses, at the point of
application, a declaration that says the other mode, naming both and the act
that changes it. Only a declaration the link delivers, signed, changes the
mode: that is how `converge` and `adopt` arrive, so the flip still works and
nothing else can do it. Genesis marks the bundle consumed, with the digest of
what it applied, so `reconcile` holds a node the mesh has spoken to against
what the mesh last said and never the bundle, and refuses the carried bytes
when they are not what genesis applied. A file is refused when it is not what
the mesh last said: a declaration carries no sequence and no issued-at, so the
host cannot tell older from newer, and says so. Both commands print what they
would change — a hold, a removal, an action named as one — before touching
anything, and --dry-run is that list and nothing more.
2026-09-23 23:15:28 +02:00
jschoubben a4e4632077 gofmt the store's firewall record 2026-09-22 18:33:13 +02:00
jschoubben 444ad8f3cf Record the forward policies before disabling ufw, so a retried retirement restores them (hq ADR 0100) 2026-09-22 18:33:09 +02:00
jschoubben facf6af46a Add the mesh's members to a list found in a file written into, and take back only those (hq ADR 0102) 2026-09-22 18:27:51 +02:00
jschoubben 824cb60cbb Hold a found directory, a found service's unit, a container that would mount found data, and a step run in a held container on an adopted node (hq ADR 0103) 2026-09-22 18:14:37 +02:00
jschoubben 0f126d137c Write into a file the machine shares instead of over it, and reload a service that re-reads its configuration instead of restarting it (hq ADR 0102) 2026-09-22 17:48:57 +02:00
jschoubben 3c90d155b3 Converge openings through the firewall an adopted node was found with, and retire it only when the node converges (hq ADR 0100) 2026-09-22 17:19:49 +02:00
jschoubben 3a613113be Keep what an adopted node was found holding until its module is taken, and report it held (hq ADR 0100) 2026-09-22 17:14:04 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben b91342a6bd A machine says which ports it already holds
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.

The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.

What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.

Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
2026-09-01 18:29:39 +02:00
jschoubben 827ce481f2 Somebody editing a managed file is now visible instead of mysterious
Asked how the mesh would know if somebody edited their hosts file. It would
not. The file was rewritten within five minutes and the outcome said
"updated" -- which is exactly what the mesh changing its own mind looks like.
So the change vanished, nothing anywhere said why, and the obvious thing to do
is edit it again.

The host now records a digest of what it wrote, which is enough to tell the two
apart on the next pass:

  the file matches the declaration          unchanged
  it matches what was last written          updated -- the mesh changed its mind
  it matches neither                        corrected -- somebody changed it here

The machine is put back either way, because holding it to what it was told is
the point. What changes is that it says so.

A digest rather than the content: the store is read on every reconcile and sits
beside the state on disk, and keeping every managed file twice would make it
grow with the size of the machine rather than with the number of resources.
2026-08-29 22:55:26 +02:00
jschoubben 980e12a850 A node that loses its mesh comes back on its own
Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.

Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.

A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.

And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.

The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.

With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.

Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
2026-08-29 20:17:49 +02:00
jschoubben fa48b5825e The bundle and the mesh stop removing each other
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.

State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.

Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.

Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
2026-08-29 16:43:46 +02:00
jschoubben ee2648188d Repoint ADR references after HQ consolidated 65 records to 23
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.

Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
2026-08-28 23:33:44 +02:00
jschoubben 9d8239afe8 Stage 2 — the host applies a declaration
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.

It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.

It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.

The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.

Three faults found by running inside a raised machine rather than by reasoning:

A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.

Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.

And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.

30 new tests, 55 in total.
2026-08-26 02:14:25 +02:00