Commit Graph
15 Commits
Author SHA1 Message Date
jschoubben c0d94e04f3 Recreate a container when the content of a file it reads at creation changes
The host decided whether a container was still the one declared by a digest
of its declaration, and the declaration names an env-file's path and a
mount's path — never what is in them. So when the store was given a new
port, the host rewrote the forge's and the analytics service's environment
files, correctly, and left both containers running with the old port in
their environment: a container reads its env-file when it is CREATED, and
`docker restart` hands it the same environment again. Both looked healthy
until they answered 502.

What a running container takes in at creation is now part of its spec, by
content: every env-file, a file bind-mounted into it, and every file this
host wrote at or under a directory bind-mounted into it — the secrets,
bindings and configs under a module's state directories. The digest is the
one the store already records for a file the host wrote (`wrote`), read
from the state as it stands when the container is reached, so a file
rewritten earlier in the same apply is already the new one; a file the host
has no record of — an env-file a predecessor left, the superuser secret
genesis writes before any declaration names it — is read from disk, which
is what keeps adopting a running store in place a reconcile and not a
recreate.

Deliberately not part of it: what else is in a bind-mounted directory,
which is the service's own data and changes while it runs; a named volume;
a seed created once, which digests as the seed the host wrote and not as
what has grown in it; and a step — a run-once or scheduled container reads
its files when it runs and runs fresh each time. On an adopted node a held
container is held before any of this is looked at.

The host records what each container was created reading, per file, so
the recreate can say which file changed — "recreated: <file> changed" in
the report and, now with its detail, in the log. A container made before
this record existed is recreated once and says so.

novox/hq 04-ISSUES/103
2026-09-23 23:12:52 +02:00
jschoubben a4e4632077 gofmt the store's firewall record 2026-09-22 18:33:13 +02:00
jschoubben 444ad8f3cf Record the forward policies before disabling ufw, so a retried retirement restores them (hq ADR 0100) 2026-09-22 18:33:09 +02:00
jschoubben facf6af46a Add the mesh's members to a list found in a file written into, and take back only those (hq ADR 0102) 2026-09-22 18:27:51 +02:00
jschoubben 824cb60cbb Hold a found directory, a found service's unit, a container that would mount found data, and a step run in a held container on an adopted node (hq ADR 0103) 2026-09-22 18:14:37 +02:00
jschoubben 0f126d137c Write into a file the machine shares instead of over it, and reload a service that re-reads its configuration instead of restarting it (hq ADR 0102) 2026-09-22 17:48:57 +02:00
jschoubben 3c90d155b3 Converge openings through the firewall an adopted node was found with, and retire it only when the node converges (hq ADR 0100) 2026-09-22 17:19:49 +02:00
jschoubben 3a613113be Keep what an adopted node was found holding until its module is taken, and report it held (hq ADR 0100) 2026-09-22 17:14:04 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben b91342a6bd A machine says which ports it already holds
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.

The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.

What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.

Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
2026-09-01 18:29:39 +02:00
jschoubben 827ce481f2 Somebody editing a managed file is now visible instead of mysterious
Asked how the mesh would know if somebody edited their hosts file. It would
not. The file was rewritten within five minutes and the outcome said
"updated" -- which is exactly what the mesh changing its own mind looks like.
So the change vanished, nothing anywhere said why, and the obvious thing to do
is edit it again.

The host now records a digest of what it wrote, which is enough to tell the two
apart on the next pass:

  the file matches the declaration          unchanged
  it matches what was last written          updated -- the mesh changed its mind
  it matches neither                        corrected -- somebody changed it here

The machine is put back either way, because holding it to what it was told is
the point. What changes is that it says so.

A digest rather than the content: the store is read on every reconcile and sits
beside the state on disk, and keeping every managed file twice would make it
grow with the size of the machine rather than with the number of resources.
2026-08-29 22:55:26 +02:00
jschoubben 980e12a850 A node that loses its mesh comes back on its own
Disconnection is an ordinary situation and not a failure, and until now the
host treated it as the end: the link dropped and the process returned. A laptop
shut for a week would have come back needing somebody to start it again.

Now it reconnects, with a backoff that starts at two seconds and slows to two
minutes. The two common reasons differ in how long they last -- a broker
restarting is back in seconds, a machine that has moved to a network with no
route may be hours -- so it starts fast and slows down, and resets once a
connection has actually held for thirty seconds. Without that reset, a node
that reconnects and immediately drops climbs to the maximum and stays there
long after the cause is gone.

A wrong certificate is said in full every time rather than folded into a retry
count. That does not mean the network is down; it means what answered is not
the mesh this node joined, and no waiting fixes it.

And it says when it gets back in. It logged every failure and nothing on
success, so a log full of "trying again" followed by silence read as still
broken when it meant the opposite.

The other half: a node now keeps what it was told, not only what it applied.
The record of what was applied holds an id, a type and a target -- what removal
needs, not what creation needs -- so it could not be re-applied. The
declaration is kept whole, signed, and verified again every time it is read
back, so the file on disk is trusted for the same reason the message was rather
than for being local. A tampered one is refused, and so is one signed by
another mesh.

With both, the host reconciles against what it was last told every five
minutes, connected or not. That is not polling for changes -- changes are
pushed -- it is the answer to a machine drifting: a file edited by hand, a
container somebody stopped, a service that died.

Verified in the lab. The broker was stopped: the node retried at 2s, 4s, 8s,
saying why each time, and kept its overlay up throughout. The broker came back
and the node rejoined without being touched. A declaration published while a
node was away was waiting on the broker and applied the moment it connected,
which is the buffer ADR 0006 describes doing its job.
2026-08-29 20:17:49 +02:00
jschoubben fa48b5825e The bundle and the mesh stop removing each other
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.

State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.

Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.

Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
2026-08-29 16:43:46 +02:00
jschoubben ee2648188d Repoint ADR references after HQ consolidated 65 records to 23
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.

Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
2026-08-28 23:33:44 +02:00
jschoubben 9d8239afe8 Stage 2 — the host applies a declaration
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.

It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.

It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.

The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.

Three faults found by running inside a raised machine rather than by reasoning:

A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.

Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.

And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.

30 new tests, 55 in total.
2026-08-26 02:14:25 +02:00