Commit Graph
16 Commits
Author SHA1 Message Date
jschoubben 982b84310e Look at what a container mounts directly, accept a pre-upgrade label, and write the genesis secret without a newline
Review of the first cut found four things.

A directory mounted into a container is no longer looked inside, not even
for the files this host wrote there. The controller records every
provider's received and contributions file as a plain file under a mounted
directory, so folding those in would have recreated the route proxy — which
re-reads its routes live, by design — on every route change, and killed
every provisioner sidecar, which polls what it receives, mid-reconcile on
every grant. Whether a service reads a file under its directory once or
watches it is the service's; restart-on is how a module says "once", and it
stays the opt-in. Env-files and files mounted directly remain by content.

Genesis wrote the superuser secret as `value\n`; `secret accept` strips the
line ending by design, so the postgres module declared `value` — and with
a mounted file's content in the spec, phase three would have recreated the
store it meant to adopt in place, with the temporary control plane
connected to it. Genesis now writes the value alone. readCredentialFile
tolerated both endings already. Pinned with the bytes the genesis code
path writes, then the module's declaration of the same container: it must
reconcile.

A container carrying a label from before the host folded in what it reads
is accepted rather than recreated, when that label matches the spec as it
used to be computed: what it reads is recorded then, a change is caught
from that record from the next apply on, and the label is renewed at the
next genuine recreate. Recreating them all would have been a restart storm
across the mesh in declaration order, the store first. The trade-off is
stated in the code: a container already stale at upgrade time is not
caught, and could not have been either way.

The record of what a container read is looked up by its name when its
declared id has none — the bundle's `store` becomes `postgres.server` for
the same container — so a change on the day it is adopted still names the
file. The by-target lookup takes the most recently applied record, since
the bundle's record for the same target is never removed by the mesh's.

novox/hq 04-ISSUES/103
2026-09-23 23:40:27 +02:00
jschoubben c60228e719 Recreate a container when the content of a file it reads at creation changes
The host decided whether a container was still the one declared by a digest
of its declaration, and the declaration names an env-file's path and a
mount's path — never what is in them. So when the store was given a new
port, the host rewrote the forge's and the analytics service's environment
files, correctly, and left both containers running with the old port in
their environment: a container reads its env-file when it is CREATED, and
`docker restart` hands it the same environment again. Both looked healthy
until they answered 502.

What a running container takes in at creation is now part of its spec, by
content: every env-file, a file bind-mounted into it, and every file this
host wrote at or under a directory bind-mounted into it — the secrets,
bindings and configs under a module's state directories. The digest is the
one the store already records for a file the host wrote (`wrote`), read
from the state as it stands when the container is reached, so a file
rewritten earlier in the same apply is already the new one; a file the host
has no record of — an env-file a predecessor left, the superuser secret
genesis writes before any declaration names it — is read from disk, which
is what keeps adopting a running store in place a reconcile and not a
recreate.

Deliberately not part of it: what else is in a bind-mounted directory,
which is the service's own data and changes while it runs; a named volume;
a seed created once, which digests as the seed the host wrote and not as
what has grown in it; and a step — a run-once or scheduled container reads
its files when it runs and runs fresh each time. On an adopted node a held
container is held before any of this is looked at.

The host records what each container was created reading, per file, so
the recreate can say which file changed — "recreated: <file> changed" in
the report and, now with its detail, in the log. A container made before
this record existed is recreated once and says so.

novox/hq 04-ISSUES/103
2026-09-23 23:40:27 +02:00
jschoubben 27c4b765b2 Refuse a declaration for the other mode, or older than the mesh's last, and say what an apply would change first
An operator ran `mesh-host reconcile` on an adopted control-node with twelve
modules assigned. It applied the bundle the host carries — the genesis
declaration, foundation only, converged: recreated the store, failed on the
broker's held port, wrote the converged base filter and started its service,
and stopped at the first failing action. The filter closed the machine for
forty-five minutes. The host reported the node adopted in every report, the
declaration said converged, and nothing compared the two; nothing was printed
before acting (hq issue 104).

The host now records the node's mode — from every declaration the mesh sends,
and at genesis from what the operator said — and refuses, at the point of
application, a declaration that says the other mode, naming both and the act
that changes it. Only a declaration the link delivers, signed, changes the
mode: that is how `converge` and `adopt` arrive, so the flip still works and
nothing else can do it. Genesis marks the bundle consumed, with the digest of
what it applied, so `reconcile` holds a node the mesh has spoken to against
what the mesh last said and never the bundle, and refuses the carried bytes
when they are not what genesis applied. A file is refused when it is not what
the mesh last said: a declaration carries no sequence and no issued-at, so the
host cannot tell older from newer, and says so. Both commands print what they
would change — a hold, a removal, an action named as one — before touching
anything, and --dry-run is that list and nothing more.
2026-09-23 23:15:28 +02:00
jschoubben a4e4632077 gofmt the store's firewall record 2026-09-22 18:33:13 +02:00
jschoubben 444ad8f3cf Record the forward policies before disabling ufw, so a retried retirement restores them (hq ADR 0100) 2026-09-22 18:33:09 +02:00
jschoubben facf6af46a Add the mesh's members to a list found in a file written into, and take back only those (hq ADR 0102) 2026-09-22 18:27:51 +02:00
jschoubben 824cb60cbb Hold a found directory, a found service's unit, a container that would mount found data, and a step run in a held container on an adopted node (hq ADR 0103) 2026-09-22 18:14:37 +02:00
jschoubben 0f126d137c Write into a file the machine shares instead of over it, and reload a service that re-reads its configuration instead of restarting it (hq ADR 0102) 2026-09-22 17:48:57 +02:00
jschoubben 3c90d155b3 Converge openings through the firewall an adopted node was found with, and retire it only when the node converges (hq ADR 0100) 2026-09-22 17:19:49 +02:00
jschoubben 3a613113be Keep what an adopted node was found holding until its module is taken, and report it held (hq ADR 0100) 2026-09-22 17:14:04 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben b91342a6bd A machine says which ports it already holds
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.

The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.

What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.

Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
2026-09-01 18:29:39 +02:00
jschoubben 827ce481f2 Somebody editing a managed file is now visible instead of mysterious
Asked how the mesh would know if somebody edited their hosts file. It would
not. The file was rewritten within five minutes and the outcome said
"updated" -- which is exactly what the mesh changing its own mind looks like.
So the change vanished, nothing anywhere said why, and the obvious thing to do
is edit it again.

The host now records a digest of what it wrote, which is enough to tell the two
apart on the next pass:

  the file matches the declaration          unchanged
  it matches what was last written          updated -- the mesh changed its mind
  it matches neither                        corrected -- somebody changed it here

The machine is put back either way, because holding it to what it was told is
the point. What changes is that it says so.

A digest rather than the content: the store is read on every reconcile and sits
beside the state on disk, and keeping every managed file twice would make it
grow with the size of the machine rather than with the number of resources.
2026-08-29 22:55:26 +02:00
jschoubben fa48b5825e The bundle and the mesh stop removing each other
04-ISSUES/010. The store now records where each resource came from -- carried,
or declared -- and each origin removes only its own. A declaration removes what
the mesh previously declared and never what the bundle raised.

State written before the field existed reads as carried, because everything a
host had applied by then came from its bundle: there was no other way to tell
it anything. Guessing the other way would have the first upgrade remove the
substrate, which is this fault arriving through the change that fixes it.

Verified on the scenario that caused it, and on the property that had to
survive it: a later declaration dropping a resource still removes that
resource, so removal by omission still means what it meant.

Also stops swallowing a publish failure. A node that applied a declaration and
could not tell the mesh looked exactly like one that had -- the mesh believing
it never answered, the node believing it did, and nothing anywhere saying so.
Reports are published mandatory now, so anything the broker cannot route comes
back and is said out loud rather than dropped in silence.
2026-08-29 16:43:46 +02:00
jschoubben ee2648188d Repoint ADR references after HQ consolidated 65 records to 23
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.

Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
2026-08-28 23:33:44 +02:00
jschoubben 9d8239afe8 Stage 2 — the host applies a declaration
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.

It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.

It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.

The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.

Three faults found by running inside a raised machine rather than by reasoning:

A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.

Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.

And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.

30 new tests, 55 in total.
2026-08-26 02:14:25 +02:00