Commit Graph
7 Commits
Author SHA1 Message Date
jschoubben 1bc97ed50d A service can be declared to reflect a file
Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.

That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.

Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.

Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.

Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
2026-08-29 18:04:16 +02:00
jschoubben a4445f5c0a A machine joins the mesh it raised
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.

Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.

The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.

An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.

Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.

The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.

Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.
2026-08-29 16:03:15 +02:00
jschoubben ee2648188d Repoint ADR references after HQ consolidated 65 records to 23
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.

Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
2026-08-28 23:33:44 +02:00
jschoubben f04294c3c1 A service can be enabled at boot, and a container uses the runtime the machine has
Two gaps found by testing podman rather than reasoning about it.

The service shape could not say "starts at boot". It ran `systemctl start`, so
`service: docker.service, running` started docker now and it would not come
back after a reboot unless something else had enabled it. A declaration that
reports success and stops being true at the next power cut.

`boot: enabled|disabled` is now a separate field, not a fourth value of
`state`, because the two are orthogonal: a unit can be enabled and stopped (it
returns at boot) or disabled and running (started by hand, gone after one).
Absent means the host asserts nothing, so a machine whose operator enabled
something is not silently disabled by a declaration that never mentioned it.

Boot state is made true BEFORE the unit is started. When an apply fails part
way, enabled-and-stopped comes back at the next boot and running-and-disabled
does not, so the more durable half goes first.

`is-enabled` has the same trap as `is-active` had. Its exit code is non-zero
for nearly everything, and `static` is neither enabled nor disabled -- the unit
has no install section and CANNOT be enabled. Reading it as "disabled" would
have the host try, fail, and blame the wrong thing, which is the same shape as
reading a missing unit as "stopped".

The container applier no longer calls `docker` literally. Verified on this
machine against podman 6.1.0:

  docker info --format '{{.ServerVersion}}'    -> 29.7.2
  podman info --format '{{.ServerVersion}}'    -> Error: can't evaluate field
                                                  ServerVersion
  podman info --format '{{.Version.Version}}'  -> 6.1.0

So one probe cannot find both, and a host using docker's would report a machine
running podman as having no container runtime at all. Everything else IS
compatible -- run, rm -f, and docker's own Go template syntax for reading state
and labels all work unchanged on podman, confirmed by running them. That is why
this is a two-entry lookup rather than an interface: only the probe differs.

Detected rather than declared, because adoption keeps what the machine already
has (research 012), which hardcoding one runtime contradicts.

A machine with neither now says so, naming both: "docker: command not found" on
a machine deliberately running podman sends the reader after the wrong thing.

Verified end to end against real docker (container created, running, labelled)
and against an empty PATH (refused, naming both runtimes).

Two injections per behaviour, all confirmed to bite. One injection produced a
build failure that my check read as "no bite" for the third time, so the check
now distinguishes them.
2026-08-27 23:58:44 +02:00
jschoubben 9a9937b7e6 A struct per resource kind, instead of one struct with every field
Jochen asked why we don't simply have dedicated structs. We should, and the
flat struct was me extending an existing pattern rather than questioning it.

Before: one Resource struct carrying path, content, mode, unit, state, package,
image, name, env, ports, volumes, args, command, verify and in. Because a file
and a container shared it, nothing stopped {"type":"file","image":"postgres"},
so a `uses` map listed which fields each kind was allowed to carry -- a second
place to keep current, and the kind nobody updates is the one that silently
accepts a field the host will never read.

Now: Directory, File, Service, Package, Container and Action are separate
structs behind a Resource interface. File has no Image field, so the mistake is
not detected -- it is unrepresentable. Adding a field to a kind is the whole of
adding it; there is nowhere else that has to agree.

Parsing is two passes: read the envelope and each resource's raw bytes, peek at
"type" to choose the struct, then decode into it. Peeking is lenient on purpose
-- reading strictly there would report an unknown field before knowing which
fields are known.

Unknown fields are found by comparing the JSON keys against the struct's own
json tags rather than by catching the decoder's error. The decoder stops at the
first unknown field, and RefusalError promises every problem at once: a caller
fixing one field at a time learns the next only by running again. Caught by
testing the refactor against a real declaration -- a container carrying both
`unit` and `mode` reported only one of them.

apply.go switches on the concrete type instead of a string, so a new kind that
has no applier is a compile error rather than a runtime default branch.

No behaviour change otherwise. All existing tests pass unmodified except two
that reached for fields the interface no longer exposes.
2026-08-27 21:03:59 +02:00
jschoubben 337126603e Complete the host's vocabulary: package, container, action
The three shapes the substrate bootstrap needs and the host did not have. Until
now tier 1 could not be raised at all -- step 0 is a package, step 1 a
container, steps 2 and 3 actions -- so every line of the tier 1 and 2 designs
was unbuildable.

package -- present, never upgraded, never uninstalled. Removal is "forgotten",
not "removed": the host cannot know what else needs the package, uninstalling a
container runtime because a declaration changed would stop every container on
the node, and the machine may have had it before the mesh saw it. Reporting it
removed would claim an effect the host declined to have.

container -- identified by a label carrying a digest of the declaration that
made it. Comparing every field the runtime reports cannot be done reliably: a
runtime normalises, defaults and reorders what it is given, and that is
indistinguishable from real drift. There is no in-place update; a container's
configuration is fixed at creation, so any change is a replacement, and saying
so beats a partial update that leaves the running thing half-declared. This is
the one shape the host removes, because it is the one the host created.

action -- bundle-only, per ADR 0047. Verify is mandatory and does double duty:
it is the idempotency check as well as the read-back. The host does not know
what a database is, so "is it already there" is a question only the declaration
can ask. `in` runs the action inside a named container, which steps 2 and 3
need.

Parse now refuses actions; ParseTrusted permits them. The safe path is the
default and the permissive one has to be named. The bundle and a local file
handed to a root process use ParseTrusted; the link will use Parse.

Also replaced the per-type "fields this type ignores" check with a field-set
diff stated as what each type USES. The negative form needs every type revisited
whenever a field is added, and the one nobody revisits silently accepts a field
it will never read.

Images must be pinned by digest (ADR 0046). A bundle naming a tag pins nothing.

Verified against a real machine, not only fakes: an action ran and was
idempotent on the second apply; an action that exits zero and satisfies nothing
fails the apply; a real container was created, labelled, replaced when its
declaration changed, exec'd into, and removed; a real package query round-
tripped. Each new test was also confirmed to fail on an injected fault -- five
injections, each breaking exactly its own test.

One existing test changed: a vanished unit is now reported "forgotten" rather
than "removed", which is what actually happened.
2026-08-27 20:36:27 +02:00
jschoubben 9d8239afe8 Stage 2 — the host applies a declaration
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.

It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.

It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.

The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.

Three faults found by running inside a raised machine rather than by reasoning:

A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.

Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.

And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.

30 new tests, 55 in total.
2026-08-26 02:14:25 +02:00