Commit Graph
32 Commits
Author SHA1 Message Date
jschoubben 7283924a35 Take over the found tunnel: its key, its port, its peers; stop it, never flush
On an adopted machine the private network takes the predecessor's tunnel
over in place (hq ADR 0105). Genesis finds the one interface up besides the
mesh's own, settles the hub's port and the mesh's range on it, and skips
ADR 0100's non-overlap check for a range that is now the tunnel's; a
--hub-port or --overlay-range that disagrees is refused naming the tunnel's.

At enrolment the found interface's private key becomes this node's overlay
key — the one credential the mesh takes rather than mints — stored where a
generated one is stored, never printed and never sent; the tunnel (port,
address, range, peers) travels with the keys so the mesh composes from it
before the first declaration.

The interface's service may say what it takes over. Before the mesh's unit
starts, the found configuration is kept like any held file and the found
unit is stopped and disabled; nothing is flushed, and an interface still up
after its unit stopped refuses the takeover rather than half-working. The
report says what was carried: interface, port, range, peer count, taken or
not, and where the original was kept.
2026-09-23 23:26:35 +02:00
jschoubben 824cb60cbb Hold a found directory, a found service's unit, a container that would mount found data, and a step run in a held container on an adopted node (hq ADR 0103) 2026-09-22 18:14:37 +02:00
jschoubben 0f126d137c Write into a file the machine shares instead of over it, and reload a service that re-reads its configuration instead of restarting it (hq ADR 0102) 2026-09-22 17:48:57 +02:00
jschoubben 3c90d155b3 Converge openings through the firewall an adopted node was found with, and retire it only when the node converges (hq ADR 0100) 2026-09-22 17:19:49 +02:00
jschoubben fcc447c216 Read a node's adoption from every declaration, so the host knows which modules are untaken (hq ADR 0100) 2026-09-22 17:11:26 +02:00
jschoubben 8da78de210 A run-once step names what it reads and runs again when it changed (novox/hq ADR 0099, issue 077) 2026-09-21 23:25:06 +02:00
jschoubben b72b71a989 The host applies the newest declaration, a file may be created once, the foundation filters first
031: a window of unacknowledged declarations is drained to the newest; the
rest are set aside and reported as superseded. 035: a file resource may say
create-once — written when absent, kept untouched when present (ADR 0087).
054: the bundle installs nftables and loads a base ruleset before the store
and broker, in the table the filter module later replaces (ADR 0088).
2026-09-21 12:11:52 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben de5160de4a A unit file reinterprets an environment value; a container does not
Found by being asked whether processes and containers handle environment the
same way. They do not, and the difference is not cosmetic.

Docker passes --env through literally. A unit file reads three things out of a
value that nothing else does, and a module's environment routinely contains all
three because a generated password is arbitrary bytes:

  - % begins a specifier. %H is the hostname. A password containing one is
    silently replaced, and it fails later as an authentication error nobody can
    explain by reading the declaration.
  - whitespace separates assignments. Unquoted, K=a b sets K to "a" and reads
    "b" as another assignment.
  - a newline ends the line, and what follows is read as a unit DIRECTIVE.

The first two are escaped: quoted, with quotes and backslashes escaped and
percent doubled. The third cannot be — a unit's environment has no way to carry
a line break — so it is refused in validation, near whoever wrote it. Without
that, an environment value could write ExecStart= and have the machine run
something nobody declared.

Ordinary awkward values stay accepted, because refusing those too would leave a
module unable to hold a generated password.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 12:40:36 +02:00
jschoubben f5cf9510c1 One kind for the module's own code, with three modes
The first cut of this added a `daemon` for the long-running case alone. That
would have meant a new vocabulary entry for each of the others — a scheduled
task, a run-once migration, a health check — when they are one thing run at
different cadences. That is a field, not four entries in a vocabulary where every
entry widens what a compromised control plane can express.

So it mirrors a container exactly, because it IS a container's twin: the same
intent, hosted by the machine's own supervisor instead of a runtime. Stays up,
runs once, or runs on a schedule.

Tools, hooks and event consumers are not further modes. They are loaded by a tool
host, which is itself a process that stays up — so the generic case already
covers them, which is the test of whether it is generic.

A scheduled process gets a timer and a unit that finishes; a long-running one
gets a unit that is restarted when it exits. Getting that wrong either way is a
second copy running continuously between fires, or a schedule that never fires.
The modes are exclusive and validation says so near the author: something that
runs once does not run on a schedule, and something not running between fires
cannot be restarted when a file changes.

A missed fire happens when the machine comes back rather than being skipped,
which is the difference between a machine that was down and a schedule that
quietly stopped.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 10:26:31 +02:00
jschoubben 1f0fb85128 A daemon says what to run, not how it is hosted
The mechanism was leaking into every module. Code of one's own meant a container
and therefore an image; a script meant a service and a unit somebody else had to
install. One intent — run this and keep it running — expressed two unrelated
ways, with the hosting chosen before anything could be declared.

A daemon names a bundle and a command. The host fetches it, refuses it unless it
hashes to what was declared, unpacks it where the mesh keeps such things, writes
the unit and puts it in the state asked for. The unit is the mesh's, generated
whole and saying so, because an edit that survives until the next declaration and
then vanishes is worse than one that is refused.

Its identity is the bytes AND how it is run: two daemons from one bundle
differing only in their command are different daemons, and tracking the digest
alone would call the second unchanged and leave the first running. The unit is
rendered deterministically for the same reason — environment from a map would be
written in Go's iteration order, so every apply would see a different unit and
restart an unchanged daemon for ever.

restart-on is honoured as a service's is: a running process does not re-read its
configuration, so replacing a file and finding the daemon already up leaves the
machine behaving as before while every check passes.

A full-host shape, not a portable one: it needs a process supervisor to install
into. It does NOT need a container runtime, which is the point.

Two guards caught this properly and both were updated deliberately rather than
silenced: the vocabulary count, which exists because every addition widens what a
compromised control plane can express, and the shape test that catches a kind the
language has and a host cannot apply — added after `network` did exactly that.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 02:29:33 +02:00
jschoubben 4af9483219 declaration: an image may be named by the digest of its own configuration
A manifest digest is assigned by a registry on push, so insisting on one meant a
registry had to exist before the thing that lets a mesh have a registry could
start — a dependency the pinning rule created by accident, not a pin. The mesh's
own control plane is built from source and lives in no public registry.

A bare sha256:... names an image the machine already holds, by the digest of its
own configuration: immutable and unforgeable in exactly the way the rule asks
for. Absent, it says so plainly rather than failing at a pull nothing serves.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 22:40:02 +02:00
jschoubben 9d1f001dcc apply: a scheduled step is a container run on a cadence (ADR 0053)
The recurring twin of run-once, one modifier over: a container marked
schedule: "<cron>" is run to completion on its cadence, not started as a
service and not run once as a gate.

The gating rule is deliberately reversed. Installing a schedule records it
as present state and reports the node current at once (applySchedule) --
it never runs the container and does not gate what follows. A Scheduler,
held for the life of the daemon and re-established from each applied
declaration (the declaration is the source of truth, ADR 0018), fires the
container off an injected clock. A run that exits non-zero is logged and
never fails the apply or flips the node's state, because it happens
outside the apply and the store entirely. Runs never stack: a run still
going when the next is due is skipped, not started as a second copy.

No new host shape and no new action -- schedule is a string on the
container the host already has, and the host process runs the container
itself rather than installing a system timer (the rejected option 1). A
minimal five-field cron (declaration/cron.go) validates on arrival and
computes the next due minute; time is injected so the scheduler is tested
without the wall clock.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 14:08:44 +02:00
jschoubben 19e5dd83ea apply: a run-once container is a step the host runs to completion (ADR 0052)
A module can declare state but not a step that runs at first boot. This adds
`run-once: true` to the container shape: the host runs it in the foreground,
requires it to exit 0, and records that it did — as the digest of the
declaration, so a re-apply does not re-run it unless the declaration changed.

Because the declaration is applied in order and a failed run-once step gates the
apply the way a failed action does, whatever is declared after the step starts
only once it has completed. That is how "before the broker starts" is enforced,
with no dependency graph the host must resolve (ADR 0005): the step is declared
first, and the container that needs it is never reached until it is done.

No new host shape and no arbitrary host command — a run-once container is
strictly less powerful than an action. Validation refuses run-once with
restart-on (contradictory lifecycles). Six unit tests; go test ./... green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:55:45 +02:00
jschoubben f06eea5fa3 declaration: an access is mounted, and the host owns nothing about it
The tenth shape (novox/hq ADR 0051). Shared, pre-existing data — a media
library, a download spool several modules use — is the operator's, not
the mesh's. A `directory` resource is the host's own: it creates it,
chowns it, sets its mode and removes it when empty. An access is the
opposite on every axis.

Add the `access` type to the vocabulary. Its applier confirms the path is
present and changes nothing: it does not create, chown, reconcile or set
a mode. Absent is refused clearly — the operator must provide it — rather
than created, because a bind mount whose source is missing is made as
root by the container runtime with the wrong ownership (04-ISSUES/026).
Undeclaring an access forgets the record and never touches the path,
which is the data loss ADR 0030 prevents, on a directory the mesh never
made.

Full hosts speak it (it gates a bind mount, which needs the container
runtime); the vocabulary guard test records the decision that made it the
tenth shape. Unit tests cover present, absent-refused, and
undeclared-left-alone.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:19:33 +02:00
jschoubben aa441bac19 A container reflects its config: restart-on for containers (04-ISSUES/009)
A container reads a mounted file once, at start; its spec (image, env, volumes)
does not include a mounted file's content, so a settings change that re-renders the
file left the running process holding the old value while every check passed. Give
Container the restart-on field a Service already has, and recreate the container
when a named resource changed this pass. Unit-tested (recreated on change, left
alone otherwise) and proven in the mesh-lab: a running grafana runtime picked up a
token change on the next push.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:39:47 +02:00
jschoubben 8c248e3d7f A secret can reach a container's environment, and sit inside a config file
Two gaps found by writing the first real module's manifest rather than
by reasoning about one. Both are fields on existing shapes, so the
vocabulary is still nine.

**env-file on a container.** A declaration reaches a node over the
broker and `env` is plain text in it, so a password there is a password
the broker sees — the transitive trust refused everywhere else. A sealed
file arrives unreadable, the host writes it, the runtime reads it. It is
also simply how third-party software takes credentials: nothing shipping
in a container will read a path the mesh invented, and every one of them
reads its environment.

**secrets in a file's content.** A program wanting its token inside a
JSON document cannot be handed a file that is entirely a token, and the
mesh cannot compose the document because it discarded the value. So the
module supplies the document with `${secret:name}` in it, the mesh
delivers the value sealed, and the host is the only thing that ever
holds both.

Substitution is textual and the host learns no formats. Deliberate: a
mechanism that understood JSON would be asked to understand YAML next,
and then INI, which is how the arrangement this replaces became
something nobody could hold in their head. The module knows its own
format because it wrote the rest of the file. The sharp edge is stated
rather than left to be discovered — a value containing a quote is not
escaped for whatever surrounds it.

Refused in both directions, because both are somebody being wrong about
where a credential is: a placeholder with nothing to fill it would write
`${secret:x}` into a config file, and a secret the content never uses
means somebody believes a credential is in a file where it is not.

A file that carries one is 0600 unless the module said otherwise.
2026-08-31 22:25:57 +02:00
jschoubben 4a43e21794 A network is a shape, so that it can be removed
novox/hq ADR 0029, and work breakdown 1.3. A module of several
containers had no way to let them reach each other by name: a container
declaration could join a network and nothing could create one.

An action was the obvious alternative and is refused on removal —
"an action has no footprint the host can undo", so a network made that
way outlives every module that is ever unassigned, and the mesh cannot
tell. A resource the mesh can create and never clean up is one it should
not create.

A name and nothing else. Not a driver, a subnet or a gateway: each is
something a module would have to know about the machine it lands on, and
a module naming a subnet collides with whatever else chose the same one.

It needs no new ordering rule. Resources apply in declaration order and
orphans are removed in reverse, so a network written before the
containers that join it is created first and removed last — after they
are gone. A runtime refusing to remove one still in use is reported
rather than swallowed, because that means something undeclared is
holding it.

The vocabulary guard fired on the change, as designed, and now names the
record instead of a number: nine shapes, with the argument beside the
count.

Creation reads back rather than trusting an exit status (ADR 0018): a
runtime that reports success and made nothing leaves every container
that joins it failing to start, one step from the cause.
2026-08-31 18:55:06 +02:00
jschoubben c3d6f240fe Give a container the names, rather than a resolver to ask
The commit before this said "told where to resolve names" and passed --dns,
which is not what it ended up doing. This is that correction: a container is
given the names themselves, written into its own hosts file by the runtime.

The reason for the change is the decision the mesh already made about names — a
file rather than a resolver, because it works on every runtime, needs no
package and has no failure mode of its own. Passing a resolver address would
have required a resolver to exist, which at that point none did.

A resolver is coming, for the case a file genuinely cannot express: a service
named under a machine, postgres.novox.internal, where the wildcard cannot be
enumerated in advance. When it arrives it will need this field back under its
own name. It is not being kept in the meantime — a field nothing fills is a
field nobody can trust, and the vocabulary is asserted by a count for exactly
that reason.
2026-08-31 12:05:52 +02:00
jschoubben 0e2b288bb6 A container can be told where to resolve names
A container does not inherit the machine's names. It gets its own /etc/hosts
holding its own hostname, and a runtime rewrites resolv.conf — so every
internal name the mesh wrote for that machine is invisible to what the machine
is running.

That was hit for real, in the lab: a database client on one node could not
resolve another node, on a mesh where both names were correct and present on
both machines. It was worked around by resolving on the host and passing an
address, which is the kind of workaround that should not be needed twice.

A field on an existing shape, not a ninth shape — the vocabulary is still the
eight the count asserts.

Per container rather than by editing the machine's resolver configuration: that
file belongs to something else on most machines, and a host that edited it
would be fighting whatever owns it on every boot — the fault this host exists
to avoid, in the place it would be hardest to see.

A container told nothing is run exactly as before. Most containers should
resolve whatever the machine resolves, and passing an empty flag would be a
change of behaviour dressed up as a default.
2026-08-31 11:09:49 +02:00
jschoubben f9c70a7c97 Something after the declaration is refused whole
A JSON decoder reads one value and stops, so a file holding a declaration and
then anything else parsed as the declaration and the rest was never looked at.
The machine applies something, reports success, and what it applied is not what
the file says — the same fault this host refuses everywhere else, in its
quietest form.

Not hypothetical. A test harness had been appending a line to the substrate
bundle by accident; every apply kept working and nothing said so for as long as
it was wrong. That is how the bug was found, and it is the argument for the
refusal: a file with something after it may be a truncated rewrite or two
declarations run together, and applying the first would be applying something
nobody wrote.

Trailing whitespace is not "something after it".
2026-08-31 00:51:29 +02:00
jschoubben c57087d75d A user, bytes, and an archive — because most of what people install is
not a service

A shell, a terminal, a chat client, a desktop are a package plus
configuration in somebody's home. A mesh with no notion of a user can own
/etc and nothing anybody looks at, which is most of the reason to manage
a machine at all.

Three shapes, and the vocabulary test asserts the count precisely because
widening it widens what a compromised control plane can express:

  user     a login, its shell and its groups
  archive  a set of files, fetched by digest and unpacked
  (file)   gains `bytes` for what is not text, and `owner`

`user` also makes "zsh is my login shell" declared state. chsh is a
command, the link may not carry one, and a shell settable only by hand is
a shell the mesh cannot manage.

Groups are additive and never pruned — usermod without --append REPLACES
them, which would silently remove every group that makes a login able to
use the machine. A machine's own groups are not the mesh's to know about.

The archive is the one place this host reaches out on its own; everywhere
else it holds one outbound connection and fetches nothing. So it carries
the discipline the bootstrap already uses for images: pinned by digest,
and the digest checked before a single file is written.

Two decisions in the unpacker worth naming:

- an entry naming a path outside the archive is REFUSED, not sanitised.
  Rewriting it to land inside would put a file somewhere nobody asked for
  and report success. Found by the test: the first version quietly
  relocated it.
- symlinks and device nodes are refused rather than skipped, or an
  archive that needed one arrives silently incomplete.

A partial host does archives and refuses users: an archive needs a
filesystem and a way to fetch; a user needs a user database it is allowed
to write.
2026-08-30 03:22:38 +02:00
jschoubben bc5b6e2143 One reader for a declaration file, because there were three
Found raising two machines: `apply <file>` refused the bundle example in
this repository with `invalid character '/'`. The bundle strips whole-line
comments; apply handed the raw bytes to the parser. So a file this repo
ships could be built into a binary and not applied from disk.

This is the third instance of one fault. There is already a test here
named "what validates is what is applied", written when `mesh-host
bundle` said yes and `reconcile` said no about the same artefact — two
paths to one thing, disagreeing. Fixing that instance left the shape
intact, so it came back somewhere else.

So the fix is structural rather than local: `declaration.ParseFileTrusted`
is the one way to read a declaration from disk, and the bundle and apply
both use it. Comment handling and its test now live in one place, since
having them in two is how it came to be done in two.

The wire format is untouched — over the link it stays exactly JSON,
because a format with a second thing to strip is a format with a second
thing to disagree about. Asserted, and confirmed to fail if the link
starts stripping.
2026-08-30 02:54:25 +02:00
jschoubben a752fc514b A file the mesh can deliver and cannot read
Everything else in a declaration is visible to whatever carried it. The
message is signed so it cannot be forged, and signing does not make it
unreadable — a password in `content` is a password the broker sees, which
is the transitive trust this design refuses everywhere else.

So a node generates a third key at enrolment and reports the public half,
exactly as it does for its identity and its overlay key. A file may
arrive `sealed` instead of `content`; the host opens it with that key and
writes the result. The control plane can then store a credential it
cannot use, and the broker relays a blob it cannot read.

A third key rather than reusing one of the two. The identity key signs
and is Ed25519; the overlay key is WireGuard's and is tied to being on
the private network, which a machine may not be. A key used for two
purposes is one rotation away from breaking the other.

Details that are not incidental:

- sealed and content together is refused, so "was this the secret or the
  placeholder" is answerable by looking
- a sealed file defaults to 0600 rather than 0644, because the
  consequence differs; an explicit mode still wins
- a node with no sealing key refuses the file rather than skipping it. A
  machine that quietly omits the one resource carrying a credential looks
  configured and cannot connect
- what is recorded is a digest of what was written, so drift on a
  credential is still detected without the node keeping the value, and
  the report that goes back over the broker carries neither

The key is made at enrolment rather than on first use. One made later is
one the mesh was never told about, so nothing could ever be sealed to it,
and the node would look fine and receive nothing.

This is why sealing was borrowed from another mesh's mistakes rather than
its design: there, credentials sit encrypted in the control plane's
database — which guards the database file and nothing else, since the
same value is also in each node's environment file in plain text and
inside every connection string composed from it. Its own tooling has to
search by value rather than by name to find the copies, and says the ones
inside composed URLs are usually the only copies in use.
2026-08-30 00:12:22 +02:00
jschoubben d81f826089 A check that what the control plane emits is what this host accepts
"The host needs no new vocabulary" is the load-bearing claim behind every
computed and contributed resource on the other side, and it had never
been tested against this parser — only asserted.

Skipped unless MESH_EMITTED names a file, so it stays a check somebody
runs deliberately rather than a dependency between two repositories.

  mesh-control plan <node> --json > /tmp/d.json
  MESH_EMITTED=/tmp/d.json go test ./internal/declaration/ -v

Confirmed to fail when the declaration carries an action, which is the
thing this parser exists to refuse.
2026-08-29 23:35:11 +02:00
jschoubben 1bc97ed50d A service can be declared to reflect a file
Because a running service does not re-read its configuration. Replace the file,
find the service running, do nothing -- and the machine keeps behaving as it
did while every check passes, because the file is right and the service is up.

That is not hypothetical. It is how a third node joining a mesh left the first
two carrying a private network that no longer existed, with every part of it
reporting success.

Declared state rather than a command: the declaration says the running service
must reflect these files, and the host works out that it does not. A command to
restart would be an action, and the link may not carry one -- the host refused
precisely that when I tried it, correctly, which is how this shape was arrived
at rather than the other.

Scoped to one apply. A change from an earlier one has already been reflected,
and restarting for it every time would make a steady machine bounce its
services for ever.

Also: the node generates its overlay key at enrolment and reports the public
half, and the store waits three minutes rather than one for the database --
sixty seconds is not enough for a cold machine running initdb, and it failed
that way three times, which is the worst kind of flake because a second run
always fixed it.
2026-08-29 18:04:16 +02:00
jschoubben a4445f5c0a A machine joins the mesh it raised
The last step of the first-node path, and the bundle now carries all of it: a
container runtime, the store, a database per context, their schemas, the broker
with a certificate it generated itself, and the control plane running.

Then the machine enrols against the mesh on its own disk. It dials the broker
over TLS, refuses anything but the pinned certificate, presents the one-time
secret with a public key it generated, and is told the name the mesh has for
it. Its specialness lasted two commands, which is what ADR 0004 asked for.

The identity is saved only after the mesh says it knows this node. A node
holding an identity the mesh never recorded would believe it had joined and be
believed by nobody, which is worse than not joining because nothing looks wrong.

An already-enrolled machine refuses a valid token rather than quietly acquiring
a second identity, and a spent token is refused by the mesh. Both checked.

Containers gained a network field. The control plane must reach the store and
the broker on the machine it was raised on, before there is any mesh to arrange
that; the alternative was publishing ports and guessing an address that works
from inside a container, which fails in a worse way.

The control plane talks to the broker over loopback in plaintext, deliberately.
The TLS on 5671 exists so a node crossing a network can pin a certificate, not
for a hop that never leaves the machine.

Verified on a sealed lab machine: eleven resources applied from bare, the
control plane consuming, a token issued from inside it, and the machine
enrolled -- with the recorded public key matching what the host printed, the
token marked spent, and the profile stored.
2026-08-29 16:03:15 +02:00
jschoubben ee2648188d Repoint ADR references after HQ consolidated 65 records to 23
96 comments across the two repos named records that no longer exist. Each now
points at the consolidated record that holds its reasoning -- ADR 0034 (a test
defends a decision) is 0017, the eight host records are 0005, the four lab
records are 0016.

Worth noting for next time: these are references from outside HQ, so renumbering
there is not free. It cost 38 files here.
2026-08-28 23:33:44 +02:00
jschoubben f04294c3c1 A service can be enabled at boot, and a container uses the runtime the machine has
Two gaps found by testing podman rather than reasoning about it.

The service shape could not say "starts at boot". It ran `systemctl start`, so
`service: docker.service, running` started docker now and it would not come
back after a reboot unless something else had enabled it. A declaration that
reports success and stops being true at the next power cut.

`boot: enabled|disabled` is now a separate field, not a fourth value of
`state`, because the two are orthogonal: a unit can be enabled and stopped (it
returns at boot) or disabled and running (started by hand, gone after one).
Absent means the host asserts nothing, so a machine whose operator enabled
something is not silently disabled by a declaration that never mentioned it.

Boot state is made true BEFORE the unit is started. When an apply fails part
way, enabled-and-stopped comes back at the next boot and running-and-disabled
does not, so the more durable half goes first.

`is-enabled` has the same trap as `is-active` had. Its exit code is non-zero
for nearly everything, and `static` is neither enabled nor disabled -- the unit
has no install section and CANNOT be enabled. Reading it as "disabled" would
have the host try, fail, and blame the wrong thing, which is the same shape as
reading a missing unit as "stopped".

The container applier no longer calls `docker` literally. Verified on this
machine against podman 6.1.0:

  docker info --format '{{.ServerVersion}}'    -> 29.7.2
  podman info --format '{{.ServerVersion}}'    -> Error: can't evaluate field
                                                  ServerVersion
  podman info --format '{{.Version.Version}}'  -> 6.1.0

So one probe cannot find both, and a host using docker's would report a machine
running podman as having no container runtime at all. Everything else IS
compatible -- run, rm -f, and docker's own Go template syntax for reading state
and labels all work unchanged on podman, confirmed by running them. That is why
this is a two-entry lookup rather than an interface: only the probe differs.

Detected rather than declared, because adoption keeps what the machine already
has (research 012), which hardcoding one runtime contradicts.

A machine with neither now says so, naming both: "docker: command not found" on
a machine deliberately running podman sends the reader after the wrong thing.

Verified end to end against real docker (container created, running, labelled)
and against an empty PATH (refused, naming both runtimes).

Two injections per behaviour, all confirmed to bite. One injection produced a
build failure that my check read as "no bite" for the third time, so the check
now distinguishes them.
2026-08-27 23:58:44 +02:00
jschoubben 9a9937b7e6 A struct per resource kind, instead of one struct with every field
Jochen asked why we don't simply have dedicated structs. We should, and the
flat struct was me extending an existing pattern rather than questioning it.

Before: one Resource struct carrying path, content, mode, unit, state, package,
image, name, env, ports, volumes, args, command, verify and in. Because a file
and a container shared it, nothing stopped {"type":"file","image":"postgres"},
so a `uses` map listed which fields each kind was allowed to carry -- a second
place to keep current, and the kind nobody updates is the one that silently
accepts a field the host will never read.

Now: Directory, File, Service, Package, Container and Action are separate
structs behind a Resource interface. File has no Image field, so the mistake is
not detected -- it is unrepresentable. Adding a field to a kind is the whole of
adding it; there is nowhere else that has to agree.

Parsing is two passes: read the envelope and each resource's raw bytes, peek at
"type" to choose the struct, then decode into it. Peeking is lenient on purpose
-- reading strictly there would report an unknown field before knowing which
fields are known.

Unknown fields are found by comparing the JSON keys against the struct's own
json tags rather than by catching the decoder's error. The decoder stops at the
first unknown field, and RefusalError promises every problem at once: a caller
fixing one field at a time learns the next only by running again. Caught by
testing the refactor against a real declaration -- a container carrying both
`unit` and `mode` reported only one of them.

apply.go switches on the concrete type instead of a string, so a new kind that
has no applier is a compile error rather than a runtime default branch.

No behaviour change otherwise. All existing tests pass unmodified except two
that reached for fields the interface no longer exposes.
2026-08-27 21:03:59 +02:00
jschoubben 337126603e Complete the host's vocabulary: package, container, action
The three shapes the substrate bootstrap needs and the host did not have. Until
now tier 1 could not be raised at all -- step 0 is a package, step 1 a
container, steps 2 and 3 actions -- so every line of the tier 1 and 2 designs
was unbuildable.

package -- present, never upgraded, never uninstalled. Removal is "forgotten",
not "removed": the host cannot know what else needs the package, uninstalling a
container runtime because a declaration changed would stop every container on
the node, and the machine may have had it before the mesh saw it. Reporting it
removed would claim an effect the host declined to have.

container -- identified by a label carrying a digest of the declaration that
made it. Comparing every field the runtime reports cannot be done reliably: a
runtime normalises, defaults and reorders what it is given, and that is
indistinguishable from real drift. There is no in-place update; a container's
configuration is fixed at creation, so any change is a replacement, and saying
so beats a partial update that leaves the running thing half-declared. This is
the one shape the host removes, because it is the one the host created.

action -- bundle-only, per ADR 0047. Verify is mandatory and does double duty:
it is the idempotency check as well as the read-back. The host does not know
what a database is, so "is it already there" is a question only the declaration
can ask. `in` runs the action inside a named container, which steps 2 and 3
need.

Parse now refuses actions; ParseTrusted permits them. The safe path is the
default and the permissive one has to be named. The bundle and a local file
handed to a root process use ParseTrusted; the link will use Parse.

Also replaced the per-type "fields this type ignores" check with a field-set
diff stated as what each type USES. The negative form needs every type revisited
whenever a field is added, and the one nobody revisits silently accepts a field
it will never read.

Images must be pinned by digest (ADR 0046). A bundle naming a tag pins nothing.

Verified against a real machine, not only fakes: an action ran and was
idempotent on the second apply; an action that exits zero and satisfies nothing
fails the apply; a real container was created, labelled, replaced when its
declaration changed, exec'd into, and removed; a real package query round-
tripped. Each new test was also confirmed to fail on an injected fault -- five
injections, each breaking exactly its own test.

One existing test changed: a vanished unit is now reported "forgotten" rather
than "removed", which is what actually happened.
2026-08-27 20:36:27 +02:00
jschoubben 9d8239afe8 Stage 2 — the host applies a declaration
A declaration is JSON, versioned, and an ordered list of resources with stable
identities (novox/hq ADR 0043). The vocabulary is directory, file and service,
and anything outside it — an unknown version, type or field — refuses the WHOLE
declaration. A host that skipped what it did not understand would apply most of
what it was sent and report success.

It converges rather than executes: applying twice changes nothing the second
time, and applying to a drifted machine returns it. A mode is maintained rather
than set, because a permission applied at creation is not a permission held —
this repository has paid for that once already.

It owns a footprint and only that. What it applied and is no longer declared is
removed; what it did not create is never touched. Removal runs FIRST, because a
resource leaving a declaration while another arrives at the same path is an
ordinary rename, and removing afterwards would delete the file just written.

The store arrives here rather than at stage 3, as ADR 0043 predicted: nothing
can be removed without knowing what was applied. It is written atomically,
refuses to start empty when it exists and cannot be read — believing it owns
nothing would leave everything behind forever — and is saved even when an apply
fails, because what was applied before the failure is on the machine either way.

Three faults found by running inside a raised machine rather than by reasoning:

A unit that DOES NOT EXIST reads as `inactive` from `systemctl is-active`,
exactly as a stopped one does. So declaring a unit stopped reported success for
a unit the host cannot manage at all — absence read as satisfaction, which is
04-ISSUES/007 wearing a different hat. LoadState separates them.

Removing an orphaned service whose unit has since been uninstalled failed the
whole apply, and a host holding such a record could then apply NOTHING, ever,
with no way out but editing its state by hand. Removal is now idempotent for the
same reason os.RemoveAll is.

And the flag parser was wrong in the same way twice: fixing `mesh-host inventory
--json` by taking the subcommand off the front left `mesh-host apply decl.json
--dry-run` broken identically, because the standard library stops at the first
non-flag argument wherever that argument is. Parsed in a loop now.

30 new tests, 55 in total.
2026-08-26 02:14:25 +02:00