Commit Graph
83 Commits
Author SHA1 Message Date
jschoubben c5c42376c1 The installer delivers any MESH_…_FILE own secret, not only the store's (ADR 0086) 2026-09-21 10:10:33 +02:00
jschoubben d756effc33 The broker-admin marker ends in a newline, and the transcript never says a credential
The verify reads the marker with the shell's read, which fails at end of file
without a line ending; the action ran and its verify said no. And the applier
reports each action with its command line, two of which now carry the real
store and broker passwords — the installer masks the values it made in
everything it says.
2026-09-21 01:32:51 +02:00
jschoubben 70d0f36896 Installer review: secrets are staged privately, and a bundle is 0600 whether or not it existed
From review: the store and broker passwords genesis makes were carried into
the controller through a world-readable file in /tmp, a bundle left at 0644 by
an earlier installer kept that mode while now holding them, a mesh raised by
the old installer would have been handed new passwords its servers do not have,
and the broker-admin action's marker did not depend on the value. Secrets now
stage in a 0700 directory owned by the controller's account; the bundle is
chmod'd; an existing store or broker volume with no credential file is refused
by name; the marker holds the password's fingerprint. Also: one install path
for the store, broker and vault, no error-string matching for the operator
key, and no unreachable fallback for the superuser.
2026-09-21 01:26:35 +02:00
jschoubben 036af3cfdc The export is its own installer step, and the result names the operator files 2026-09-21 01:11:52 +02:00
jschoubben ee0c8b856e Genesis makes the root secrets, the operator key, and installs the vault
The template raises the store with the password 'bootstrap' and the broker
with its image's default administrator, and the installer carried both into
the mesh as accepted secrets — permanent, and not secret (novox/hq issue 071).

Now the installer makes both credentials, once, at the paths the postgres and
lavinmq modules declare as their own secrets, rewrites the produced bundle to
use them (the store reads its password from a file; the broker's default
account is given the new password by an action before anything dials it), and
writes the bundle at 0600 since it now carries them.

Before the first secret is accepted it makes the operator's sealing key beside
the bundle and gives the mesh the public half, so everything minted from there
is sealed to it too (ADR 0085, amended). Phase three adopts the broker as the
lavinmq module beside the store and installs mesh-vault as a foundation module;
the run ends by writing the operator-sealed export beside the key.
2026-09-21 00:12:55 +02:00
jschoubben 1232031fb9 Correct store phrasing — a module gets a database only if it asks
Not "every module's database"; a module requests one via requires
postgres-database. The one server holds the controller's contexts and the
database of each module that asks for one.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:20:00 +02:00
jschoubben 9109a8c178 Phase 3.1: adopt the foundation store as the postgres module
InstallStore turns the mesh-store the foundation raised at genesis into the
postgres module, adopted in place: it verifies the module's server names the
same container and the same image the foundation is running (fail-fast on a
drift, rather than tearing down the mesh's store), then registers, builds the
provisioner, and carries the superuser in via secret accept — the mesh cannot
invent a credential that already made the databases (mirroring the control
plane's store-connection delivery, control.go). pinImage generalised to any
module for reuse.

Issue 051 (WBS 3.1). One server holds the controller's contexts and every
module's database.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:04:06 +02:00
jschoubben 121367319d Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 01c7730fb3 Genesis raises gitea correctly: host network, honest SQL, matched ROOT_URL
Fixes found raising the package registry end-to-end in the lab: seed gitea's DB
with plain psql statements (no \gexec, no $$ DO-blocks that clash with the
shell); run gitea on the host network so it reaches the substrate store and
answers where the builder looks; set gitea ROOT_URL to the machine's loopback so
npm's stored credential matches the tarball host; keep the pivot's passwords so a
re-run is the same run; create the admin without re-enabling must-change-password.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 14:11:36 +02:00
jschoubben 1a7c9fdac3 gitea runs on the host network at genesis
So it reaches the substrate store's loopback-published postgres and answers where
mesh-bootstrap and the builder look for it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:33:07 +02:00
jschoubben 79863068fc Genesis raises the package registry before it builds the base
The base (mesh-tools) resolves the SDK by version from the mesh's package
registry rather than cloning it from a git URL (hq ADR 0076, issue 053), so the
registry has to answer and the SDK has to be in it before the base build runs.

New steps, before base: seed gitea's database in the substrate store, raise
gitea's server on it, create the admin/org/team and the builder's account, seal
the builder its registry credential, and publish the SDK on a public base. gitea
is adopted as an ordinary module after the base, so its provisioner image can be
built. A minimal Go gitea admin client stands in until that module exists.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:27:11 +02:00
jschoubben 21474b0144 The installer goes as far as it can, and asks where a human must choose
Twelve steps made a mesh that RUNS and then said "what remains is somebody
else's". The seven things that turn it into a mesh that WORKS — the shared base,
a database provider, the catalogue, the private network, the packet filter — were
typed afterwards, which is how they went missing for weeks without anything
complaining.

Six more steps now: base, store, catalogue, network, filter, extras. Everything
in them is module add, build, assign and push — the same verbs a person types,
through the same commands, so the installer and an operator remain one act.

Where a human must choose, the installer asks. A choice resolves in the order a
person expects: the flag wins; a lone option answers itself ALOUD, because "it
chose for me" and "there was nothing to choose" read identically afterwards
unless one speaks; a terminal is asked; a default fills in; and a required
choice nothing answered refuses naming its flag — a guessed packet filter is a
machine somebody else configured. The filter is required, so the question is
which, not whether. A run without a terminal (the lab, --json) is never left
waiting on a prompt nobody will answer.

Placement is part of the network step, not a separate act — a lesson paid for:
the module installed, the names file was written with no names in it, and
everything reported success because nobody had said where the machine IS. The
hub endpoint derives from the broker address when unsaid: the host other
machines dial is one fact, not two that drift.

Extras fail the run rather than soft-fail: somebody asked for them by name, and
a mesh reporting success minus one thing is reporting the wrong thing.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 21:56:23 +02:00
jschoubben 7a223464e5 Genesis installs distribution, which is what the software is called
The module that provides artifact-store runs Distribution, the OCI reference
implementation. It was called registry, which named neither the software nor the
provision.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:51:00 +02:00
jschoubben de5160de4a A unit file reinterprets an environment value; a container does not
Found by being asked whether processes and containers handle environment the
same way. They do not, and the difference is not cosmetic.

Docker passes --env through literally. A unit file reads three things out of a
value that nothing else does, and a module's environment routinely contains all
three because a generated password is arbitrary bytes:

  - % begins a specifier. %H is the hostname. A password containing one is
    silently replaced, and it fails later as an authentication error nobody can
    explain by reading the declaration.
  - whitespace separates assignments. Unquoted, K=a b sets K to "a" and reads
    "b" as another assignment.
  - a newline ends the line, and what follows is read as a unit DIRECTIVE.

The first two are escaped: quoted, with quotes and backslashes escaped and
percent doubled. The third cannot be — a unit's environment has no way to carry
a line break — so it is refused in validation, near whoever wrote it. Without
that, an environment value could write ExecStart= and have the machine run
something nobody declared.

Ordinary awkward values stay accepted, because refusing those too would leave a
module unable to hold a generated password.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 12:40:36 +02:00
jschoubben f5cf9510c1 One kind for the module's own code, with three modes
The first cut of this added a `daemon` for the long-running case alone. That
would have meant a new vocabulary entry for each of the others — a scheduled
task, a run-once migration, a health check — when they are one thing run at
different cadences. That is a field, not four entries in a vocabulary where every
entry widens what a compromised control plane can express.

So it mirrors a container exactly, because it IS a container's twin: the same
intent, hosted by the machine's own supervisor instead of a runtime. Stays up,
runs once, or runs on a schedule.

Tools, hooks and event consumers are not further modes. They are loaded by a tool
host, which is itself a process that stays up — so the generic case already
covers them, which is the test of whether it is generic.

A scheduled process gets a timer and a unit that finishes; a long-running one
gets a unit that is restarted when it exits. Getting that wrong either way is a
second copy running continuously between fires, or a schedule that never fires.
The modes are exclusive and validation says so near the author: something that
runs once does not run on a schedule, and something not running between fires
cannot be restarted when a file changes.

A missed fire happens when the machine comes back rather than being skipped,
which is the difference between a machine that was down and a schedule that
quietly stopped.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 10:26:31 +02:00
jschoubben 1f0fb85128 A daemon says what to run, not how it is hosted
The mechanism was leaking into every module. Code of one's own meant a container
and therefore an image; a script meant a service and a unit somebody else had to
install. One intent — run this and keep it running — expressed two unrelated
ways, with the hosting chosen before anything could be declared.

A daemon names a bundle and a command. The host fetches it, refuses it unless it
hashes to what was declared, unpacks it where the mesh keeps such things, writes
the unit and puts it in the state asked for. The unit is the mesh's, generated
whole and saying so, because an edit that survives until the next declaration and
then vanishes is worse than one that is refused.

Its identity is the bytes AND how it is run: two daemons from one bundle
differing only in their command are different daemons, and tracking the digest
alone would call the second unchanged and leave the first running. The unit is
rendered deterministically for the same reason — environment from a map would be
written in Go's iteration order, so every apply would see a different unit and
restart an unchanged daemon for ever.

restart-on is honoured as a service's is: a running process does not re-read its
configuration, so replacing a file and finding the daemon already up leaves the
machine behaving as before while every check passes.

A full-host shape, not a portable one: it needs a process supervisor to install
into. It does NOT need a container runtime, which is the point.

Two guards caught this properly and both were updated deliberately rather than
silenced: the vocabulary count, which exists because every addition widens what a
compromised control plane can express, and the shape test that catches a kind the
language has and a host cannot apply — added after `network` did exactly that.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 02:29:33 +02:00
jschoubben e7f94e0402 A one-shot service that finished is not stopped, and a container is what it reads
Two faults that both reported success while being wrong, found while proving
the firewall module actually delivers.

A unit whose job is to apply something and exit — load a rule set, set a
sysctl — is inactive the instant it succeeds. Reading that as stopped made it
permanently unsatisfiable: the host started it, it worked, the host read back
stopped and reported the machine as not doing what it was told, on every apply,
for ever, with the rules correctly in place the whole time. That is what the
firewall has been doing on every machine it was assigned to, and why the
four-machine bed was red.

And a container took its identity from its own fields, not from the files it
reads. A file written in an earlier apply — or before the container declared it
as a dependency — left a process holding a credential the mesh had already
replaced, with everything reporting success (novox/hq 04-ISSUES/045). What a
container reads is now part of what it is, so the comparison is a standing one
rather than a tripwire that fires during one apply and never again.
2026-09-14 16:51:57 +02:00
jschoubben 8eeb28f00b Installation sets up the builder, so a raised mesh can produce
Genesis ended with a mesh that runs and cannot make anything: every module in
the catalogue names artifacts and nothing had built them, so the first thing
anybody had to do was install a builder by hand.

The installer already carries one — it is what built the control plane — so
this is the same two acts the control plane goes through, in the same order:
publish it, so the mesh names it by a digest its own registry assigned rather
than a local identity nothing else can fetch, then install it as an ordinary
module pinned to that. And then the part only it needs, a broker account, issued
before the push so it arrives with the declaration rather than after it.

Verified on a bare machine: the install ends with a builder running, and that
mesh then built the shared base images and a module on top of them with nobody
helping it.
2026-09-14 12:31:43 +02:00
jschoubben 163a44a49a Preflight names the carried image for what it is
It said 'control plane' beside a builder's tag, which is the sort of line that
teaches a reader the wrong thing about what the installer carries.
2026-09-13 04:38:44 +02:00
jschoubben 432e3edcfd Publish the image this mesh built, not the one the installer carried
The carried image is the builder now. The publish step still pushed it, so the
registry got a builder under the control plane's name and the mesh installed it
as the control plane — which presented as a control plane that started, printed
a builder's usage, exited cleanly, and did it again. Caught by the lab on the
first genesis run, at the step that waits for it to answer.
2026-09-13 04:24:18 +02:00
jschoubben e1a2fe7323 The installer carries a builder and builds the control plane it raises
It carried the thing it was going to run; it now carries the thing that makes
it. One artifact either way — but a mesh raised this way holds a control plane
it built from a repository and a commit it can name, and can therefore build
again. A mesh handed a finished image could not, and had no way to find that
out until somebody needed it to.

A build step sits between load and bundle, because the bundle must name an
image and that image no longer arrives finished. Everything after it is
unchanged: a locally built image is named by the digest of its own
configuration, which is exactly what the carried one was named by.

Refused in preflight when nothing says what to build, so a run that cannot
finish says so before it has changed anything.
2026-09-13 04:08:58 +02:00
jschoubben 26ff447aa3 bootstrap: the mesh hearing from a machine is not an agent running on it
Enrolling IS the machine speaking to the mesh, so straight after it the mesh has
always heard from this node — and the step took that as proof an agent was
running and skipped starting one.

The cost is silent and total. Everything after is the control plane being told
things, and nothing it is told reaches a machine with no agent to collect it: the
registry push at step 7 was accepted, the module recorded, and no container ever
created. It surfaced three minutes later as 'the registry is not there at all',
one step from its cause and looking nothing like it.

Both halves are asked now. A process may be wedged and collect nothing, which is
why the mesh is asked at all; and the mesh may have heard once from a machine
running nothing, which is why the machine is asked too.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 12:11:57 +02:00
jschoubben 7e3481f025 bootstrap: the slot the installer fills is not a registry to reach for
The first real run of mesh-bootstrap stopped in preflight, dialling 192.0.2.250:5000
for ninety seconds on a machine whose network was fine. That address is the registry
the lab used to raise; the substrate template still names the control plane by it,
and step 3 replaces that reference with the id of the image this installer carries.
Nothing ever pulls it.

So preflight excludes the control plane's resource by identity, rather than by the
happy accident of the template filling its slot with something that needs no registry.
Every other container's registry is still dialled, because those are somebody else's
images at somebody else's registry and a machine that cannot reach one fails inside a
pull, which says the wrong thing.

Also: `make bootstrap` takes BOOTSTRAP_OUT. The lab now builds the installer from
source before every raise, into a path it chooses, and a caller that could not say
where the output goes would have to copy it afterwards.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 11:38:15 +02:00
jschoubben cb5e137297 bootstrap: read the image id back from the runtime, never predict it
An image id does not survive `docker save` -> transfer -> `docker load`. The id is
the digest of the image's *configuration*, and a runtime rewrites that
configuration as it loads: a newer Docker saves in one format, an older one stores
it in another. Same layers, same program, different name. Measured on a live raise:

  saved on the workstation  sha256:b86bb81ca2f9691f24f4725f50962d1e49c98c5ffe211113241243d42d18ceea
  loaded on the machine     sha256:2dc219046c73702fc640317f0342a28ec962ef1e9ef547b2f02861c508ca78fb

`internal/image`.ID read the id out of the carried tar and its comment said that
was the id the runtime would assign. That is true on the machine the image was
built on and false on every machine it is carried to — which is every machine this
program exists for. The installer then either stopped at step 2 refusing the
runtime's answer, or would have written a bundle naming an image the machine does
not hold; and nothing serves an image named by the digest of its own configuration,
which is the whole point of naming one that way, so the apply would have died
inside a pull that cannot succeed. The lab hit this.

So the image is identified by its TAG, which is ordinary metadata the tar carries
through unchanged. The runtime is asked what that tag resolves to before the load
(already held, nothing to do) and again after (this is what the bundle names). The
tag never reaches the bundle — a pinned bundle may not rely on one, ADR 0006 — it
is how the id is obtained, not what is written down.

  - image.ID becomes image.ArchiveID, and says plainly that it is a fact about the
    file and not a prediction about any machine. It is kept for reports, and printed
    beside the runtime's answer whenever the two differ.
  - Idempotence is decided from what the runtime holds under the tag, not from a
    predicted id, which cannot answer the question at all here.
  - An untagged archive is refused, in preflight and again at the load: there would
    be no portable name to ask about, and the only thing left is scraping a sentence
    `docker load` writes for a person. `make bootstrap` refuses an id or an untagged
    image, so it is caught in front of whoever can fix it.
  - A dry run cannot know the id and says so rather than pretending. Run refuses to
    write a bundle carrying an unconfirmed id at all.

Tests: the injected Runner now answers with an id DIFFERING from the tar's, and the
runtime's answer is what must be used. The test that refused a differing id encoded
the mistake and is replaced by one refusing an answer that is not an id at all.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:12:36 +02:00
jschoubben 0a88dc1f9d bootstrap: follow the mount from the variable to the secret
The catalogue's mesh-control manifest landed while this was being written, and it
does what the ordinary case does: it keeps its secrets under /var/lib/mesh and
mounts them into the container at /run/secrets, so MESH_STORE_INVENTORY_FILE names
a path that no own-secret writes. Matching on the path alone found nothing and
would have refused a correct manifest.

So the lookup follows the volumes. It also reads the other shape the manifest uses
— `VAR=${secret:name}` inside the environment file a container reads — which is
how a value that is not a path gets in at all, and which is where the broker's two
credentials live.

That generalises what is delivered: every variable the module fills from a secret
is looked up in the substrate's control plane. What the substrate names is accepted
through `secret accept`; what it does not is left for the mesh to generate, and
said so. A store connection the substrate does not name stays an error — a control
plane that cannot open a context is not one.

Checked against the real manifest (mesh-catalog feat/control-plane-module): five
variables resolve, the placeholder pins in one place, and the container it waits
for is `mesh-control`.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-11 00:03:14 +02:00
jschoubben f534cf8b42 bootstrap: the rest of the pivot — enrol, registry, publish, reinstall, retire
Steps 6 to 10, which turn a substrate into a mesh that can maintain itself
(novox/hq ADR 0067).

 6 enrol      a node record, a token, `mesh-host enrol`, and the host agent
              running. Proved by the mesh having HEARD from the node, not by a
              process existing: a host that cannot reach the broker looks exactly
              like a successful install until the first push applies nothing.
 7 registry   the module that gives this mesh an image store, registered from a
              --catalog checkout, assigned and pushed. Its image is upstream and
              never built (04-ISSUES/029) — a placeholder digest there is refused.
              Verified by asking `/v2/`, because a container that is up is not a
              registry that serves.
 8 publish    the carried image pushed into that registry, which assigns it the
              first manifest digest it has ever had. This is the hinge: without
              it the mesh works and can never upgrade itself.
 9 control    the control plane registered as an ordinary module pinned to that
              digest, with the substrate's own store connections delivered
              through `secret accept` — read out of the bundle that made them,
              because the mesh cannot invent a credential that predates it.
10 retire     the temporary control plane dropped from the bundle and removed by
              the host's ordinary removal pass.

Every step asks before it acts and reports "already done". No step leaves the
machine without a control plane: steps 9 and 10 overlap deliberately, and two
stateless control planes are untidy rather than broken.

mesh-control's `internal/builder`.PublishImage is mirrored rather than imported —
tier 0 depends on nothing that must be installed first — with one correction: the
digest is chosen from RepoDigests by repository instead of taken as element zero,
so an image pushed to two registries cannot silently pin this mesh to the wrong
one.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:59:57 +02:00
jschoubben af953dbb0d bootstrap: the temporary control plane gets a temporary name
The substrate raises a control plane and a module will later declare one. If both
are called `mesh-control` then for one moment two owners hold one container, and
the host — which tracks what it owns — has no way to stop owning something without
destroying it. That looked like a missing mechanism.

It is a naming problem. The substrate's container becomes `temp-mesh-control` and
the module's keeps the plain name: two containers, two owners, nothing to hand
over. Dropping the temporary one from the bundle at the end is then destruction by
omission, which is what the host already does to anything that leaves a
declaration — and the right end for something named "temp" (novox/hq ADR 0067).

The rename is textual and matches the QUOTED name, so the `mesh-control` inside
the image reference is not caught by it. Read back afterwards: the produced bundle
must call it the temporary name, and no other container may have been renamed.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:59:32 +02:00
jschoubben 9c9e02ba1d mesh-bootstrap: drop an unread parameter
A parameter nothing reads is a claim the function makes about what it needs, and
this one was wrong.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:18:07 +02:00
jschoubben b82ab95f74 mesh-bootstrap: the first-node procedure, as a program rather than a test
The only complete written-down copy of how a mesh is stood up was an integration
test in the lab. That is why every bootstrap gap kept being found late: an install
procedure that lives as a test fixture is exercised by whoever writes tests, never
by whoever installs. This is that procedure.

A separate binary, not a mesh-host subcommand. mesh-host says of itself that it
connects to nothing and listens on nothing and that what it applies comes from a
file, and that sentence is what makes an always-running root daemon auditable. An
installer loads images and interrogates a control plane. Same tier, different
program.

The control plane's image is carried, not built and not fetched. The forge that
holds its source runs on the mesh, so a bootstrap that had to fetch it would need
a mesh in order to raise one. Embedding breaks that cycle the way the carried
bundle breaks "copy it onto a machine and run it". The image id is read out of the
saved tar before the runtime is asked anything, which is what makes the load
idempotent: the installer can ask whether the machine already holds exactly this.

Five steps, each idempotent and each saying whether it found or changed something,
because this is run over and over by somebody getting a machine working. It stops
at a running substrate with a control plane that replies — enrolment, the module
catalogue and assignment are the next stage and are deliberately absent.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:17:30 +02:00
jschoubben 4af9483219 declaration: an image may be named by the digest of its own configuration
A manifest digest is assigned by a registry on push, so insisting on one meant a
registry had to exist before the thing that lets a mesh have a registry could
start — a dependency the pinning rule created by accident, not a pin. The mesh's
own control plane is built from source and lives in no public registry.

A bare sha256:... names an image the machine already holds, by the digest of its
own configuration: immutable and unforgeable in exactly the way the rule asks
for. Absent, it says so plainly rather than failing at a pull nothing serves.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 22:40:02 +02:00
jschoubben e5af6bb58e apply: ensure a scheduled container's image is present at apply, without running it
A schedule: container (ADR 0053) is installed as present state and never run at apply — the
Scheduler fires it later on its cadence. But a service or run-once container only gets its
image as a side effect of docker run, so a scheduled step's image was not pulled until its
first scheduled fire: absent from the node right after a successful apply, so the first run
paid the whole pull latency and tooling that expects the image present after apply found it
missing.

applyContainer now probes the runtime and ensures the pinned image present for a scheduled
step before recording it. A new ensureImage helper inspects the image and pulls it only if
absent, then reads back (ADR 0018). Ensuring an image is not running it: no docker run fires
the container, so the no-run invariant of ADR 0053 holds. The runtime probe, previously
skipped for a schedule, now runs because a pull needs it — the schedule.go comment is updated
to match.

Tests: the install-does-not-run test is extended to allow the image-ensure while asserting no
fire and no needless pull; a new test applies a scheduled container whose image is absent and
asserts it is pulled and still not started. go build, go vet, go test ./... all pass.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 02:23:46 +02:00
jschoubben 9d1f001dcc apply: a scheduled step is a container run on a cadence (ADR 0053)
The recurring twin of run-once, one modifier over: a container marked
schedule: "<cron>" is run to completion on its cadence, not started as a
service and not run once as a gate.

The gating rule is deliberately reversed. Installing a schedule records it
as present state and reports the node current at once (applySchedule) --
it never runs the container and does not gate what follows. A Scheduler,
held for the life of the daemon and re-established from each applied
declaration (the declaration is the source of truth, ADR 0018), fires the
container off an injected clock. A run that exits non-zero is logged and
never fails the apply or flips the node's state, because it happens
outside the apply and the store entirely. Runs never stack: a run still
going when the next is due is skipped, not started as a second copy.

No new host shape and no new action -- schedule is a string on the
container the host already has, and the host process runs the container
itself rather than installing a system timer (the rejected option 1). A
minimal five-field cron (declaration/cron.go) validates on arrival and
computes the next due minute; time is injected so the scheduler is tested
without the wall clock.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-06 14:08:44 +02:00
jschoubben 19e5dd83ea apply: a run-once container is a step the host runs to completion (ADR 0052)
A module can declare state but not a step that runs at first boot. This adds
`run-once: true` to the container shape: the host runs it in the foreground,
requires it to exit 0, and records that it did — as the digest of the
declaration, so a re-apply does not re-run it unless the declaration changed.

Because the declaration is applied in order and a failed run-once step gates the
apply the way a failed action does, whatever is declared after the step starts
only once it has completed. That is how "before the broker starts" is enforced,
with no dependency graph the host must resolve (ADR 0005): the step is declared
first, and the container that needs it is never reached until it is done.

No new host shape and no arbitrary host command — a run-once container is
strictly less powerful than an action. Validation refuses run-once with
restart-on (contradictory lifecycles). Six unit tests; go test ./... green.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 23:55:45 +02:00
jschoubben f06eea5fa3 declaration: an access is mounted, and the host owns nothing about it
The tenth shape (novox/hq ADR 0051). Shared, pre-existing data — a media
library, a download spool several modules use — is the operator's, not
the mesh's. A `directory` resource is the host's own: it creates it,
chowns it, sets its mode and removes it when empty. An access is the
opposite on every axis.

Add the `access` type to the vocabulary. Its applier confirms the path is
present and changes nothing: it does not create, chown, reconcile or set
a mode. Absent is refused clearly — the operator must provide it — rather
than created, because a bind mount whose source is missing is made as
root by the container runtime with the wrong ownership (04-ISSUES/026).
Undeclaring an access forgets the record and never touches the path,
which is the data loss ADR 0030 prevents, on a directory the mesh never
made.

Full hosts speak it (it gates a bind mount, which needs the container
runtime); the vocabulary guard test records the decision that made it the
tenth shape. Unit tests cover present, absent-refused, and
undeclared-left-alone.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-05 22:19:33 +02:00
jschoubben aa441bac19 A container reflects its config: restart-on for containers (04-ISSUES/009)
A container reads a mounted file once, at start; its spec (image, env, volumes)
does not include a mounted file's content, so a settings change that re-renders the
file left the running process holding the old value while every check passed. Give
Container the restart-on field a Service already has, and recreate the container
when a named resource changed this pass. Unit-tested (recreated on change, left
alone otherwise) and proven in the mesh-lab: a running grafana runtime picked up a
token change on the next push.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-04 23:39:47 +02:00
jschoubben 8211d8b6fb A report says which declaration it is about
The mesh decided "has this machine caught up" by comparing its send
time to the report's arrival, and lost the race it invited: an apply
started under the previous declaration finishes after the next one is
sent, its report lands newer than the send, and the machine reads as
caught up with words it has not read yet. The lab hit exactly that —
one test's closing push was still being applied when the next test's
push recorded its send, and the next test then read files that were
never going to be there yet.

Clocks cannot answer "which". The report now carries the digest of the
exact bytes it applied — the same bytes, hashed the same way, that the
mesh recorded when it sent them — and which-declaration becomes an
equality the mesh checks rather than an ordering it hopes.
2026-09-02 00:01:12 +02:00
jschoubben aca9eb37ff An owner may be a number the machine has never heard of
A directory a module mounts into its container belongs to whoever runs
inside — grafana's 472, redis's 999, www-data's 33 — and none of those
has a row in the machine's passwd. Owner-by-name refused them all,
which looked principled and meant every module whose container drops
privileges could not own its own data.

The lab showed both coats of it in one run: the store's config file was
unreadable to the store, restarting forever on permission denied, and
the forge could not traverse into the 0700 root-owned directory that
held its files — a directory that had only become root-owned when
declaring it fixed 04-ISSUES/026, because Docker used to create it
0755. A fix that tightens ownership without a way to say whose it
should be moves the fault, not removes it.

"uid:gid" and bare "uid" are numeric and chowned as given; a name still
resolves as before, and a name with a colon is refused rather than
half-read.
2026-09-01 23:45:02 +02:00
jschoubben b91342a6bd A machine says which ports it already holds
novox/hq ADR 0038 and 04-ISSUES/028. The substrate is not a module: a
node raises it from the bundle it carries before any mesh exists, so the
control plane has never heard of the store, the broker, or the control
plane's own container. A module assigned afterwards is handed a port one
of them holds, and finds out from a container runtime three layers down.

The host already recorded which resources it carried and which the mesh
sent — that distinction exists so the two never remove each other. It
now also records what each one binds, and reports the carried ones.

What the declaration binds, not what is open. A machine's open ports are
a moving target — something a person started, a connection the kernel
handed out — and assigning around those would mean a port that was free
when it was asked for and taken when it was used. What a resource
declares is stable, and it is the half the mesh can be responsible for.

Only the carried ones are reported. What the mesh put here it already
knows, and reporting it back would make the machine an authority on the
mesh's own bookkeeping.
2026-09-01 18:29:39 +02:00
jschoubben a7a2a48615 A full host implements the network shape, and something checks that it does
The shape was added everywhere except the one list that decides whether
a host can actually apply it, so every declaration carrying a network
was refused whole — correctly, and with the reason stated:

  resource "umami.net" is a network, and the arch host does not
  implement that shape

The mesh behaved as designed throughout. A host that applied the parts
it understood would leave a machine that looks configured and is not, so
it refused the lot and said why. What was missing was anybody reading
the host's log.

The vocabulary test did not catch it because it checks what the language
has, not what a host can do — those are different lists and only one of
them was updated. There is now a test that a host claiming to do
everything implements every shape the language has. It fails with the
message above when the registration is removed.

A network needs the same runtime a container does, so it belongs to a
full host and not to the portable floor.
2026-09-01 00:36:17 +02:00
jschoubben 8c248e3d7f A secret can reach a container's environment, and sit inside a config file
Two gaps found by writing the first real module's manifest rather than
by reasoning about one. Both are fields on existing shapes, so the
vocabulary is still nine.

**env-file on a container.** A declaration reaches a node over the
broker and `env` is plain text in it, so a password there is a password
the broker sees — the transitive trust refused everywhere else. A sealed
file arrives unreadable, the host writes it, the runtime reads it. It is
also simply how third-party software takes credentials: nothing shipping
in a container will read a path the mesh invented, and every one of them
reads its environment.

**secrets in a file's content.** A program wanting its token inside a
JSON document cannot be handed a file that is entirely a token, and the
mesh cannot compose the document because it discarded the value. So the
module supplies the document with `${secret:name}` in it, the mesh
delivers the value sealed, and the host is the only thing that ever
holds both.

Substitution is textual and the host learns no formats. Deliberate: a
mechanism that understood JSON would be asked to understand YAML next,
and then INI, which is how the arrangement this replaces became
something nobody could hold in their head. The module knows its own
format because it wrote the rest of the file. The sharp edge is stated
rather than left to be discovered — a value containing a quote is not
escaped for whatever surrounds it.

Refused in both directions, because both are somebody being wrong about
where a credential is: a placeholder with nothing to fill it would write
`${secret:x}` into a config file, and a secret the content never uses
means somebody believes a credential is in a file where it is not.

A file that carries one is 0600 unless the module said otherwise.
2026-08-31 22:25:57 +02:00
jschoubben f48e06473d A directory holding anything the mesh did not put there is never removed
Found by asking what the conversion needs, and it is the one failure in
this system that cannot be undone.

Unassigning a module made its directory an orphan, and an orphan
directory was deleted with everything under it — os.RemoveAll — while
the report said "removed". A database's files, a mail spool, somebody's
uploads. Reproduced before fixing: assign a module, let a service write
into its directory, unassign the module, and the file is gone.

Now a directory that still holds something is kept and said so, naming
how many items are in it.

What makes that safe rather than merely cautious is the removal order,
which was already right. Everything the mesh puts in a directory is
itself a declared resource, and orphans are removed in reverse
declaration order — so what the mesh wrote is already gone by the time
the directory is reached. Anything still there was put there by
something else, which is the definition of data.

It is the host's own line applied to the one shape where getting it
wrong does not recover: it removes what it made and leaves what it
merely configured. An empty directory is what it made; a full one is
not, and an empty one is still removed so nothing accumulates.

Files are unchanged. A declared file is the mesh's own, and losing a
config file is not the failure this is about.
2026-08-31 19:53:41 +02:00
jschoubben 4a43e21794 A network is a shape, so that it can be removed
novox/hq ADR 0029, and work breakdown 1.3. A module of several
containers had no way to let them reach each other by name: a container
declaration could join a network and nothing could create one.

An action was the obvious alternative and is refused on removal —
"an action has no footprint the host can undo", so a network made that
way outlives every module that is ever unassigned, and the mesh cannot
tell. A resource the mesh can create and never clean up is one it should
not create.

A name and nothing else. Not a driver, a subnet or a gateway: each is
something a module would have to know about the machine it lands on, and
a module naming a subnet collides with whatever else chose the same one.

It needs no new ordering rule. Resources apply in declaration order and
orphans are removed in reverse, so a network written before the
containers that join it is created first and removed last — after they
are gone. A runtime refusing to remove one still in use is reported
rather than swallowed, because that means something undeclared is
holding it.

The vocabulary guard fired on the change, as designed, and now names the
record instead of a number: nine shapes, with the argument beside the
count.

Creation reads back rather than trusting an exit status (ADR 0018): a
runtime that reports success and made nothing leaves every container
that joins it failing to start, one step from the cause.
2026-08-31 18:55:06 +02:00
jschoubben af9d316258 Resources are applied in the order they were declared, and now something says so
Half of novox/hq work breakdown 1.3, and it needed no change: the apply
loop walks d.Resources and sorts nothing, so a module that needs one
thing before another says so by writing it first.

Asserted because it is the kind of property a later change breaks
silently. Sorting the resources for any good reason at all — by type,
by identity, for a tidier report — would still pass every other test in
this package.

It is sequence, not readiness. A container started is not a container
ready, and nothing here waits: what depends on something being usable
retries, which is what both example provisioners do and is the more
robust answer anyway, because a dependency can restart long after
everything was applied.

Two mistakes worth keeping in the test's own comments. The first
version stubbed the runner to always succeed, so verify passed, every
action counted as already done, and nothing ran — the assertion was
measuring an empty list. The second declared the actions over the link,
which refuses them: only a bundle may carry an action (ADR 0005).
2026-08-31 18:50:10 +02:00
jschoubben 8e12b3c9e4 Name the decisions these tests defend, and check the bundle at all
From auditing the decision records: of 28, only 12 were named by any
test, so "which decisions are defended" could not be answered without
reading everything. ADR 0017 says a test names the decision it defends —
that rule was itself unenforced.

Most of the gap was citation, not coverage. Drift detection was tested
in several places without naming ADR 0011; the archive refusal without
naming 0012; forged declarations without naming 0002. Named now, so the
question is answerable by grep.

The bundle was the real gap: nothing tested substrate-first-node.lock at
all. It is what a machine becomes when there is no mesh to ask — the one
declaration applied with nothing to verify it against — and it was
edited by hand and read by nothing but a running host.

Two tests now assert what it carries: exactly postgres, lavinmq and the
control plane. That defends ADR 0028, which removed the object store
from the substrate after it had been a member for months on the strength
of "it cannot grant itself a bucket" — true, and the answer to only half
the test. Nothing counted what the bundle held.

Fault-injected, and the first attempt did not bite: the injection landed
on a comment line, which stripComments discards. Injecting into the
image field fails as it should.
2026-08-31 17:33:34 +02:00
jschoubben b4d2e851a3 Name a stale package index, rather than reporting a failed install
novox/hq 04-ISSUES/002, which was recorded against HAL and is present here: a
machine asking the mirrors for a version they have already replaced gets a 404
from every one of them. The package exists and the declaration is correct — it
is the machine's view that is old — and reported as a generic install failure
it sends somebody to check the manifest, which is the one thing that is right.

It is deliberately not fixed by syncing. `pacman -Sy <pkg>` installs a package
built against libraries the machine does not have: a partial upgrade, which
this distribution does not support and which surfaces much later as something
apparently unrelated. The remedy is a full upgrade, which is a decision about
the whole machine rather than something a host does silently while applying one
resource. So this says which of the two it is looking at, and leaves the
decision where it belongs.

Every mirror, not one: a single mirror timing out is transient and retrying is
the answer.

And the package manager's own words were being discarded entirely — the output
was read into `_`. Whatever it said is now part of the failure, which is the
rule everywhere else here and was not being followed in the one place the
reason only exists in the output.
2026-08-31 12:59:53 +02:00
jschoubben c3d6f240fe Give a container the names, rather than a resolver to ask
The commit before this said "told where to resolve names" and passed --dns,
which is not what it ended up doing. This is that correction: a container is
given the names themselves, written into its own hosts file by the runtime.

The reason for the change is the decision the mesh already made about names — a
file rather than a resolver, because it works on every runtime, needs no
package and has no failure mode of its own. Passing a resolver address would
have required a resolver to exist, which at that point none did.

A resolver is coming, for the case a file genuinely cannot express: a service
named under a machine, postgres.novox.internal, where the wildcard cannot be
enumerated in advance. When it arrives it will need this field back under its
own name. It is not being kept in the meantime — a field nothing fills is a
field nobody can trust, and the vocabulary is asserted by a count for exactly
that reason.
2026-08-31 12:05:52 +02:00
jschoubben 0e2b288bb6 A container can be told where to resolve names
A container does not inherit the machine's names. It gets its own /etc/hosts
holding its own hostname, and a runtime rewrites resolv.conf — so every
internal name the mesh wrote for that machine is invisible to what the machine
is running.

That was hit for real, in the lab: a database client on one node could not
resolve another node, on a mesh where both names were correct and present on
both machines. It was worked around by resolving on the host and passing an
address, which is the kind of workaround that should not be needed twice.

A field on an existing shape, not a ninth shape — the vocabulary is still the
eight the count asserts.

Per container rather than by editing the machine's resolver configuration: that
file belongs to something else on most machines, and a host that edited it
would be fighting whatever owns it on every boot — the fault this host exists
to avoid, in the place it would be hardest to see.

A container told nothing is run exactly as before. Most containers should
resolve whatever the machine resolves, and passing an empty flag would be a
change of behaviour dressed up as a default.
2026-08-31 11:09:49 +02:00
jschoubben 8fcfa88fe0 A machine that wakes or moves says so, instead of waiting to be told
A suspended laptop's connection is dead the moment it wakes, and the socket
looks perfectly healthy from inside the process — no error, no close, because
nothing has tried to send anything. Heartbeats find out twenty or thirty
seconds later. For that time the node believes it is in a mesh it has left,
which is the one state this design says must never be indistinguishable from
being connected. The machine knew immediately.

So being roused ends the current attempt rather than only shortening the wait
after it: shortening the wait would do nothing at all, because the process is
not waiting — it is sitting inside a connection that will not return.

A signal, because nothing may listen on a node (novox/hq ADR 0004). A socket
for this would be a control surface on every machine, reachable by anything
that can reach the machine, in exchange for saving twenty seconds — and the
whole security argument rests on there not being one.

Two rouses in the same instant are one: a machine suspending and resuming
repeatedly must not build a backlog of reconnections to work through. And the
backoff is not reset by being roused — that says the machine changed, not that
whatever was refusing the connection has stopped, and a laptop woken on a
network with no route would otherwise retry at full speed for as long as
somebody keeps opening the lid.

The dispatcher acts on the events that change where packets go and not on
`down`: the link is already gone there, reconnecting will fail, and the backoff
exists for exactly that.
2026-08-31 10:21:26 +02:00
jschoubben f9c70a7c97 Something after the declaration is refused whole
A JSON decoder reads one value and stops, so a file holding a declaration and
then anything else parsed as the declaration and the rest was never looked at.
The machine applies something, reports success, and what it applied is not what
the file says — the same fault this host refuses everywhere else, in its
quietest form.

Not hypothetical. A test harness had been appending a line to the substrate
bundle by accident; every apply kept working and nothing said so for as long as
it was wrong. That is how the bug was found, and it is the argument for the
refusal: a file with something after it may be a truncated rewrite or two
declarations run together, and applying the first would be applying something
nobody wrote.

Trailing whitespace is not "something after it".
2026-08-31 00:51:29 +02:00
jschoubben c83ed4eca9 A node's serving key is stored in the format a server reads
PKCS#8 PEM, not this host's own base64. The mesh delivers a PEM certificate
beside it and every TLS server there is reads PEM: nginx's ssl_certificate_key,
Go's LoadX509KeyPair, openssl s_server. Stored the other way the file was
intact, present, correctly permissioned, and unusable — the machine failed at
the moment something connected, which the lab found by connecting.

A key in the old encoding is refused by name rather than called corrupt: it is
replaced by enrolling again, and that is a different remedy from a damaged
file.
2026-08-31 00:41:10 +02:00