Review of the first cut found four things.
A directory mounted into a container is no longer looked inside, not even
for the files this host wrote there. The controller records every
provider's received and contributions file as a plain file under a mounted
directory, so folding those in would have recreated the route proxy — which
re-reads its routes live, by design — on every route change, and killed
every provisioner sidecar, which polls what it receives, mid-reconcile on
every grant. Whether a service reads a file under its directory once or
watches it is the service's; restart-on is how a module says "once", and it
stays the opt-in. Env-files and files mounted directly remain by content.
Genesis wrote the superuser secret as `value\n`; `secret accept` strips the
line ending by design, so the postgres module declared `value` — and with
a mounted file's content in the spec, phase three would have recreated the
store it meant to adopt in place, with the temporary control plane
connected to it. Genesis now writes the value alone. readCredentialFile
tolerated both endings already. Pinned with the bytes the genesis code
path writes, then the module's declaration of the same container: it must
reconcile.
A container carrying a label from before the host folded in what it reads
is accepted rather than recreated, when that label matches the spec as it
used to be computed: what it reads is recorded then, a change is caught
from that record from the next apply on, and the label is renewed at the
next genuine recreate. Recreating them all would have been a restart storm
across the mesh in declaration order, the store first. The trade-off is
stated in the code: a container already stale at upgrade time is not
caught, and could not have been either way.
The record of what a container read is looked up by its name when its
declared id has none — the bundle's `store` becomes `postgres.server` for
the same container — so a change on the day it is adopted still names the
file. The by-target lookup takes the most recently applied record, since
the bundle's record for the same target is never removed by the mesh's.
novox/hq 04-ISSUES/103
Review of the fix for hq issue 104 found three faults in it. A file applied
on an enrolled node — the mesh's own last declaration included — is applied
as the bundle is, so its resources are recorded as the machine's own and
what the mesh declared reads as undeclared: the plan removed the foundation.
`apply FILE` is for a machine the mesh has not spoken to, and is now refused
saying so whenever declared.json exists. The plan looked at what is held
before what the declaration says is taken, so the one cutover ADR 0100 says
must be previewed read as a hold; it now decides in holdOnAdopted's order,
models a step run inside a held container, and a test holds the plan's
sequence to the apply's outcomes. Genesis wrote the mode on every run, so a
re-run after `converge` left the state saying adopted while the kept,
signed declaration said converged, and the reconcile loop refused every five
minutes with no delivery coming to end it: genesis now writes the mode only
when none is recorded, and where the state and the verified kept declaration
disagree, the kept declaration wins and the repair is said.
Also: a file lock beside the state, taken by the link service, the host's
own commands and the installer alike, so a `reconcile` run by hand no
longer races the loop's save — chosen over refusing while a named service is
active, which would miss a `mesh-host run` started by hand; `--json
--dry-run` emits {plan} like an apply emits {plan, report}; the README's
duplicate flag line; and the bundle refusal is about the digest, not a claim
the carried bytes can never match what genesis applied.
An operator ran `mesh-host reconcile` on an adopted control-node with twelve
modules assigned. It applied the bundle the host carries — the genesis
declaration, foundation only, converged: recreated the store, failed on the
broker's held port, wrote the converged base filter and started its service,
and stopped at the first failing action. The filter closed the machine for
forty-five minutes. The host reported the node adopted in every report, the
declaration said converged, and nothing compared the two; nothing was printed
before acting (hq issue 104).
The host now records the node's mode — from every declaration the mesh sends,
and at genesis from what the operator said — and refuses, at the point of
application, a declaration that says the other mode, naming both and the act
that changes it. Only a declaration the link delivers, signed, changes the
mode: that is how `converge` and `adopt` arrive, so the flip still works and
nothing else can do it. Genesis marks the bundle consumed, with the digest of
what it applied, so `reconcile` holds a node the mesh has spoken to against
what the mesh last said and never the bundle, and refuses the carried bytes
when they are not what genesis applied. A file is refused when it is not what
the mesh last said: a declaration carries no sequence and no issued-at, so the
host cannot tell older from newer, and says so. Both commands print what they
would change — a hold, a removal, an action named as one — before touching
anything, and --dry-run is that list and nothing more.
Registering a module is an overwrite. Recording the forge's port ran `module add`
every time, so a genesis re-run pointed at an older catalogue would replace the
manifest of a forge that is built and assigned — with a push a few lines later.
Registering is only here so a settings row has a module row to hang on, and that
row is already there on a mesh that knows the forge. So: ask first, and skip.
Two comments narrowed to what is true. What follows the node's setting is what
the mesh derives from a module's ports — its container mapping, its filter rule,
its opening and what it serves. The forge's own address in its runtime's
environment (hq 088) and its route contribution's port do not, and are already
wrong for any port the mesh assigned. And a settings layer is the module's, not
one resource's: a second mergeable file on the builder would be given `serves`
too.
novox/hq 04-ISSUES/085
Every foundation port given at genesis became a per-node setting of the module
that binds it, except the package registry's: that one was fixed by rewriting
the builder's manifest when the installer registered it. Registering the builder
again from the catalogue undid it, and the forge's own module, when it took the
bootstrap forge over, came up on the catalogue's port — which on a machine where
a predecessor holds 3000 points the builder at the predecessor's forge.
So the rewrite is gone, and the port is recorded twice as a setting, both from
the one input:
- the forge's module is registered at genesis — not assigned, nothing of it runs
— so the controller has something to hold `{"ports": {"3000": <given>}}`
against. Assigning the forge later raises it on the port this machine was
given, and its container, its filter rule, its opening, what it serves and
what consumers are told all read it from there.
- the builder is given `{"serves": {"port": <given>}}`, which merges into the
binding it carries in place of one nothing can resolve yet.
A genesis on the catalogue's port records nothing and registers nothing, so it
does exactly what it did before.
novox/hq 04-ISSUES/085, ADR 0100
The control plane's manifest existed twice: at the root of its repository, read
whenever the mesh rebuilds it from source, and as a copy in the catalogue, read by
genesis. Nothing kept them equal, and the first rebuild replaced the mesh's record
with the repository's shape while every later push was refused (novox/hq
04-ISSUES/072). The builder's one-shot result already carries the manifest it built,
artifact resolved to the image; step 3 keeps it and step 9 registers it, re-pinning
the built image's bare id to the reference the registry assigned. The catalogue is
still read for the registry's and the builder's manifests and for phase two.
A container on the machine dialling a port the machine publishes reaches it
through the runtime's proxy — input, not forward — and the builder could not
reach the broker. The derived ruleset opens the mesh's own ports in both
chains; the base one now does the same.
031: a window of unacknowledged declarations is drained to the newest; the
rest are set aside and reported as superseded. 035: a file resource may say
create-once — written when absent, kept untouched when present (ADR 0087).
054: the bundle installs nftables and loads a base ruleset before the store
and broker, in the table the filter module later replaces (ADR 0088).
The verify reads the marker with the shell's read, which fails at end of file
without a line ending; the action ran and its verify said no. And the applier
reports each action with its command line, two of which now carry the real
store and broker passwords — the installer masks the values it made in
everything it says.
From review: the store and broker passwords genesis makes were carried into
the controller through a world-readable file in /tmp, a bundle left at 0644 by
an earlier installer kept that mode while now holding them, a mesh raised by
the old installer would have been handed new passwords its servers do not have,
and the broker-admin action's marker did not depend on the value. Secrets now
stage in a 0700 directory owned by the controller's account; the bundle is
chmod'd; an existing store or broker volume with no credential file is refused
by name; the marker holds the password's fingerprint. Also: one install path
for the store, broker and vault, no error-string matching for the operator
key, and no unreachable fallback for the superuser.
The template raises the store with the password 'bootstrap' and the broker
with its image's default administrator, and the installer carried both into
the mesh as accepted secrets — permanent, and not secret (novox/hq issue 071).
Now the installer makes both credentials, once, at the paths the postgres and
lavinmq modules declare as their own secrets, rewrites the produced bundle to
use them (the store reads its password from a file; the broker's default
account is given the new password by an action before anything dials it), and
writes the bundle at 0600 since it now carries them.
Before the first secret is accepted it makes the operator's sealing key beside
the bundle and gives the mesh the public half, so everything minted from there
is sealed to it too (ADR 0085, amended). Phase three adopts the broker as the
lavinmq module beside the store and installs mesh-vault as a foundation module;
the run ends by writing the operator-sealed export beside the key.
Not "every module's database"; a module requests one via requires
postgres-database. The one server holds the controller's contexts and the
database of each module that asks for one.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
InstallStore turns the mesh-store the foundation raised at genesis into the
postgres module, adopted in place: it verifies the module's server names the
same container and the same image the foundation is running (fail-fast on a
drift, rather than tearing down the mesh's store), then registers, builds the
provisioner, and carries the superuser in via secret accept — the mesh cannot
invent a credential that already made the databases (mirroring the control
plane's store-connection delivery, control.go). pinImage generalised to any
module for reuse.
Issue 051 (WBS 3.1). One server holds the controller's contexts and every
module's database.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Fixes found raising the package registry end-to-end in the lab: seed gitea's DB
with plain psql statements (no \gexec, no $$ DO-blocks that clash with the
shell); run gitea on the host network so it reaches the substrate store and
answers where the builder looks; set gitea ROOT_URL to the machine's loopback so
npm's stored credential matches the tarball host; keep the pivot's passwords so a
re-run is the same run; create the admin without re-enabling must-change-password.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The base (mesh-tools) resolves the SDK by version from the mesh's package
registry rather than cloning it from a git URL (hq ADR 0076, issue 053), so the
registry has to answer and the SDK has to be in it before the base build runs.
New steps, before base: seed gitea's database in the substrate store, raise
gitea's server on it, create the admin/org/team and the builder's account, seal
the builder its registry credential, and publish the SDK on a public base. gitea
is adopted as an ordinary module after the base, so its provisioner image can be
built. A minimal Go gitea admin client stands in until that module exists.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Twelve steps made a mesh that RUNS and then said "what remains is somebody
else's". The seven things that turn it into a mesh that WORKS — the shared base,
a database provider, the catalogue, the private network, the packet filter — were
typed afterwards, which is how they went missing for weeks without anything
complaining.
Six more steps now: base, store, catalogue, network, filter, extras. Everything
in them is module add, build, assign and push — the same verbs a person types,
through the same commands, so the installer and an operator remain one act.
Where a human must choose, the installer asks. A choice resolves in the order a
person expects: the flag wins; a lone option answers itself ALOUD, because "it
chose for me" and "there was nothing to choose" read identically afterwards
unless one speaks; a terminal is asked; a default fills in; and a required
choice nothing answered refuses naming its flag — a guessed packet filter is a
machine somebody else configured. The filter is required, so the question is
which, not whether. A run without a terminal (the lab, --json) is never left
waiting on a prompt nobody will answer.
Placement is part of the network step, not a separate act — a lesson paid for:
the module installed, the names file was written with no names in it, and
everything reported success because nobody had said where the machine IS. The
hub endpoint derives from the broker address when unsaid: the host other
machines dial is one fact, not two that drift.
Extras fail the run rather than soft-fail: somebody asked for them by name, and
a mesh reporting success minus one thing is reporting the wrong thing.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The module that provides artifact-store runs Distribution, the OCI reference
implementation. It was called registry, which named neither the software nor the
provision.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Genesis ended with a mesh that runs and cannot make anything: every module in
the catalogue names artifacts and nothing had built them, so the first thing
anybody had to do was install a builder by hand.
The installer already carries one — it is what built the control plane — so
this is the same two acts the control plane goes through, in the same order:
publish it, so the mesh names it by a digest its own registry assigned rather
than a local identity nothing else can fetch, then install it as an ordinary
module pinned to that. And then the part only it needs, a broker account, issued
before the push so it arrives with the declaration rather than after it.
Verified on a bare machine: the install ends with a builder running, and that
mesh then built the shared base images and a module on top of them with nobody
helping it.
The carried image is the builder now. The publish step still pushed it, so the
registry got a builder under the control plane's name and the mesh installed it
as the control plane — which presented as a control plane that started, printed
a builder's usage, exited cleanly, and did it again. Caught by the lab on the
first genesis run, at the step that waits for it to answer.
It carried the thing it was going to run; it now carries the thing that makes
it. One artifact either way — but a mesh raised this way holds a control plane
it built from a repository and a commit it can name, and can therefore build
again. A mesh handed a finished image could not, and had no way to find that
out until somebody needed it to.
A build step sits between load and bundle, because the bundle must name an
image and that image no longer arrives finished. Everything after it is
unchanged: a locally built image is named by the digest of its own
configuration, which is exactly what the carried one was named by.
Refused in preflight when nothing says what to build, so a run that cannot
finish says so before it has changed anything.
Enrolling IS the machine speaking to the mesh, so straight after it the mesh has
always heard from this node — and the step took that as proof an agent was
running and skipped starting one.
The cost is silent and total. Everything after is the control plane being told
things, and nothing it is told reaches a machine with no agent to collect it: the
registry push at step 7 was accepted, the module recorded, and no container ever
created. It surfaced three minutes later as 'the registry is not there at all',
one step from its cause and looking nothing like it.
Both halves are asked now. A process may be wedged and collect nothing, which is
why the mesh is asked at all; and the mesh may have heard once from a machine
running nothing, which is why the machine is asked too.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The first real run of mesh-bootstrap stopped in preflight, dialling 192.0.2.250:5000
for ninety seconds on a machine whose network was fine. That address is the registry
the lab used to raise; the substrate template still names the control plane by it,
and step 3 replaces that reference with the id of the image this installer carries.
Nothing ever pulls it.
So preflight excludes the control plane's resource by identity, rather than by the
happy accident of the template filling its slot with something that needs no registry.
Every other container's registry is still dialled, because those are somebody else's
images at somebody else's registry and a machine that cannot reach one fails inside a
pull, which says the wrong thing.
Also: `make bootstrap` takes BOOTSTRAP_OUT. The lab now builds the installer from
source before every raise, into a path it chooses, and a caller that could not say
where the output goes would have to copy it afterwards.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An image id does not survive `docker save` -> transfer -> `docker load`. The id is
the digest of the image's *configuration*, and a runtime rewrites that
configuration as it loads: a newer Docker saves in one format, an older one stores
it in another. Same layers, same program, different name. Measured on a live raise:
saved on the workstation sha256:b86bb81ca2f9691f24f4725f50962d1e49c98c5ffe211113241243d42d18ceea
loaded on the machine sha256:2dc219046c73702fc640317f0342a28ec962ef1e9ef547b2f02861c508ca78fb
`internal/image`.ID read the id out of the carried tar and its comment said that
was the id the runtime would assign. That is true on the machine the image was
built on and false on every machine it is carried to — which is every machine this
program exists for. The installer then either stopped at step 2 refusing the
runtime's answer, or would have written a bundle naming an image the machine does
not hold; and nothing serves an image named by the digest of its own configuration,
which is the whole point of naming one that way, so the apply would have died
inside a pull that cannot succeed. The lab hit this.
So the image is identified by its TAG, which is ordinary metadata the tar carries
through unchanged. The runtime is asked what that tag resolves to before the load
(already held, nothing to do) and again after (this is what the bundle names). The
tag never reaches the bundle — a pinned bundle may not rely on one, ADR 0006 — it
is how the id is obtained, not what is written down.
- image.ID becomes image.ArchiveID, and says plainly that it is a fact about the
file and not a prediction about any machine. It is kept for reports, and printed
beside the runtime's answer whenever the two differ.
- Idempotence is decided from what the runtime holds under the tag, not from a
predicted id, which cannot answer the question at all here.
- An untagged archive is refused, in preflight and again at the load: there would
be no portable name to ask about, and the only thing left is scraping a sentence
`docker load` writes for a person. `make bootstrap` refuses an id or an untagged
image, so it is caught in front of whoever can fix it.
- A dry run cannot know the id and says so rather than pretending. Run refuses to
write a bundle carrying an unconfirmed id at all.
Tests: the injected Runner now answers with an id DIFFERING from the tar's, and the
runtime's answer is what must be used. The test that refused a differing id encoded
the mistake and is replaced by one refusing an answer that is not an id at all.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The catalogue's mesh-control manifest landed while this was being written, and it
does what the ordinary case does: it keeps its secrets under /var/lib/mesh and
mounts them into the container at /run/secrets, so MESH_STORE_INVENTORY_FILE names
a path that no own-secret writes. Matching on the path alone found nothing and
would have refused a correct manifest.
So the lookup follows the volumes. It also reads the other shape the manifest uses
— `VAR=${secret:name}` inside the environment file a container reads — which is
how a value that is not a path gets in at all, and which is where the broker's two
credentials live.
That generalises what is delivered: every variable the module fills from a secret
is looked up in the substrate's control plane. What the substrate names is accepted
through `secret accept`; what it does not is left for the mesh to generate, and
said so. A store connection the substrate does not name stays an error — a control
plane that cannot open a context is not one.
Checked against the real manifest (mesh-catalog feat/control-plane-module): five
variables resolve, the placeholder pins in one place, and the container it waits
for is `mesh-control`.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Steps 6 to 10, which turn a substrate into a mesh that can maintain itself
(novox/hq ADR 0067).
6 enrol a node record, a token, `mesh-host enrol`, and the host agent
running. Proved by the mesh having HEARD from the node, not by a
process existing: a host that cannot reach the broker looks exactly
like a successful install until the first push applies nothing.
7 registry the module that gives this mesh an image store, registered from a
--catalog checkout, assigned and pushed. Its image is upstream and
never built (04-ISSUES/029) — a placeholder digest there is refused.
Verified by asking `/v2/`, because a container that is up is not a
registry that serves.
8 publish the carried image pushed into that registry, which assigns it the
first manifest digest it has ever had. This is the hinge: without
it the mesh works and can never upgrade itself.
9 control the control plane registered as an ordinary module pinned to that
digest, with the substrate's own store connections delivered
through `secret accept` — read out of the bundle that made them,
because the mesh cannot invent a credential that predates it.
10 retire the temporary control plane dropped from the bundle and removed by
the host's ordinary removal pass.
Every step asks before it acts and reports "already done". No step leaves the
machine without a control plane: steps 9 and 10 overlap deliberately, and two
stateless control planes are untidy rather than broken.
mesh-control's `internal/builder`.PublishImage is mirrored rather than imported —
tier 0 depends on nothing that must be installed first — with one correction: the
digest is chosen from RepoDigests by repository instead of taken as element zero,
so an image pushed to two registries cannot silently pin this mesh to the wrong
one.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The substrate raises a control plane and a module will later declare one. If both
are called `mesh-control` then for one moment two owners hold one container, and
the host — which tracks what it owns — has no way to stop owning something without
destroying it. That looked like a missing mechanism.
It is a naming problem. The substrate's container becomes `temp-mesh-control` and
the module's keeps the plain name: two containers, two owners, nothing to hand
over. Dropping the temporary one from the bundle at the end is then destruction by
omission, which is what the host already does to anything that leaves a
declaration — and the right end for something named "temp" (novox/hq ADR 0067).
The rename is textual and matches the QUOTED name, so the `mesh-control` inside
the image reference is not caught by it. Read back afterwards: the produced bundle
must call it the temporary name, and no other container may have been renamed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF