It carried the thing it was going to run; it now carries the thing that makes
it. One artifact either way — but a mesh raised this way holds a control plane
it built from a repository and a commit it can name, and can therefore build
again. A mesh handed a finished image could not, and had no way to find that
out until somebody needed it to.
A build step sits between load and bundle, because the bundle must name an
image and that image no longer arrives finished. Everything after it is
unchanged: a locally built image is named by the digest of its own
configuration, which is exactly what the carried one was named by.
Refused in preflight when nothing says what to build, so a run that cannot
finish says so before it has changed anything.
Enrolling IS the machine speaking to the mesh, so straight after it the mesh has
always heard from this node — and the step took that as proof an agent was
running and skipped starting one.
The cost is silent and total. Everything after is the control plane being told
things, and nothing it is told reaches a machine with no agent to collect it: the
registry push at step 7 was accepted, the module recorded, and no container ever
created. It surfaced three minutes later as 'the registry is not there at all',
one step from its cause and looking nothing like it.
Both halves are asked now. A process may be wedged and collect nothing, which is
why the mesh is asked at all; and the mesh may have heard once from a machine
running nothing, which is why the machine is asked too.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The first real run of mesh-bootstrap stopped in preflight, dialling 192.0.2.250:5000
for ninety seconds on a machine whose network was fine. That address is the registry
the lab used to raise; the substrate template still names the control plane by it,
and step 3 replaces that reference with the id of the image this installer carries.
Nothing ever pulls it.
So preflight excludes the control plane's resource by identity, rather than by the
happy accident of the template filling its slot with something that needs no registry.
Every other container's registry is still dialled, because those are somebody else's
images at somebody else's registry and a machine that cannot reach one fails inside a
pull, which says the wrong thing.
Also: `make bootstrap` takes BOOTSTRAP_OUT. The lab now builds the installer from
source before every raise, into a path it chooses, and a caller that could not say
where the output goes would have to copy it afterwards.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
An image id does not survive `docker save` -> transfer -> `docker load`. The id is
the digest of the image's *configuration*, and a runtime rewrites that
configuration as it loads: a newer Docker saves in one format, an older one stores
it in another. Same layers, same program, different name. Measured on a live raise:
saved on the workstation sha256:b86bb81ca2f9691f24f4725f50962d1e49c98c5ffe211113241243d42d18ceea
loaded on the machine sha256:2dc219046c73702fc640317f0342a28ec962ef1e9ef547b2f02861c508ca78fb
`internal/image`.ID read the id out of the carried tar and its comment said that
was the id the runtime would assign. That is true on the machine the image was
built on and false on every machine it is carried to — which is every machine this
program exists for. The installer then either stopped at step 2 refusing the
runtime's answer, or would have written a bundle naming an image the machine does
not hold; and nothing serves an image named by the digest of its own configuration,
which is the whole point of naming one that way, so the apply would have died
inside a pull that cannot succeed. The lab hit this.
So the image is identified by its TAG, which is ordinary metadata the tar carries
through unchanged. The runtime is asked what that tag resolves to before the load
(already held, nothing to do) and again after (this is what the bundle names). The
tag never reaches the bundle — a pinned bundle may not rely on one, ADR 0006 — it
is how the id is obtained, not what is written down.
- image.ID becomes image.ArchiveID, and says plainly that it is a fact about the
file and not a prediction about any machine. It is kept for reports, and printed
beside the runtime's answer whenever the two differ.
- Idempotence is decided from what the runtime holds under the tag, not from a
predicted id, which cannot answer the question at all here.
- An untagged archive is refused, in preflight and again at the load: there would
be no portable name to ask about, and the only thing left is scraping a sentence
`docker load` writes for a person. `make bootstrap` refuses an id or an untagged
image, so it is caught in front of whoever can fix it.
- A dry run cannot know the id and says so rather than pretending. Run refuses to
write a bundle carrying an unconfirmed id at all.
Tests: the injected Runner now answers with an id DIFFERING from the tar's, and the
runtime's answer is what must be used. The test that refused a differing id encoded
the mistake and is replaced by one refusing an answer that is not an id at all.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The catalogue's mesh-control manifest landed while this was being written, and it
does what the ordinary case does: it keeps its secrets under /var/lib/mesh and
mounts them into the container at /run/secrets, so MESH_STORE_INVENTORY_FILE names
a path that no own-secret writes. Matching on the path alone found nothing and
would have refused a correct manifest.
So the lookup follows the volumes. It also reads the other shape the manifest uses
— `VAR=${secret:name}` inside the environment file a container reads — which is
how a value that is not a path gets in at all, and which is where the broker's two
credentials live.
That generalises what is delivered: every variable the module fills from a secret
is looked up in the substrate's control plane. What the substrate names is accepted
through `secret accept`; what it does not is left for the mesh to generate, and
said so. A store connection the substrate does not name stays an error — a control
plane that cannot open a context is not one.
Checked against the real manifest (mesh-catalog feat/control-plane-module): five
variables resolve, the placeholder pins in one place, and the container it waits
for is `mesh-control`.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Steps 6 to 10, which turn a substrate into a mesh that can maintain itself
(novox/hq ADR 0067).
6 enrol a node record, a token, `mesh-host enrol`, and the host agent
running. Proved by the mesh having HEARD from the node, not by a
process existing: a host that cannot reach the broker looks exactly
like a successful install until the first push applies nothing.
7 registry the module that gives this mesh an image store, registered from a
--catalog checkout, assigned and pushed. Its image is upstream and
never built (04-ISSUES/029) — a placeholder digest there is refused.
Verified by asking `/v2/`, because a container that is up is not a
registry that serves.
8 publish the carried image pushed into that registry, which assigns it the
first manifest digest it has ever had. This is the hinge: without
it the mesh works and can never upgrade itself.
9 control the control plane registered as an ordinary module pinned to that
digest, with the substrate's own store connections delivered
through `secret accept` — read out of the bundle that made them,
because the mesh cannot invent a credential that predates it.
10 retire the temporary control plane dropped from the bundle and removed by
the host's ordinary removal pass.
Every step asks before it acts and reports "already done". No step leaves the
machine without a control plane: steps 9 and 10 overlap deliberately, and two
stateless control planes are untidy rather than broken.
mesh-control's `internal/builder`.PublishImage is mirrored rather than imported —
tier 0 depends on nothing that must be installed first — with one correction: the
digest is chosen from RepoDigests by repository instead of taken as element zero,
so an image pushed to two registries cannot silently pin this mesh to the wrong
one.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The substrate raises a control plane and a module will later declare one. If both
are called `mesh-control` then for one moment two owners hold one container, and
the host — which tracks what it owns — has no way to stop owning something without
destroying it. That looked like a missing mechanism.
It is a naming problem. The substrate's container becomes `temp-mesh-control` and
the module's keeps the plain name: two containers, two owners, nothing to hand
over. Dropping the temporary one from the bundle at the end is then destruction by
omission, which is what the host already does to anything that leaves a
declaration — and the right end for something named "temp" (novox/hq ADR 0067).
The rename is textual and matches the QUOTED name, so the `mesh-control` inside
the image reference is not caught by it. Read back afterwards: the produced bundle
must call it the temporary name, and no other container may have been renamed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The only complete written-down copy of how a mesh is stood up was an integration
test in the lab. That is why every bootstrap gap kept being found late: an install
procedure that lives as a test fixture is exercised by whoever writes tests, never
by whoever installs. This is that procedure.
A separate binary, not a mesh-host subcommand. mesh-host says of itself that it
connects to nothing and listens on nothing and that what it applies comes from a
file, and that sentence is what makes an always-running root daemon auditable. An
installer loads images and interrogates a control plane. Same tier, different
program.
The control plane's image is carried, not built and not fetched. The forge that
holds its source runs on the mesh, so a bootstrap that had to fetch it would need
a mesh in order to raise one. Embedding breaks that cycle the way the carried
bundle breaks "copy it onto a machine and run it". The image id is read out of the
saved tar before the runtime is asked anything, which is what makes the load
idempotent: the installer can ask whether the machine already holds exactly this.
Five steps, each idempotent and each saying whether it found or changed something,
because this is run over and over by somebody getting a machine working. It stops
at a running substrate with a control plane that replies — enrolment, the module
catalogue and assignment are the next stage and are deliberately absent.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF