Files
hq/01-RESEARCH/012-the-minimum-viable-node/00-overview.md
T
jschoubben 5b3d0ebd4f ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version
that blocked assumed the machine might have no network. That assumption came
from the LAB: a scenario is a closed address space by design, which is what lets
two scenarios hold the same addresses without meeting. Production is not sealed
— a machine being adopted has a network, and one that does not is a machine
where very little works anyway.

So substrate.lock carries references, not payload: an image name and a digest,
fetched at apply time. A first node pulls from upstream because no mesh registry
exists yet; every node after that pulls from the mesh's own. The lab is the
exception and places images itself, the way it already places the host binary —
a property of a test environment, and letting it dictate the production design
would be the tail wagging the dog.

Pinned by DIGEST rather than tag. Reproducibility comes from pinning the
identity of a thing, not from carrying its bytes, which is what makes fetching
acceptable rather than a compromise.

ADR 0041 survives untouched, which was the point. "Copy it onto a machine and
run it" stays literally true — one binary, a few megabytes, which then fetches
what it was told to. Carrying images would have quietly redefined the property
that decision rests on.

Costs accepted and named: an apply can now fail because something is
unreachable, which a self-contained artifact could not, so it must fail legibly
— naming what it could not fetch and from where. And the lab needs a way to
place images into a machine that also has no container runtime, both of which
are lab-installation concerns and neither solved here.

Research 012's build-time-versus-apply-time reframing narrows accordingly: it
still holds for what a tailored installer contains, and no longer has to hold
for images.
2026-08-26 23:52:46 +02:00

13 KiB

status, initiated, touches
status initiated touches
active 2026-08-26
02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
02-DECISIONS/0004-managed-files-are-generated-never-edited.md
03-DESIGN/01-to-be/05-the-node-host.md
03-DESIGN/00-as-is/05-runtime-and-installation.md
01-RESEARCH/011-the-module-graph/00-overview.md

012 — The minimum viable node, and adopting what is already there

What is being investigated

Two questions that turn out to be one:

What is the bare minimum to run a one-node mesh? Not the tiers as asserted, but the actual closure — take the thing that must run, walk what it needs, and the set that comes back is the answer.

And how does a machine that is already in use become that? A candidate node is not empty. It has a package manager, probably a container runtime, possibly a git installation, each with configuration somebody chose. The mesh must own those, and owning is not the same as finding them present.

Why

Building tier 0 reached a wall that looked like a packaging problem and is not.

The host can be told to run a container or install a package. Both need a file — an image, an archive — and the question was where the host gets it. That framing produced a bad trilemma: carry everything in the bundle, download at apply time, or have something push the files in first. Downloading fails on the first node, which cannot fetch the image registry from the image registry it is trying to start.

Qualified by ADR 0046. The reframing below still holds for what a tailored installer contains — the missing pieces for a given machine. It does not have to hold for container images: the installer fetches those by digest, because a real machine has a network and the sealed case is the lab.

The reframing: the machine is not offline. What matters is when the fetching happens. Move it from apply time to build time — build the installer on a machine that has a network, tailored to the target, and apply it on a target that then needs nothing. That is the same move the lab already made for its router image, and the same property the delivery design already claims: what ships is self-contained and a deploy touches no network.

Which makes the interesting question not where do artifacts come from but what is missing from this particular machine, and that needs both of the questions above answered.

What tailoring implies

  • The binary stays generic; the payload is tailored. One static host per architecture. What is machine-specific is the bundle it applies. Rebuilding the host per machine would buy nothing.
  • Detection is the input, not a report. What the host already reports about a machine — its profile and inventory — is what the difference is computed against. This is the first use of stage 1 by something other than a person reading it.

Adoption

Simply having a package installed is not enough. If the mesh manages the container runtime, it decides that runtime's configuration; a runtime found already installed carries settings somebody chose, and those cannot be discovered by noticing the binary exists.

So a machine already in use is adopted: what is there is read, taken over, and thereafter generated.

This was the original path. 00-as-is/05 records adoption of a pre-existing machine's configuration as the original mechanism, since made legacy and explicitly out of scope for the lab. It returns here for a different reason than it was dropped for, which is a thing to notice rather than to gloss.

It creates a state that does not exist today. ADR 0004 has managed files generated and never edited; adoption needs a one-time import before that rule starts applying. Three states, and the middle one is new:

unmanaged → adopted once → generated

And it crosses a boundary just drawn. ADR 0043 says the host never touches what it did not create — the rule that stops a converger deleting what the mesh never put there. Adoption is the deliberate act of taking ownership of exactly that. The rule needs a companion rather than an exception: never, unless adoption made it the host's, with adoption being explicit, recorded, and visible in what the host says it owns.

Nothing is taken over without keeping what was there

Before adoption touches a file, the original is kept. Adoption happens on machines somebody is already using, and the configuration being taken over is configuration somebody chose. A one-way door on a working machine is not an installation, it is a risk nobody agreed to.

This is a never rule rather than a courtesy, and it earns that by the same incident the mesh's strongest rule already carries: the worst loss in this record came from a tool acting on a path it did not own. Adoption is that act, made deliberate — which makes the safeguard obligatory rather than optional.

What that requires, and what remains open: where the copy lives, whether it is recorded in what the node knows about itself so that adoption is visibly reversible, and whether the mesh keeps it forever or hands it back when it stops managing the thing.

Adoption produces a briefing, not just a result

Proposed by the operator, and it answers a question this effort had open with two bad answers.

Adoption meets things a script cannot decide. A container runtime configured with one storage driver and a mesh wanting another. A package pinned to a version somebody chose for a reason. Local settings the mesh has no opinion about and no business discarding. Silently winning is wrong in both directions; refusing outright makes a machine in use unadoptable.

So adoption has two outputs. What it did — mechanical, recorded, in the node's state. And a briefing: what it found, what it took over, and what it could not resolve, written to be read by a person or an agent, which is the first thing a session on that node has to work with.

Conflicts are flagged, not resolved. That is the same principle the declaration parser already applies — name every problem at once, to somebody who can act on it — at a larger scale, and applied to a case where refusing wholesale would be worse than proceeding.

On conflict, the machine's configuration is kept

Decided by the operator, after first deciding the opposite — recorded that way because the reasoning for each direction is the useful part.

Where the existing configuration and the mesh's disagree, what is already on the machine stays, the conflict is flagged, and it is reconciled afterwards. Adoption always completes — flags inform, they do not block — and adopted with open questions prevents nothing. The node is a node.

What this buys. Adoption becomes non-destructive by construction. The class of conflict that made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — cannot arise, because nothing tied to the machine's physical reality is ever overwritten. A machine in use keeps working exactly as it did.

What it exposes, which is the mirror of what it fixes. The mesh's configuration is not only preference. Some of it is what a module needs in order to function at all. Keeping the machine's version there produces a module that is installed and does not work — an installed package is not a capability (04-ISSUES/007) arriving from a direction that issue did not anticipate. And a fleet where every node kept its own settings is a fleet where a module works on one node and fails on another with nothing in the mesh able to say why.

The distinction that dissolves both rules. Neither direction is right as a blanket, because the question is not whose configuration wins. It is whether the module requires the setting or merely prefers it — required contradictions cannot be kept without breaking the module, and preferences should always yield to what is already there.

That is a property of the module's own declaration rather than of the adoption algorithm, which makes it one more thing the graph would carry (research 011). Until modules can say which of their settings are load-bearing, adoption is choosing a default in the dark, and the default chosen here is the one that does not break the machine it is adopting.

The briefing carries an outcome, and the outcome is derived

Proposed by the operator: the report states plainly whether adoption succeeded, partly succeeded or failed, and each line carries its own severity.

Each line is marked, and the overall is the worst mark present. Derived rather than stated alongside, because two fields written independently drift — and a briefing reading full success while carrying a failed line is exactly the fault this record keeps cataloguing. An outcome computed from its lines cannot disagree with them.

Mark Means
ok done, and verified
kept a disagreement; the machine's value was kept and somebody should look
unknown could not be determined
failed could not be done — the node is not what was asked for

unknown is not a shade of success. Adoption will meet configuration it cannot parse and state it cannot read, and folding those into fine is the same move as reporting an installed package as a capability. A thing nobody could determine is a thing nobody can rely on, and it gets its own mark for the same reason a capability detector reports why.

And this opens something the earlier rule did not cover. Flags inform, they do not block was decided about conflicts — where the mesh chose, deliberately, and the machine still works. A failure is different in kind: not we chose but we could not. Treating both the same makes a node where something the mesh needed never happened indistinguishable from one where a log level differed. Whether a failed line still lets adoption complete is therefore reopened by adding severity, and is not decided here.

Open questions

Question Why it is open
What is the closure for a one-node mesh? The skeleton asserts four pinned services. Research 006 already asks whether it is four or five and does not answer. A graph gives a computed answer instead of an asserted one, which is research 011.
Is "tier" the same thing as a graph level? Tiers were named as a bootstrap order. If the closure is computed, tiers may be a derived view of the graph rather than a separate concept — or they may be a coarser boundary that survives for a different reason.
What happens when existing configuration contradicts what the mesh needs? Answered — the machine's configuration is kept, the conflict is flagged, and it is reconciled afterwards.
Do flags block, or only inform? Answered — they inform. Adoption always completes, and the node is a node.
Can a module say which of its settings are load-bearing? The question that dissolves the conflict rule rather than choosing a side. A setting the module requires cannot be kept from the machine without producing something installed and broken; a setting it merely prefers should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph.
Does a failed line still let adoption complete? Flags inform, they do not block was decided about conflicts, where the mesh chose and the machine works. A failure is we could not, which is different in kind — and treating them alike hides the worse one behind the commoner one.
How is a flagged conflict reconciled, and by whom? The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided.
Where does the kept original live, and for how long? Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing.
What shape is a briefing? Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log.
Does owning a package mean owning its version? Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it.
How does a bundle stay true between building and applying? It is built against a scan of the target. The machine can move between the scan and the apply, so the host has to verify rather than assume — and fail plainly when the bundle no longer fits.
What cannot be precomputed at all? Anything built from source on the target still needs a toolchain and a network at that moment. Tailoring moves that cost rather than removing it, and minimal viable has to be honest about what it cannot ship ahead.
Does presence differencing understate the gap? Knowing a package manager is installed does not say it is the version the mesh needs. A difference computed on presence alone is optimistic.