Files
hq/01-RESEARCH/012-the-minimum-viable-node/00-overview.md
T
jschoubben bcb18c7329 012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's
disagree, the mesh's version is installed, the conflict is flagged, and it is
reconciled afterwards — the mesh's configuration is known to work, the machine's
is not, and a half-adopted machine is a state nobody understands.

So adoption always completes and flags inform rather than block, which also
settles what 'adopted with open questions' prevents: nothing. The node is a
node. The original is kept, so nothing is unrecoverable.

One class left open rather than folded in, because it is the one place the
oldest rule in this record argues the other way. 'Known to work' is true of the
mesh's configuration in isolation, not on this machine. Most disagreements are
preference and overwriting them is right. A few are tied to what is physically
present — a storage driver against the filesystem it is actually on, a data
directory pointing at a mount that exists — and installing ours there does not
discard a preference, it can make existing data unreadable. Restoring the
configuration file afterwards does not undo that.

The default is settled. The exception is not 'there is a conflict' but 'applying
ours would destroy something a configuration backup cannot restore', and
identifying that class is open.
2026-08-26 21:37:04 +02:00

10 KiB

status, initiated, touches
status initiated touches
active 2026-08-26
02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
02-DECISIONS/0004-managed-files-are-generated-never-edited.md
03-DESIGN/01-to-be/05-the-node-host.md
03-DESIGN/00-as-is/05-runtime-and-installation.md
01-RESEARCH/011-the-module-graph/00-overview.md

012 — The minimum viable node, and adopting what is already there

What is being investigated

Two questions that turn out to be one:

What is the bare minimum to run a one-node mesh? Not the tiers as asserted, but the actual closure — take the thing that must run, walk what it needs, and the set that comes back is the answer.

And how does a machine that is already in use become that? A candidate node is not empty. It has a package manager, probably a container runtime, possibly a git installation, each with configuration somebody chose. The mesh must own those, and owning is not the same as finding them present.

Why

Building tier 0 reached a wall that looked like a packaging problem and is not.

The host can be told to run a container or install a package. Both need a file — an image, an archive — and the question was where the host gets it. That framing produced a bad trilemma: carry everything in the bundle, download at apply time, or have something push the files in first. Downloading fails on the first node, which cannot fetch the image registry from the image registry it is trying to start.

The reframing: the machine is not offline. What matters is when the fetching happens. Move it from apply time to build time — build the installer on a machine that has a network, tailored to the target, and apply it on a target that then needs nothing. That is the same move the lab already made for its router image, and the same property the delivery design already claims: what ships is self-contained and a deploy touches no network.

Which makes the interesting question not where do artifacts come from but what is missing from this particular machine, and that needs both of the questions above answered.

What tailoring implies

  • The binary stays generic; the payload is tailored. One static host per architecture. What is machine-specific is the bundle it applies. Rebuilding the host per machine would buy nothing.
  • Detection is the input, not a report. What the host already reports about a machine — its profile and inventory — is what the difference is computed against. This is the first use of stage 1 by something other than a person reading it.

Adoption

Simply having a package installed is not enough. If the mesh manages the container runtime, it decides that runtime's configuration; a runtime found already installed carries settings somebody chose, and those cannot be discovered by noticing the binary exists.

So a machine already in use is adopted: what is there is read, taken over, and thereafter generated.

This was the original path. 00-as-is/05 records adoption of a pre-existing machine's configuration as the original mechanism, since made legacy and explicitly out of scope for the lab. It returns here for a different reason than it was dropped for, which is a thing to notice rather than to gloss.

It creates a state that does not exist today. ADR 0004 has managed files generated and never edited; adoption needs a one-time import before that rule starts applying. Three states, and the middle one is new:

unmanaged → adopted once → generated

And it crosses a boundary just drawn. ADR 0043 says the host never touches what it did not create — the rule that stops a converger deleting what the mesh never put there. Adoption is the deliberate act of taking ownership of exactly that. The rule needs a companion rather than an exception: never, unless adoption made it the host's, with adoption being explicit, recorded, and visible in what the host says it owns.

Nothing is taken over without keeping what was there

Before adoption touches a file, the original is kept. Adoption happens on machines somebody is already using, and the configuration being taken over is configuration somebody chose. A one-way door on a working machine is not an installation, it is a risk nobody agreed to.

This is a never rule rather than a courtesy, and it earns that by the same incident the mesh's strongest rule already carries: the worst loss in this record came from a tool acting on a path it did not own. Adoption is that act, made deliberate — which makes the safeguard obligatory rather than optional.

What that requires, and what remains open: where the copy lives, whether it is recorded in what the node knows about itself so that adoption is visibly reversible, and whether the mesh keeps it forever or hands it back when it stops managing the thing.

Adoption produces a briefing, not just a result

Proposed by the operator, and it answers a question this effort had open with two bad answers.

Adoption meets things a script cannot decide. A container runtime configured with one storage driver and a mesh wanting another. A package pinned to a version somebody chose for a reason. Local settings the mesh has no opinion about and no business discarding. Silently winning is wrong in both directions; refusing outright makes a machine in use unadoptable.

So adoption has two outputs. What it did — mechanical, recorded, in the node's state. And a briefing: what it found, what it took over, and what it could not resolve, written to be read by a person or an agent, which is the first thing a session on that node has to work with.

Conflicts are flagged, not resolved. That is the same principle the declaration parser already applies — name every problem at once, to somebody who can act on it — at a larger scale, and applied to a case where refusing wholesale would be worse than proceeding.

On conflict, the mesh's version is installed

Decided by the operator: where the existing configuration and the mesh's disagree, the mesh's version is installed, the conflict is flagged, and it is reconciled afterwards.

The reasoning is that the mesh's configuration is known to work, the machine's is not, and a half-adopted machine is a state nobody understands. Adoption therefore always completes — flags inform, they do not block — and the original is kept, so nothing is unrecoverable.

That also answers what adopted with open questions prevents: nothing. The node is a node.

One class of conflict this should not cover, and it is the open part. Known to work is true of the mesh's configuration in isolation, not on this machine. Most disagreements are preference — tuning, logging, a mirror — and overwriting them is right. A few are tied to what is physically present: a container runtime's storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists. Installing the mesh's version there does not discard a preference; it can make existing data unreadable, and restoring the configuration file afterwards does not undo that.

So the default is settled and the exception is not: the exception is not there is a conflict but applying ours would destroy something a configuration backup cannot restore. Identifying that class is open, and it is the one place where this repository's oldest rule — the incident that came from a tool acting on a path it did not own — argues for refusing rather than proceeding.

Open questions

Question Why it is open
What is the closure for a one-node mesh? The skeleton asserts four pinned services. Research 006 already asks whether it is four or five and does not answer. A graph gives a computed answer instead of an asserted one, which is research 011.
Is "tier" the same thing as a graph level? Tiers were named as a bootstrap order. If the closure is computed, tiers may be a derived view of the graph rather than a separate concept — or they may be a coarser boundary that survives for a different reason.
What happens when existing configuration contradicts what the mesh needs? Answered — the mesh's version is installed, the conflict is flagged, and it is reconciled afterwards. The original is kept, so nothing is unrecoverable.
Do flags block, or only inform? Answered — they inform. Adoption always completes, and the node is a node.
Which conflicts are destructive rather than merely disagreements? The one exception to the rule above. Overwriting a preference is right; overwriting something tied to the machine's filesystem or existing data can make that data unreadable, and a configuration backup does not undo it. Identifying that class is open.
Where does the kept original live, and for how long? Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing.
What shape is a briefing? Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log.
Does owning a package mean owning its version? Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it.
How does a bundle stay true between building and applying? It is built against a scan of the target. The machine can move between the scan and the apply, so the host has to verify rather than assume — and fail plainly when the bundle no longer fits.
What cannot be precomputed at all? Anything built from source on the target still needs a toolchain and a network at that moment. Tailoring moves that cost rather than removing it, and minimal viable has to be honest about what it cannot ship ahead.
Does presence differencing understate the gap? Knowing a package manager is installed does not say it is the version the mesh needs. A difference computed on presence alone is optimistic.