77f3a4cea7514f89b1f11873a487d0e88b0abc34
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
77f3a4cea7 |
Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
|
||
|
|
5b3d0ebd4f |
ADR 0046 — the installer fetches what it pins
The blocking question was where a container image comes from, and the version that blocked assumed the machine might have no network. That assumption came from the LAB: a scenario is a closed address space by design, which is what lets two scenarios hold the same addresses without meeting. Production is not sealed — a machine being adopted has a network, and one that does not is a machine where very little works anyway. So substrate.lock carries references, not payload: an image name and a digest, fetched at apply time. A first node pulls from upstream because no mesh registry exists yet; every node after that pulls from the mesh's own. The lab is the exception and places images itself, the way it already places the host binary — a property of a test environment, and letting it dictate the production design would be the tail wagging the dog. Pinned by DIGEST rather than tag. Reproducibility comes from pinning the identity of a thing, not from carrying its bytes, which is what makes fetching acceptable rather than a compromise. ADR 0041 survives untouched, which was the point. "Copy it onto a machine and run it" stays literally true — one binary, a few megabytes, which then fetches what it was told to. Carrying images would have quietly redefined the property that decision rests on. Costs accepted and named: an apply can now fail because something is unreachable, which a self-contained artifact could not, so it must fail legibly — naming what it could not fetch and from where. And the lab needs a way to place images into a machine that also has no container runtime, both of which are lab-installation concerns and neither solved here. Research 012's build-time-versus-apply-time reframing narrows accordingly: it still holds for what a tailored installer contains, and no longer has to hold for images. |
||
|
|
278f7427ed |
012: the briefing carries an outcome, derived from its lines
Proposed by the operator: state plainly whether adoption succeeded, partly succeeded or failed, with a severity per line. Taken with one change — the overall is DERIVED as the worst mark present, never written alongside. Two fields maintained independently drift, and a briefing reading "full success" while carrying a failed line is exactly the fault this record keeps cataloguing. An outcome computed from its lines cannot disagree with them. Four marks: ok, kept, unknown, failed. "unknown" is not a shade of success — adoption will meet configuration it cannot parse and state it cannot read, and folding those into "fine" is the same move as reporting an installed package as a capability. And adding severity reopens something the earlier rule did not cover. "Flags inform, they do not block" was decided about CONFLICTS, where the mesh chose deliberately and the machine still works. A failure is not "we chose" but "we could not". Treating both the same makes a node where something the mesh needed never happened indistinguishable from one where a log level differed. |
||
|
|
0106318bcb |
012: on conflict, keep the machine's configuration
Reversed by the operator, and both directions are recorded because the reasoning for each is the useful part. What is already on the machine stays, the conflict is flagged, adoption completes. This buys non-destructiveness by construction: the class that made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — cannot arise, because nothing tied to the machine's physical reality is overwritten. It exposes the mirror. The mesh's configuration is not only preference; some of it is what a module needs to function. Keeping the machine's version there produces a module that is installed and does not work, which is 04-ISSUES/007 arriving from a direction that issue did not anticipate. And a fleet where every node kept its own settings is one where a module works on one node and fails on another with nothing able to say why. So neither direction is right as a blanket, and the question is not whose configuration wins. It is whether the module REQUIRES the setting or merely PREFERS it — required contradictions cannot be kept without breaking the module, preferences should always yield to what is there. That is a property of the module's declaration rather than of the adoption algorithm, which makes it one more thing the graph would carry. Until modules can say which of their settings are load-bearing, adoption is defaulting in the dark, and the default chosen is the one that does not break the machine it is adopting. |
||
|
|
bcb18c7329 |
012: on conflict, install the mesh's version
Decided by the operator. Where the existing configuration and the mesh's disagree, the mesh's version is installed, the conflict is flagged, and it is reconciled afterwards — the mesh's configuration is known to work, the machine's is not, and a half-adopted machine is a state nobody understands. So adoption always completes and flags inform rather than block, which also settles what 'adopted with open questions' prevents: nothing. The node is a node. The original is kept, so nothing is unrecoverable. One class left open rather than folded in, because it is the one place the oldest rule in this record argues the other way. 'Known to work' is true of the mesh's configuration in isolation, not on this machine. Most disagreements are preference and overwriting them is right. A few are tied to what is physically present — a storage driver against the filesystem it is actually on, a data directory pointing at a mount that exists — and installing ours there does not discard a preference, it can make existing data unreadable. Restoring the configuration file afterwards does not undo that. The default is settled. The exception is not 'there is a conflict' but 'applying ours would destroy something a configuration backup cannot restore', and identifying that class is open. |
||
|
|
60736199a4 |
012: keep the original, and flag what cannot be decided
Two additions from the operator, and the second answers a question this effort had open with two bad answers. Nothing is taken over without keeping what was there. Adoption happens on machines somebody is already using, and the configuration being taken over is configuration somebody chose. This is a never rule rather than a courtesy, and it earns that by the same incident the mesh's strongest rule carries: the worst loss in this record came from a tool acting on a path it did not own. Adoption is that act made deliberate, which makes the safeguard obligatory. And adoption produces a briefing, not just a result. It meets things a script cannot decide — a runtime configured one way against a mesh wanting another, a package pinned for a reason, local settings the mesh has no opinion about. Silently winning is wrong in both directions and refusing outright makes a machine in use unadoptable. So conflicts are FLAGGED: what it found, what it took over, what it could not resolve, written to be read by a person or an agent as the first thing a session on that node has to work with. That is the declaration parser's principle at a larger scale — name every problem at once, to somebody who can act on it. The question it turns on is recorded rather than assumed away: are flags advisory or blocking? A briefing nobody opens is worse than a failure, because the machine is in service and the record says it went well — 04-ISSUES/003 again. Working position: the node is usable and the mesh KNOWS it has unresolved adoption questions, as a state something can ask about rather than a document in a log directory. What that state prevents is undecided. |
||
|
|
ddb8091f68 |
Research 012 — the minimum viable node, and adopting what is already there
Building tier 0 reached a wall that looked like a packaging problem and is not. The host can be told to run a container or install a package; both need a file, and asking where the host gets it produced a bad trilemma — carry everything, download at apply time, or push the files in first. Downloading fails on the first node, which cannot fetch the image registry from the image registry it is trying to start. The reframing came from the operator: the machine is not offline, and what matters is WHEN the fetching happens. Move it from apply time to build time — build the installer on a machine with a network, tailored to the target, apply it on a target that then needs nothing. The same move the lab already made for its router image. Which makes the question not where artifacts come from but what is missing from THIS machine, and that needs two things answered: the closure for a one-node mesh, and how a machine already in use becomes one. Adoption is the second half, and it is sharper than it sounds. Having a package installed is not owning it: a container runtime found already present carries settings somebody chose, and noticing the binary exists discovers none of them. It was also the original path — 00-as-is/05 records adoption of a pre-existing machine's configuration as the original mechanism, since made legacy and explicitly out of scope for the lab. It returns for a different reason than it was dropped for. Two collisions recorded rather than discovered later. ADR 0004 has managed files generated and never edited, and adoption needs a one-time import before that rule starts applying — three states, and the middle one is new. And ADR 0043 says the host never touches what it did not create, which is exactly what adoption does; that rule needs a companion rather than an exception. Eight open questions, including whether 'tier' is just a coarse view of a graph level, whether owning a package means owning its version, and what cannot be precomputed at all — because tailoring moves the cost of building from source rather than removing it. |