f70973e222650e7706533aab18693fc1dbbf7d0b
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1111bd84d7 |
Establish the repo for the completed Phase 0-3 build
Settles the design repository now that the self-upgrade build is on main: - Records the two decisions that shipped without a record — ADR 0077 (the controller/foundation/node vocabulary) and ADR 0078 (the store and broker are ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on. - Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation. - Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs now that the forge repo is renamed; updates the glossary note and repos.md. - Fixes the six broken links from the design-doc renames, indexes the glossary, regenerates the decisions reading order. Both checks (records.py, index.py) are green. Statuses stay honest: the build is on main and lab-proven but not deployed as the production mesh, so the to-be docs remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation to implemented + as-is belongs to deployment, not merge. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
33a00d5656 |
Adopt the glossary's vocabulary in the mutable design docs
"control plane" -> controller and "substrate" -> foundation throughout 03-DESIGN, 00-META and the README, with 06-the-control-plane.md and 07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md. The immutable 02-DECISIONS records keep their original wording (and links to them are unchanged) — a term retired here may still appear there, which the glossary explains how to read. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
1b5308c9cc |
Review of the to-be layer: check what the documents claim against what runs
First pass of a design review, done by reading documents against code and against a raised mesh rather than against each other. Every error below was invisible to a proofread. **Statuses were stale, and nothing checked them.** Ten to-be documents said `designed` while naming working, lab-proven code — several with a *What was built* or *Raised, and observed* section. Added a `status-vs-code` check: naming a file is a claim that the file implements this, so a document that points at one has stopped being merely designed. It failed on all ten before it passed, per the rule this folder sets for its own checks. **The bundle carries three images, not two.** 07 reasoned about which substrate services go in and overlooked that the control plane is in there too — it is what the substrate exists to start, and there is nothing to fetch it with yet. Counted, not deduced. **The bootstrap uses four shapes, not six.** It listed `file` and `directory`, which substrate-first-node.lock never asks for. The claim that mattered — nothing is blocked on the host — was true either way, which is why the wrong count survived. **The eight capabilities were documented nowhere.** Implemented in internal/profile/detectors.go and enumerated in no document, including the one about the host that detects them. A vocabulary modules write against, readable only by reading the code. Now written down, with the seat/graphical-session distinction that is wrong in both directions if collapsed. **MinIO swept out of the to-be layer** per 0028. The gate now fails on one thing left deliberately: ADR 0024 is `proposed` while two documents rest on it and the feature it decides is built and lab-proven. Accepting a decision is not mine to do. |
||
|
|
573a94e102 |
Correct the record: a limitation that no longer exists, and one that was never written
The connectivity design still said a hub cannot be filtered — a gap recorded in the morning and closed in the afternoon, left standing as though it were current. Worse than a stale date: it would send somebody away from something that works. `restart-on` was described nowhere, including the part added today that lets a service reflect a file another module put on the machine. A rule the host enforces and no document mentions is a rule nobody can rely on. And nine of fifteen design documents claimed an `updated:` older than their last change, some by a week. That field is what cross-cutting views are generated from, so it is not decoration. |
||
|
|
35ce23fb76 |
Record that a machine coming back is ordinary, and what waking now does
Asked whether a machine that drops off needs re-adopting: it does not, nothing expires, and the only thing that forces re-enrolment is losing its own key. The gap was the twenty or thirty seconds after a resume in which a node believes it is in a mesh it has left — recovering on its own, which made it a quality gap rather than a fault, and still a machine waiting to be told something it already knew. |
||
|
|
7f314eb399 |
Point three design documents at the code that exists for them
Their subject matter has been built and proven for days and their frontmatter still said code: [] — which is what the cross-cutting view is generated from, so it was claiming nothing existed for the substrate, the node lifecycle and delivery. |
||
|
|
004057d85c |
A node's identity is a keypair it generates. This was never open.
I have been treating "what a node presents to prove it is that node" as an undecided design question for weeks, and blocking on it. It was decided. 08-connectivity says of the overlay keys: each node generates its own keypair, the private key never leaves the machine, the public key is published to the mesh -- and says explicitly that this IS ADR 0004's "a node holds its own identity", applied. Nobody had applied it to the thing 0004 is actually about. What caused it was a word. The lifecycle said a joining node receives its own durable identity, which reads as the mesh issuing something, and then the question is what. The mesh issues nothing. A node arrives holding its identity; what it receives is being known. That line now says what happens: it presents the one-time secret and its own public key, which the mesh records. The rule above it then holds literally rather than aspirationally. The mesh stores a public key, so a copy of the mesh's database grants nothing, and compromise of a node really is compromise of only that node. Also recorded, since it was asked directly: same principle as SSH, own key, not the machine's SSH host key. Host keys are regenerated by reinstalls and image clones, which would silently un-enrol a node; their lifecycle belongs to sshd rather than the mesh; and a partial host has no SSH daemon at all, so an identity scheme resting on one excludes a supported kind of node. The good half of that idea is kept: the mesh knows every node, so it can distribute host keys the way it distributes authorised keys, and node-to-node SSH stops depending on trust-on-first-use. |
||
|
|
82a3065f82 |
Tier 2 exists, and the token was missing a quarter of itself
mesh-control is built as far as it can honestly go: one context of seven, inventory, with its schema and the command that applies it. The repos map and the control plane design say so, and point at ADR 0024 for what it took. Separately, and more importantly: this repository described the enrolment token as carrying three things when ADR 0004 says four. The missing one is the control plane's signing identity -- the reason a node does not have to trust the broker it dials. Without it the control plane's authority is transitive through the broker, and 0004 spells out what that costs: a compromised broker could forge declarations, and since the host applies whatever the link delivers, that is the whole machine. The record has the argument in full; the design doc had dropped the conclusion. Found by reading the two together while deciding what the control plane must store, which is roughly the only way it would have been found -- both documents are internally consistent and only disagree with each other. |
||
|
|
333356cff3 |
Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception. |
||
|
|
e1febe8e0f |
Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18, 19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only the archaeology of what used to be there. Renumbered contiguously. Renames run in ascending order, so every target number is already free and no two files ever collide. The reference rewrite is one simultaneous pass rather than a sequence of replacements. Numbers moved into slots other numbers were vacating -- the node host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time would have cascaded and silently pointed things at the wrong record. Seven plain-text references survived the merges as prose rather than links, naming records that no longer existed: the enrolment token, the link boundary, what a declaration is, reachability, the repository structure. Each mapped to the consolidated record that now holds it. Verified rather than assumed: every [ADR NNNN](path) link now has matching text and target, checked across the whole repository, and the checker passes. Frontmatter `consolidates:` lists dropped -- they named records that are gone, and each consolidated record already says in prose what it absorbed. |
||
|
|
77f3a4cea7 |
Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
|
||
|
|
10365f2eae |
Consolidate the design layer: one place per topic
Jochen: a jungle of specs that slightly contradict or patch each other, and what matters is a working state rather than history. Both are fair and both are mine. Measured rather than assumed. 05-the-node-host and 09-the-node-lifecycle both covered enrolment, the install commands, the unit file, the launcher and reconcile -- I wrote 09 without taking anything out of 05, so the same things were said twice and could drift apart. Split by what each document IS. 05 is the component: what the host is, its parts, the declaration vocabulary, the build order, how it is verified. 09 is what happens to it: install, enrol, run, upgrade, retire. The whole "The process" section left 05, and the unit file moved to 09 where installing is described. 05 goes from 338 lines to 245 and now points at 09 rather than restating it. 09 also carried a 105-line "Resolved" section -- six mechanisms framed as "these were open and here is the answer". The content is needed; the framing is history, and history is what makes a document read as a changelog rather than a description. Renamed to what it actually is and the was-open phrasing removed. Also added 10-delivery.md, which did not exist: four accepted decisions -- 0054, 0063, 0064, 0065 -- had no design document at all, which is the specific reason the delivery picture felt scattered. It is now one document covering modules, the three edges, the core library, and how a change becomes a running thing, with a table of what each property is designed against and what must exist before it can be built. |
||
|
|
ba0d01788e |
0062: a host may be episodic; 0060's Android gap closed
0060 named the gap and did not close it: everywhere else an init runs the launcher at boot, and Android grants neither an init to register with nor anything worth supervising, because a supervisor would be killed alongside what it supervises. Closed by narrowing what is required rather than building something. A host is resident or episodic, and both are hosts. Being killed by the platform is disconnection, which 0036 already made ordinary -- and every mechanism an episodic host needs already exists because it was built for laptops that close. A partial host can join a mesh and cannot be the first node, since every bootstrap step is a shape it refuses. Its bundle says so. Two consequences that are easy to miss: last-heard-from means much less on an episodic host, so a healthy phone reads as a dead server unless the reader knows which kind it is; and a declaration may take a long time to land, which makes 0058's outstanding-versus-failed distinction load-bearing. Still open, and in that order: what an Android node is FOR, and only then how it is started. |
||
|
|
f1b1cd9aa0 |
Review: three ADRs no longer said what we had concluded
A sweep for claims overtaken by the last few days. Annotated rather than rewritten, following the pattern already in 0049 -- what changed and why is the useful part, and an accepted record should not quietly become something else. 0057's init section was wrong on all three of its claims. It said the host needs FOUR things from an init; 0061 reduced that to one. It said every machine the mesh targets already has systemd; Alpine does not, and it is the intended first node. It said there is no second init to abstract over; there is now, and the answer is still not an abstraction -- it is a four-line file per system. What survives is the part that was always right: an init is not a dependency in 0041's sense, because it is not installed, it is what the machine already is. 0048 named Docker as the container runtime. It is now docker or podman, detected rather than chosen -- because adoption keeps what a machine already has, so naming one contradicted a rule already decided. That row is the only one of the five that names two, and the record now says why. 0060 claimed the bundle is portable across operating systems. Its mechanism is; its contents are not -- package names, unit names, service names all differ, so an Arch host embeds an Arch bundle. That was my error, and it is the exact confusion behind the question that found it. The design layer had the same drift: 07 and 09 said "Docker" where they meant a container runtime, 09 said systemd restarts the host after an upgrade when the launcher does, and both install snippets assumed Arch. They now show Alpine and Arch side by side, which makes the point better than prose did -- step 1 differs per system, step 2 never does. Checked and NOT changed: 0047's "the vocabulary grows by one shape" is a claim about the rate, not the count, and is still true. 0037 lists docker among tools the host manages, which it does. 0041 says nothing about either. |
||
|
|
e1ad39b500 |
Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch host's implementation, not abstractions the mesh has to grow. They are not independent choices: a machine has pacman because it is Arch, and the package manager, service manager and packaging format arrive together as one decision somebody made at install time. Rejected abstracting them, and the reason is correctness rather than effort. The service applier reads LoadState to tell "not installed" apart from "stopped", which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both express, and the lowest common denominator is exactly where that fault lives. Almost all of it is shared -- the vocabulary, store, apply loop, read-back discipline, refusal model, bundle and link are portable. Two appliers differ. And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so this is the seam that already existed. Android is the interesting case rather than Debian: no service manager, no package installation, usually no root. Such a host implements file, directory and action and refuses the rest -- the same refusal a host already gives an unknown type, with a different reason. Those three are the portable floor. The container runtime is deliberately left open: it is not an OS split, since Arch runs docker or podman. 0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing else. Both are expressible in OpenRC, runit, s6 and an Android init.rc. Counting failed starts and rolling back moves into a launcher, because that is the one piece which must work when the host does not, and a script with a counter can be tested where OnFailure= can only be hoped for. Supersedes 0059, keeping its reasoning in full. The checker found all six places citing 0059 and refused the commit until they named the replacement. |
||
|
|
19997d56c3 |
Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance. |
||
|
|
605c9fd441 |
Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed. |
||
|
|
aeea2a9f9a |
Resolve the host lifecycle's open items, and say how the host is delivered
The upgrade question turned out to be a delivery question, so 0058 answers both. Today's third silo runs once per node and sends each one a command to install and start. That is where the as-is records a package install that 404ed from every mirror while the job went green, an image pull failure that did not fail the deploy, and a verify stage that was built and never scheduled because it was missing from a list. The shape underneath all of those is that the thing reporting success was not the thing doing the work. Meanwhile ADR 0037 has given every node a component that applies state, reads back and reports -- so two mechanisms now change a node and only one checks its work. 0058: a pipeline ends when the declaration is updated. Deploy stops sending commands to nodes and becomes one write. The host applies it on its next reconcile, and the host cannot report success it did not verify. The verify stage disappears as a stage, which is the point -- verification stops being a step that can be left off a list. A pipeline result now means "the declaration is updated, and here is which nodes have applied it". It does not wait for every node, because a node may be legitimately switched off for a week. Outstanding is reported separately from failed, since conflating them is how the old system produced a stall with no error anywhere. The host is delivered by exactly this path and needs no new resource type: a `file` writes the package manager's config pointing at the mesh's repository, a `package` names the version. Added a step I had missed -- before exiting for a restart, the host runs the new binary once. A package can install something that does not execute here, and that turns "the node never came back" into "the apply failed and said why". Six open items resolved: re-enrolment is decided when the token is issued and revokes the previous identity; the mesh keeps a recovery copy of what each node reports it owns, which un-strands the orphans; last-contact is reported with no threshold, because a laptop off for three weeks is doing nothing wrong; adoption always completes but a failed line makes a node ineligible for assignment; a briefing is a structured document whose outcome is computed from its lines; and the token is printed once and carried by hand, which is the property that makes it worth anything. Still open and named: automatic rollback of a host version that will not start. 0057 and 0058 are both proposed. |
||
|
|
2204b01909 |
Design the node lifecycle end to end
The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed. |