b4f3eaf11b4829bfff1aa31538909a387a963cad
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
121367319d |
Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo becomes mesh-controller, the seat the-controller, and the store+broker pair the foundation (embedded base bundles, default template and example lock renamed with their go:embed directives). No behaviour change — a pure vocabulary rename. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx |
||
|
|
e1a2fe7323 |
The installer carries a builder and builds the control plane it raises
It carried the thing it was going to run; it now carries the thing that makes it. One artifact either way — but a mesh raised this way holds a control plane it built from a repository and a commit it can name, and can therefore build again. A mesh handed a finished image could not, and had no way to find that out until somebody needed it to. A build step sits between load and bundle, because the bundle must name an image and that image no longer arrives finished. Everything after it is unchanged: a locally built image is named by the digest of its own configuration, which is exactly what the carried one was named by. Refused in preflight when nothing says what to build, so a run that cannot finish says so before it has changed anything. |
||
|
|
7e3481f025 |
bootstrap: the slot the installer fills is not a registry to reach for
The first real run of mesh-bootstrap stopped in preflight, dialling 192.0.2.250:5000 for ninety seconds on a machine whose network was fine. That address is the registry the lab used to raise; the substrate template still names the control plane by it, and step 3 replaces that reference with the id of the image this installer carries. Nothing ever pulls it. So preflight excludes the control plane's resource by identity, rather than by the happy accident of the template filling its slot with something that needs no registry. Every other container's registry is still dialled, because those are somebody else's images at somebody else's registry and a machine that cannot reach one fails inside a pull, which says the wrong thing. Also: `make bootstrap` takes BOOTSTRAP_OUT. The lab now builds the installer from source before every raise, into a path it chooses, and a caller that could not say where the output goes would have to copy it afterwards. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
cb5e137297 |
bootstrap: read the image id back from the runtime, never predict it
An image id does not survive `docker save` -> transfer -> `docker load`. The id is
the digest of the image's *configuration*, and a runtime rewrites that
configuration as it loads: a newer Docker saves in one format, an older one stores
it in another. Same layers, same program, different name. Measured on a live raise:
saved on the workstation sha256:b86bb81ca2f9691f24f4725f50962d1e49c98c5ffe211113241243d42d18ceea
loaded on the machine sha256:2dc219046c73702fc640317f0342a28ec962ef1e9ef547b2f02861c508ca78fb
`internal/image`.ID read the id out of the carried tar and its comment said that
was the id the runtime would assign. That is true on the machine the image was
built on and false on every machine it is carried to — which is every machine this
program exists for. The installer then either stopped at step 2 refusing the
runtime's answer, or would have written a bundle naming an image the machine does
not hold; and nothing serves an image named by the digest of its own configuration,
which is the whole point of naming one that way, so the apply would have died
inside a pull that cannot succeed. The lab hit this.
So the image is identified by its TAG, which is ordinary metadata the tar carries
through unchanged. The runtime is asked what that tag resolves to before the load
(already held, nothing to do) and again after (this is what the bundle names). The
tag never reaches the bundle — a pinned bundle may not rely on one, ADR 0006 — it
is how the id is obtained, not what is written down.
- image.ID becomes image.ArchiveID, and says plainly that it is a fact about the
file and not a prediction about any machine. It is kept for reports, and printed
beside the runtime's answer whenever the two differ.
- Idempotence is decided from what the runtime holds under the tag, not from a
predicted id, which cannot answer the question at all here.
- An untagged archive is refused, in preflight and again at the load: there would
be no portable name to ask about, and the only thing left is scraping a sentence
`docker load` writes for a person. `make bootstrap` refuses an id or an untagged
image, so it is caught in front of whoever can fix it.
- A dry run cannot know the id and says so rather than pretending. Run refuses to
write a bundle carrying an unconfirmed id at all.
Tests: the injected Runner now answers with an id DIFFERING from the tar's, and the
runtime's answer is what must be used. The test that refused a differing id encoded
the mistake and is replaced by one refusing an answer that is not an id at all.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
|
||
|
|
b82ab95f74 |
mesh-bootstrap: the first-node procedure, as a program rather than a test
The only complete written-down copy of how a mesh is stood up was an integration test in the lab. That is why every bootstrap gap kept being found late: an install procedure that lives as a test fixture is exercised by whoever writes tests, never by whoever installs. This is that procedure. A separate binary, not a mesh-host subcommand. mesh-host says of itself that it connects to nothing and listens on nothing and that what it applies comes from a file, and that sentence is what makes an always-running root daemon auditable. An installer loads images and interrogates a control plane. Same tier, different program. The control plane's image is carried, not built and not fetched. The forge that holds its source runs on the mesh, so a bootstrap that had to fetch it would need a mesh in order to raise one. Embedding breaks that cycle the way the carried bundle breaks "copy it onto a machine and run it". The image id is read out of the saved tar before the runtime is asked anything, which is what makes the load idempotent: the installer can ask whether the machine already holds exactly this. Five steps, each idempotent and each saying whether it found or changed something, because this is run over and over by somebody getting a machine working. It stops at a running substrate with a control plane that replies — enrolment, the module catalogue and assignment are the next stage and are deliberately absent. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF |
||
|
|
8fcfa88fe0 |
A machine that wakes or moves says so, instead of waiting to be told
A suspended laptop's connection is dead the moment it wakes, and the socket looks perfectly healthy from inside the process — no error, no close, because nothing has tried to send anything. Heartbeats find out twenty or thirty seconds later. For that time the node believes it is in a mesh it has left, which is the one state this design says must never be indistinguishable from being connected. The machine knew immediately. So being roused ends the current attempt rather than only shortening the wait after it: shortening the wait would do nothing at all, because the process is not waiting — it is sitting inside a connection that will not return. A signal, because nothing may listen on a node (novox/hq ADR 0004). A socket for this would be a control surface on every machine, reachable by anything that can reach the machine, in exchange for saving twenty seconds — and the whole security argument rests on there not being one. Two rouses in the same instant are one: a machine suspending and resuming repeatedly must not build a backlog of reconnections to work through. And the backoff is not reset by being roused — that says the machine changed, not that whatever was refusing the connection has stopped, and a laptop woken on a network with no route would otherwise retry at full speed for as long as somebody keeps opening the lid. The dispatcher acts on the events that change where packets go and not on `down`: the link is already gone there, reconnecting will fail, and the backoff exists for exactly that. |
||
|
|
e09503acc7 |
A real bundle: a bare machine raises a store and a database
The first three steps of the substrate bootstrap, run on a lab machine confirmed to have no route out. It went from bare to a container runtime installed and enabled, PostgreSQL running from an image pinned by digest, and the control plane's database created inside it -- from the file the host carries, with nothing to ask. Second run changed nothing. `owned` lists all five afterwards, and `mesh` is in the store. It stops before the last two steps because there is no control plane yet: its schema cannot be loaded and its image does not exist. The bundle says so rather than naming something that cannot be applied. Two things fixed on the way. `make host BUNDLE=...` still swapped a file called substrate.lock, which the per-system split had renamed months of decisions ago -- it now takes SYSTEM and replaces that system's bundle. And the .lock files still cited ADR 0060, since the renumbering pass only covered .md, .go, .ts and .sh. One thing learned by it failing first: a directory the host creates is owned by root, and a database inside a container runs as somebody else, so it could not write and the container crash-looped. The store's data is a named volume now, which lets the image set up its own ownership and outlives the container -- which is what you want for the thing holding the mesh's state. Worth noting the failure was caught by the action's verify rather than by the container step. `docker inspect` reported the container running because it was, briefly, between restarts. Running is not working, and the thing that knew the difference was the step that asked the database whether it would answer. |
||
|
|
02f1fcc865 |
Three hosts: arch, alpine and android
ADR 0060, built. `make hosts` produces mesh-host-arch, mesh-host-alpine and mesh-host-android, each pinned to its system at link time. The claim that "almost all of it is shared" held up. All 36 existing apply tests pass unchanged -- the only edit was naming which system they run against, which was previously implicit. What moved into internal/system is two appliers' worth of code and the probes that go with them. Each system's differences are real and needed re-deriving rather than translating: apk reports absence by EMPTY OUTPUT and exits zero either way, where pacman exits non-zero. Reading apk's exit code the way pacman's is read reports every package as installed. That is the single most dangerous difference between the two and it is invisible until it bites. OpenRC has no LoadState, so "the service does not exist" is read from its prose rather than a field. Same distinction, different evidence -- and this is exactly what an interface spanning both would have had to drop, which is why 0060 rejected one. OpenRC has no is-enabled either. Boot state comes from the runlevel listing: "does it start at boot" becomes "does it appear in rc-update show default". Android is a partial host and that is the point. It implements file, directory and action -- the shapes needing only a filesystem and a way to run something -- and refuses the other three by name, before anything is applied. Its unreachable appliers return ErrUnsupported rather than a zero value, so "unreachable" fails loudly if it stops being true. A host also confirms it is on the machine it was built for, once, at the start. The alpine host on this Arch machine says "this machine is not Alpine" instead of failing later inside a package manager that is not there. And a host built without -X main.builtFor refuses everything, naming the hosts that exist. Two test problems found by injecting faults. One injection did not compile, so the check now reports that separately from a pass. The other passed with the behaviour removed: the missing-service assertion matched "does not exist", which the FALL-THROUGH error also contains because it echoes the raw output. It now asserts the diagnosis, which only the correct branch produces. Verified with the real binaries: android refuses a package naming what it does support; alpine on Arch refuses the machine; arch applies and is idempotent; a system-less build refuses everything. |
||
|
|
057f34f924 |
The init is asked for start and restart; a launcher does the rest
ADR 0061. Recovery was the most systemd-specific part of the host, and it is the part that must work on a machine where nothing else does -- which made unit-file syntax a poor place for it, because syntax cannot be tested and the one time it runs is the one time nobody can afford it wrong. So StartLimitBurst and OnFailure move into a launcher script that init starts instead of the host. The unit drops to start-at-boot and restart-on-exit, which OpenRC, runit, s6 and an Android init.rc can all express. Everything 0059 decided is kept: two watchdogs, roll back once, recovery is local, the rollback shares no code with the host. The counter is the whole mechanism, so it is what the tests are mostly about. Three real problems came out of writing them: A counter file holding "1 2" became "12" -- `tr -d [:space:]` concatenates rather than rejecting -- which is past the limit, so a HEALTHY node rolled itself back. Now it reads the first field and insists on a plain integer. The corrupt-counter test used "not-a-number", which shell arithmetic happens to evaluate to 0, so it passed with the guard removed and proved nothing. Replaced with values that discriminate: "5x" errors under set -e and kills the launcher, and "0x10" is read as HEX 16 -- past the limit, so again a healthy node rolls back. And the test harness itself was wrong. With `set -e` and a bare launcher call, removing a guard killed the script at the first corrupt case and silently skipped everything after -- reporting a full pass over tests that never ran. Every launcher call now records its failure instead of aborting. Same class as the placebo assertion found last time, and the reason to keep injecting faults rather than trusting green. Both scripts run in `make check`. 27 launcher tests, 9 rollback tests, all confirmed to bite. |
||
|
|
f4143806c2 |
Build the rollback mechanism, and test it
ADR 0059's recovery path: the pieces that run when the host will not start. internal/upgrade -- two facts, neither of them the host judging its health. Whether the executable this process started from has been replaced on disk, and which version last completed a reconcile. The first design was wrong and the tests caught it, not review. It asked /proc/self/exe whether it was marked deleted. That is Linux procfs behaviour rather than a fact about files, and it catches only unlink -- a binary swapped by rename onto the same path reads as untouched, which is exactly what a package manager does. Now the identity is captured at start and compared later: no procfs, and neither case missed. known-good is one bare line. The reader is a shell script on a machine where the host is failing to start, so it must not need a parser to be present and working. Written only after a clean apply, which is the whole claim -- not health, because a disconnected node is ordinary and a failing resource is the machine's problem rather than the binary's. packaging/ -- the unit, the rollback unit, and the rollback script. The script shares no code with the host and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, nothing that has to be installed. The unit carries Restart=always with a comment saying why on-failure would break every upgrade. Both are tested and both sets of tests were confirmed to bite. Injecting five faults broke exactly the intended tests -- except one, and chasing why it did not found a placebo assertion I had written: `check "exits zero" ... "0" "0"` compares a literal to itself and can never fail. Replaced with the real exit code, after which the injection bites. Also caught: an injection that produced a build failure rather than a test failure, which my grep read as "no failure". Re-run so it compiled, and the test did bite. The script test runs in `make check`, so it is a gate rather than something that was run once. Verified against the real binary: known-good is written beside the store after a clean apply and is NOT written after a failed one. |
||
|
|
08a1263a81 |
Stage 2 — the bundle a host carries
novox/hq ADR 0038: one behaviour, two sources of declaration. This is the source that does not need a mesh — the first node's path. The bundle is embedded in the binary rather than shipped beside it, because "copy it onto a machine and run it is the whole installation" stops being true the moment a second file has to arrive with it. `make host BUNDLE=...` builds a host carrying one; `mesh-host reconcile` applies it; `mesh-host bundle` shows it. A default build carries nothing and REFUSES to reconcile, saying why. A host that applied nothing and reported success would look exactly like one that raised a first node, and the difference would surface later as a mesh that never came up with nothing to point at. Proved on a sealed machine: no route out, no name resolution, one binary copied on, and it configured itself from what it carried. Idempotent on the second run. One bug found by running rather than reasoning, and it is a shape worth naming: `mesh-host bundle` validated the carried bundle through a path that strips comments, while `reconcile` handed the raw bytes to the parser. So the command whose whole job is to check the bundle said yes, and the command that uses it said no. Two paths to one artefact, disagreeing. There is one path now, and a test asserts that what validates is what is applied. What this does NOT prove is stated in the README rather than left implied: the claim under stage 2 is that one host can raise the substrate alone, and the substrate is four container services. There is no container type, because a container needs an image and where images come from is open; what belongs in a substrate is not known, because the closure for a one-node mesh is what research 011 and 012 exist to answer; and the machine used to test this cannot install a container runtime through a sealed network. The mechanism is finished. The claim is not, and shipping a host that claimed a substrate it has never raised would be the fault this whole project is about. 65 tests. |
||
|
|
73c010e7ef |
Stage 1 — the host reports what a machine is and can do
Tier 0's first slice, per novox/hq 03-DESIGN/01-to-be/05-the-node-host.md. It applies nothing, connects to nothing, listens on nothing. 2.9 MB, static, no dynamic dependencies: copy it onto a machine and run it is the whole install, which is the property ADR 0041 rests on. A capability is detected, never assumed. Every detector runs something that only succeeds if the thing FUNCTIONS — the daemon is asked for its version, the package database is queried, the firewall is asked to list a ruleset, which needs the privilege as well as the tool. 04-ISSUES/007 is the fault this prevents: a client on disk with its daemon down looks exactly like a working runtime, and a node assigned work on that basis fails when the work arrives. Every verdict carries the reason and the method. A capability reported absent with no reason is the same fault in a new place: something nobody can act on. Two bugs found by running rather than reasoning, both silent: systemctl is-system-running exits non-zero for every state except `running` — including `degraded`, which means units failed and the init is emphatically there. Reading the exit code reported NO service manager on a machine whose init it was. That is 007 in the mirror, and both directions place work wrongly. A verdict now reads what a tool says about itself, not only how it exited. And `mesh-host inventory --json` printed text: the standard library stops parsing at the first non-flag argument, so the flag sat unread and the command exited 0 having ignored what was asked. The parser now takes the subcommand off the front, and a stray or mistyped argument is refused rather than dropped. Detection deliberately does NOT follow ADR 0008. That rule governs applying state, where a failed step means the machine is not what was asked for. A failed probe is a finding — "absent, because the probe failed" — and aborting would replace one legible absence with total ignorance of the rest. 25 tests: structure and logic with a fake runner, and the same detectors against this machine, because a test that fakes the system under detection asserts only that the fake behaves as expected. |