e1ad39b5007374993d6db95bb545a25f1b387327
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e1ad39b500 |
Per-OS hosts, and an init asked for only start and restart
0060 -- the host is built per operating system. systemd and pacman are the Arch host's implementation, not abstractions the mesh has to grow. They are not independent choices: a machine has pacman because it is Arch, and the package manager, service manager and packaging format arrive together as one decision somebody made at install time. Rejected abstracting them, and the reason is correctness rather than effort. The service applier reads LoadState to tell "not installed" apart from "stopped", which is what stops it reporting absence as success. An interface spanning systemd and OpenRC degrades to what both express, and the lowest common denominator is exactly where that fault lives. Almost all of it is shared -- the vocabulary, store, apply loop, read-back discipline, refusal model, bundle and link are portable. Two appliers differ. And delivery was already per-OS, since a .pkg.tar.zst is an Arch artifact, so this is the seam that already existed. Android is the interesting case rather than Debian: no service manager, no package installation, usually no root. Such a host implements file, directory and action and refuses the rest -- the same refusal a host already gives an unknown type, with a different reason. Those three are the portable floor. The container runtime is deliberately left open: it is not an OS split, since Arch runs docker or podman. 0061 -- the init is asked for start-at-boot and restart-on-exit, and nothing else. Both are expressible in OpenRC, runit, s6 and an Android init.rc. Counting failed starts and rolling back moves into a launcher, because that is the one piece which must work when the host does not, and a script with a counter can be tested where OnFailure= can only be hoped for. Supersedes 0059, keeping its reasoning in full. The checker found all six places citing 0059 and refused the commit until they named the replacement. |
||
|
|
19997d56c3 |
Approve 0057-0059, with four corrections from review
Not approved as drafted -- four things came out of checking them against each other, and one was a bug that would have broken every upgrade. The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto a new binary by exiting CLEANLY. on-failure does not restart a process that exited zero, so every upgraded node would have been left stopped, having successfully upgraded. Found by reading the two records against each other rather than by either alone. Now Restart=always in all three places that mention it. The host cannot run in a container, and the reason is decisive rather than stylistic: step 0 of the substrate bootstrap installs the container runtime, so a host inside a container would need the thing it exists to install. It would also break 0041 -- copy it onto a machine and run it stops being true when the machine must already have a runtime. Everything above tier 0 is a container; the host is not. That split is the tier boundary, not an inconsistency. systemd is named rather than abstracted. An init is not a dependency in 0041's sense: 0041 is about what must be installed before the host works, and an init is not installed, it is what the machine already is. The unit file is the only systemd-specific artefact and it belongs to the package, so a machine with a different supervisor ships a different package. The mesh is a watchdog, and my first draft was half an answer. Recovery must be local -- nothing dials a node, and a host that cannot start cannot report. But detection is the mesh's, and a local supervisor structurally cannot do it: it sees one process failing and cannot tell a broken machine from a broken release. Only something watching every node can, and that distinction decides whether the response is "fix this machine" or "stop shipping this version". So a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop on silence. Local rollback still needed, because the canary nodes break and because a node offline during the rollout gets the declaration later with no batch around it. The first declaration is the overlay and nothing else. Forced, because a node's address and peers are assigned rather than chosen. But also the way back in: a node reachable over the overlay can be fixed by hand if a later declaration breaks it, and a large first declaration risks a node that is broken and unreachable at once. Also stated plainly, because it reads as a contradiction: nodes reach each other over the overlay and every node consumes from the broker; what 0039 forbids is an inbound CONTROL surface, not reachability. And in 06: no node holds a credential to any control-plane store, for reads or writes. Four ADRs already say this separately and none of them said it in one place. Nodes state over the broker; the owning context writes. With a note that most high-frequency writes are observability's, not the registry's -- routing logs into the registry would be the shared-schema mistake arriving through a door marked performance. |
||
|
|
605c9fd441 |
Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed. |