Changes are pushed, not polled; and a stuck host rolls itself back
Two corrections and one new decision, all from Jochen catching things. Pushed, not polled. I described updates as landing "on the next reconcile", which reads as polling and is not the design. A declaration arrives as a message on a link that is already open; the host applies it then. Polling over an existing connection would be slower to land AND constant traffic to learn nothing. The timer is for drift and nothing else, and it cannot be replaced by an event for a definitional reason: drift is change the mesh did not make -- somebody edited a managed file, a distribution upgrade replaced a config -- so nothing will ever publish a message about it. Only looking finds it. Separated the heartbeat from the reconcile timer, which I had been conflating. They point in opposite directions and answer different questions: the timer looks at the machine and asks whether it still matches; the heartbeat reports upward and is what makes silence mean something. A node with nothing to do sends nothing, and without a heartbeat that is indistinguishable from a node that stopped. 0059 -- a host that cannot start is rolled back by the service manager. I had left this open on the grounds that recovery meant the host judging its own health. That objection does not survive being asked properly: a keepalive is something else judging the host. The watchdog must be local, because nothing dials a node and a host that cannot start cannot report -- so it is the service manager, which is already there. The failure it prevents is sharper than "the node is down": a host that will not start looks exactly like a machine somebody switched off, which is the one condition this design has deliberately decided not to alarm on. So a bad release reaches every node, each goes quiet, and the mesh reports a fleet of sleeping laptops. Confirmed means started and completed one reconcile -- deliberately not "the link is up", or a laptop on a train would roll itself back. The rollback is a script shipped by the package, not a host subcommand, because a binary that will not start cannot be its own recovery. It rolls back once: a second failure means the machine is the problem, not the binary. Also refined the records checker, which produced a false positive: a proposed record may extend another proposed one, because decisions are drafted in chains and the alternative is marking things accepted to satisfy a check. An accepted document resting on a proposed record still fails, and that was verified. 0057, 0058 and 0059 are all proposed.
This commit is contained in:
@@ -153,6 +153,12 @@ def check_rests_on(failures, records):
|
||||
continue
|
||||
status = records[number]["front"].get("status")
|
||||
if status != "accepted":
|
||||
# A proposed record may extend another proposed one. Decisions are drafted in
|
||||
# chains -- 0059 extends 0057 while both await review -- and refusing that would
|
||||
# mean either drafting out of order or marking records accepted to satisfy a
|
||||
# check, which is the failure this repository already made once.
|
||||
if frontmatter(read(path)).get("status") == "proposed":
|
||||
continue
|
||||
# An as-is document describes what runs, and what runs was built under
|
||||
# whatever was decided at the time. ADR 0056: "as-is describing a superseded
|
||||
# decision is exactly what as-is is for."
|
||||
|
||||
Reference in New Issue
Block a user