Design the node lifecycle end to end
The host was described as a component and never as something that runs for years on a machine somebody else also uses. 09 covers every state a machine can be in and every transition between them. Four states: unmanaged, hosted, enrolled, disconnected. Only the last two are nodes, and they are the same node in two situations. `hosted` -- the host installed but never told which mesh it belongs to -- had no name before and is where a machine sits between the two adoption commands. Things that were unclear and now are not: The first node walks the same path in an unusual order: reconcile from the bundle, the control plane it just raised issues a token, enrol against it. Its specialness lasts two commands. A side effect worth having -- enrolment is exercised on node one, rather than being written and first used on node two. Enrolment reports profile and inventory BEFORE the control plane decides anything. The profile is the input to that decision, not a diagnostic; the control plane cannot decide what a machine should run without knowing what it can run. Rebooting mid-apply is safe by construction. The store records each resource after it worked, so a host that dies half way through comes back and applies the rest. The rule that stops the host lying about what it did also makes it crash-safe. Retiring splits in two. Graceful is a final empty declaration. A node that is gone will reconcile its last declaration forever -- the honest consequence of making disconnection ordinary. The answer is not to make the host expire but that the node holds nothing that outlives revocation: every grant is a per-node credential revoked at the provider. A lost node keeps running and stops being able to reach anything. Said plainly rather than implying the mesh can switch a machine off, which it cannot and should not. Losing the store is quiet and permanent, so it gets its own section. The host re-enrols and re-applies fine; what does not come back is removal, because resources it no longer has a record of become unowned and sit there indefinitely. Also corrects 0057, which said the mesh must not upgrade the host at all. That conflated two acts. Replacing the binary is safe -- Unix keeps the running inode. Stopping the unit is not. So the host may apply a package naming itself, and restarts by finishing its apply and exiting cleanly, letting the supervisor start it on the new binary. It never asks the service manager to restart it. That makes a fleet-wide host upgrade an ordinary declaration, which the first draft gave up on. 0057 remains proposed.
This commit is contained in:
@@ -18,6 +18,7 @@ document is written and this one's status becomes `implemented`.
|
||||
| [`06-the-control-plane.md`](06-the-control-plane.md) | Tier 2 — what the term means, and the test for what belongs in it | [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) |
|
||||
| [`07-the-substrate.md`](07-the-substrate.md) | Tier 1 — what the control plane consumes and cannot grant itself | [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md), [0048](../../02-DECISIONS/0048-the-substrate-is-named.md) |
|
||||
| [`08-connectivity.md`](08-connectivity.md) | One context in full — overlay, resolution, exposure, filtering, certificates | [ADR 0049](../../02-DECISIONS/0049-a-route-is-a-grant.md), [0050](../../02-DECISIONS/0050-reachability-is-a-property-of-the-address.md), [0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md), [0055](../../02-DECISIONS/0055-the-control-plane-is-the-node-coordinating-contexts.md) |
|
||||
| [`09-the-node-lifecycle.md`](09-the-node-lifecycle.md) | How a machine becomes a node, stays one, and stops being one | [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md), [0051](../../02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md) |
|
||||
|
||||
## Not yet written
|
||||
|
||||
|
||||
Reference in New Issue
Block a user