Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
Showing only changes of commit dcc4b8339c - Show all commits
@@ -108,6 +108,40 @@ what the message vocabulary can say.
Reads work the same way in reverse — a node is *told*, in declarations. It never asks.
### Who actually consumes, and who writes
**The control plane is the consumer. There is one of it, and the context that owns the data does
the write.**
```
node ──► broker ──► the control plane, consuming
├─ a node reported what it applied ─► inventory writes the registry
├─ a node reported health ─► observability writes its own store
└─ a grant was requested ─► provisioning writes its own store
```
Seven contexts, **one deployable** — they are not separate services, so this is one process
consuming and dispatching internally, not seven consumers racing. Each context then writes only
the store it exclusively owns
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)).
**One consumer is a property worth having**, not just a consequence of
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). The as-is records that
*two consumers accidentally sharing one queue silently split the traffic between them, each
receiving half of what it expects* — which has happened, between a module's daemon and its
capability server. With one consumer that class of fault cannot arise.
**And the broker is the buffer while the control plane is down.** Nodes go on publishing;
messages queue; the control plane drains them when it returns. That is what makes
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)'s single control plane
tolerable — an outage delays the mesh's *knowledge* rather than losing it.
**With one consequence that must be bounded before it is discovered:** a queue with no limit
grows until the broker's disk is full, and the broker is the one component every node depends
on. Queues carrying node reports need a maximum length or a message lifetime, and losing the
oldest health report is obviously right where losing the oldest declaration acknowledgement is
not — so the bound is per queue and is not decided here.
### On volume, which is the real worry underneath
**Most high-frequency writes are not registry writes, and that is the first thing to check