diff --git a/03-DESIGN/01-to-be/06-the-control-plane.md b/03-DESIGN/01-to-be/06-the-control-plane.md index c074a74..c4a1902 100644 --- a/03-DESIGN/01-to-be/06-the-control-plane.md +++ b/03-DESIGN/01-to-be/06-the-control-plane.md @@ -108,6 +108,40 @@ what the message vocabulary can say. Reads work the same way in reverse — a node is *told*, in declarations. It never asks. +### Who actually consumes, and who writes + +**The control plane is the consumer. There is one of it, and the context that owns the data does +the write.** + +``` +node ──► broker ──► the control plane, consuming + ├─ a node reported what it applied ─► inventory writes the registry + ├─ a node reported health ─► observability writes its own store + └─ a grant was requested ─► provisioning writes its own store +``` + +Seven contexts, **one deployable** — they are not separate services, so this is one process +consuming and dispatching internally, not seven consumers racing. Each context then writes only +the store it exclusively owns +([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). + +**One consumer is a property worth having**, not just a consequence of +[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md). The as-is records that +*two consumers accidentally sharing one queue silently split the traffic between them, each +receiving half of what it expects* — which has happened, between a module's daemon and its +capability server. With one consumer that class of fault cannot arise. + +**And the broker is the buffer while the control plane is down.** Nodes go on publishing; +messages queue; the control plane drains them when it returns. That is what makes +[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md)'s single control plane +tolerable — an outage delays the mesh's *knowledge* rather than losing it. + +**With one consequence that must be bounded before it is discovered:** a queue with no limit +grows until the broker's disk is full, and the broker is the one component every node depends +on. Queues carrying node reports need a maximum length or a message lifetime, and losing the +oldest health report is obviously right where losing the oldest declaration acknowledgement is +not — so the bound is per queue and is not decided here. + ### On volume, which is the real worry underneath **Most high-frequency writes are not registry writes, and that is the first thing to check