Files
hq/04-ISSUES/051-the-mesh-cannot-update-what-it-depends-on/00-report.md
T
jschoubben f52895a646 Issue 051 — the mesh can update everything except what it depends on
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.

That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.

The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.

Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:44:10 +02:00

4.6 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-14

051 — The mesh can update everything except what it depends on

Symptom

A mesh raised from bare metal runs its store and its broker from a bundle the installer wrote once, at install time, with the images pinned in it. Nothing can change them afterwards. There is no build for them, no version for them to be behind, no upgrade … roll-out, and no way for the mesh to report that either is out of date — because the mesh holds no record of them as modules at all.

Every other thing on that machine has all of it.

Observed on a one-node mesh, where the duplication makes it plain — the same image, twice:

mesh-store      <the postgres image>     the control plane's own records
postgres        <the same digest>        the postgres module's server
mesh-postgres   <mesh-built runtime>     that module's provisioner and tools

One of those two servers can be rebuilt from source and rolled out. The other cannot, and it is the one holding the mesh's inventory, identity and licences.

Why this matters

It is exactly backwards. The store and the broker are the components every other thing depends on, so they are the ones whose security updates matter most, and they are the only ones the mesh has no mechanism to deliver. A release that reaches every module's database does not reach the one the control plane keeps its own records in.

It is invisible rather than reported. status says what is behind its source. The substrate cannot be behind anything, because the mesh does not know it exists as something with a source. So a mesh with a year-old broker reports itself entirely current, which is worse than reporting a problem.

The pattern that would fix it already exists and is proven. The control plane is carried in, raised, and then adopted as an ordinary module pinned to the image that is running — the pivot in ADR 0067. Nothing about that pattern is specific to the control plane. Applied to the store and the broker it would make the substrate a moment rather than a kind: how a mesh starts, not what it permanently is.

It is also why the same image runs twice. A store that cannot be a module cannot provide postgres-database, so a module wanting a database needs a second server. Adoption removes the duplication as a side effect, and the control plane already reaches its three contexts through three separate credentials — which is the shape of a consumer, not an owner.

What makes this harder than it looks

The recursion is real, not incidental. The control plane learns what modules exist by reading its store. A store that is a module is a record inside the thing it is holding up. Adoption is what resolves it — the store is raised by the installer because nothing else can, and only afterwards becomes something the mesh has a record of — but the order of operations during an upgrade needs stating, not assuming: a machine being sent a new store while the control plane is reading from it is not a rollout, it is an outage.

And the broker carries the rollout itself. A declaration reaches a machine over the broker. A broker upgrade is the mesh asking the machine to replace the thing the request arrived on.

Open questions

  • Is adoption the right mechanism, extending ADR 0067 to the substrate, or should the substrate stay outside the module system and gain its own narrower update path?
  • What does a store upgrade look like when the control plane is mid-read? Drain, quiesce, or a window where the mesh accepts that it cannot be asked anything.
  • What does a broker upgrade look like when the instruction travels over it? A machine that is told to replace its broker has to complete the work without being able to report progress.
  • Does adoption also mean one server instead of two by default, with a separate one for the control plane as a choice for those who want the isolation — which is assigning a different provider, a mechanism that already exists?
  • What reports it today? Whatever is decided, status should be able to say the substrate is behind. Today it cannot form the sentence.

How this would be checked

Rule Checked by
The mesh can say its substrate is out of date The store's source moves and status reports it behind, the way it does for any module.
The mesh can deliver a substrate update A store or broker is upgraded on a running mesh and the control plane is answering afterwards, with its records intact.
Nothing runs twice without a reason A mesh with one machine runs one postgres unless somebody asked for two.