diff --git a/02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md b/02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md new file mode 100644 index 0000000..d2fbd27 --- /dev/null +++ b/02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md @@ -0,0 +1,59 @@ +--- +topic: the mesh +status: accepted +date: 2026-09-18 +deciders: jochen +reconstructed: false +extends: 0010-delivery.md +--- + +# 83. One push leaves the mesh consistent + +## Context + +A provision is minted while composing the *consumer's* node; the provider's grant list is a pure +read of secrets already issued from it. So assigning a cross-node consumer and pushing its node +produced a consumer that retried forever against a provider that had never heard of it, until the +provider's node was pushed a second time — an action with no signal to take, documented nowhere, +and invisible whenever consumer and provider share a machine +([issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)). + +Two remedies were on the table: **cascade** — a push also delivers to the machines its compose +changed — or **report** — a push says "now push the provider" and leaves the act to the operator. + +## Decision + +A push finishes what it starts: after composing and sending the named node, the controller flushes +every *other* machine that is now behind — whose declaration differs from what it was last sent — +by name, in the push's own output, converging over a bounded number of rounds (a flushed send may +itself mint). + +Behind is measured against what a machine was last *sent*, not against a before/after snapshot of +this push. The mint that makes a provider behind happens when the consumer is assigned or its +account issued — before `push` runs at all — so by push time the provider already differs from +what it holds, with no in-command delta to detect. The only durable signal is "what it should be" +versus "what it last received", which is the same comparison `push --behind` already makes. + +A machine behind for an unrelated reason is flushed by this too, and that is correct rather than a +cost: a named push that knew a machine was behind and left it so would be the very silence this +decision removes. The narrower reading — flush only what this push provably changed — was +rejected because it cannot see a mint that a prior command performed, which is precisely the 057 +case. + +Reporting alone was rejected because it converts a derived fact the controller already holds into +an operator obligation, and an obligation enforced by nothing is issue 057 restated. The +declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and +merely saying so would make "push succeeded" mean less than it says. + +## Consequences + +- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same + act that minted the provision. The undocumented rule "push the provider node too" ceases to + exist rather than becoming documentation. +- A named push delivers to every machine that is behind, not only the one named — each named in + the output, never silent. `push --behind` remains the way to reconcile the mesh without naming + a node; a named push now carries the same guarantee for the machines its work touched and any + others already waiting. +- How this is checked: the built-store-cross-node bed registers a cross-node consumer, pushes + only the consumer's node, and asserts the provider minted its vhost — the workaround push is + removed, so a regression fails the bed. diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 5a587cb..d61a1e6 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -85,6 +85,7 @@ python3 00-META/checks/index.py fail if stale - **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md) - **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md) - **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md) +- **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md) ### Its tiers, from the bottom up diff --git a/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md index a818050..7c8e7ea 100644 --- a/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md +++ b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md @@ -1,9 +1,10 @@ --- -status: open +status: located opened: 2026-09-17 -located-in: [] +located-in: + - mesh-controller fixed-by: -amended-design: +amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md --- # 057 — A cross-node consumer is provisioned only when the provider is pushed again diff --git a/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/01-diagnosis.md b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/01-diagnosis.md new file mode 100644 index 0000000..ebe8e00 --- /dev/null +++ b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/01-diagnosis.md @@ -0,0 +1,16 @@ +# Diagnosis — 2026-09-18 + +The mechanism was already understood when the issue was opened: a provision secret is minted as a +side-effect of composing the consumer's node, and the provider's grant list is a pure read of +secrets already issued — so a provider composed before the consumer existed, and never again, is +blind to it. The open question was the remedy: cascade the push, or report the obligation. + +Decided as [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): a push +finishes what it starts. The controller captures the digest of what every machine should be +before composing the named node, recomputes after, and delivers to exactly the machines whose +declaration changed because of this push — named in the output, bounded rounds, converging. +Machines behind for unrelated reasons stay the business of `push --behind`. + +**Located in:** mesh-controller (the push command). Checked by the built-store-cross-node bed, +which now pushes only the consumer's node and asserts the provider minted the vhost — the +workaround push is removed, so a regression fails the bed. diff --git a/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md b/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md index 62feb5c..b10816a 100644 --- a/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md +++ b/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md @@ -1,7 +1,8 @@ --- -status: open +status: located opened: 2026-09-17 -located-in: [] +located-in: + - mesh-tools fixed-by: amended-design: --- diff --git a/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/01-diagnosis.md b/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/01-diagnosis.md new file mode 100644 index 0000000..b30706e --- /dev/null +++ b/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/01-diagnosis.md @@ -0,0 +1,17 @@ +# Diagnosis — 2026-09-18 + +1. The exit is the shared runtime's: serve mode connected the broker once and treated any failure + as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a + joined node — became an exit, a container-runtime restart, and a visible crash-loop. +2. The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the + shape is general: any reachability failure at startup had the same consequence. +3. Answering the report's open questions: the shared runtime now retries the broker connection + in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it + is fatal" is answered *never*, because the failure modes that do not heal are not reachability: + a pinned-certificate mismatch still refuses immediately (an impostor does not become the + broker by being asked again), and configuration errors still exit at once. One-shot commands + (emit, invoke, run) still fail fast — a person is waiting. +4. Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime + restart count of zero. + +**Located in:** mesh-tools (the shared runtime's serve mode). diff --git a/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/00-report.md b/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/00-report.md index 9719020..75f66c4 100644 --- a/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/00-report.md +++ b/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/00-report.md @@ -1,8 +1,10 @@ --- -status: open +status: resolved opened: 2026-09-17 -located-in: [] -fixed-by: +located-in: + - mesh-catalog + - mesh-controller +fixed-by: mesh-catalog feat/the-mesh-builds-its-catalogue (965c58f); proven by the built-store-cross-node bed and direct builds amended-design: --- diff --git a/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/01-diagnosis.md b/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/01-diagnosis.md new file mode 100644 index 0000000..387be9d --- /dev/null +++ b/04-ISSUES/060-the-mesh-cannot-build-most-of-its-own-catalogue/01-diagnosis.md @@ -0,0 +1,32 @@ +# Diagnosis — 2026-09-18 + +The gap was mechanical for most of the catalogue and structural for a few. + +**The mechanical part (fixed).** 44 of 70 modules now carry a `build` section and a Dockerfile +mirroring the eight that already had one: two named bases (build and runtime), the module's own +source compiled against the sdk in the base, and every serve-time entrypoint named in the runtime +image so tools serve, events flow, and a provider's provisioner reconciles in the same process +(the convention hq issue 061 settled — postgres, gitea and lavinmq were retrofitted to it from +running their provisioner as a separate `args` command). Proven end to end: redis, keycloak, +mongodb, mosquitto and openai-consumer build through the mesh's own builder from their own +directory; the built-store-cross-node bed builds and runs the retrofitted lavinmq and postgres. + +**The structural part (deferred, each with an owner).** + +- Three modules need an artifact the mesh build environment cannot fetch — an npm package + (model-usage: `pg`; anthropic-manager: `tweetnacl`) or a binary from a public image (minio: + `mc`). The build's npm and container runtime reach the mesh's own registry, not the public one; + apt works, npm and image pulls do not. This is its own question — [issue 064](../064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md). +- route-proxy's build context is the mesh-controller repository, not its own module directory (it + ships the packaging of a proxy whose canonical source lives with the contract). The build + section names a Dockerfile relative to the module; a cross-repository context is a shape it + cannot yet express. Deferred pending a design decision on cross-repo build inputs. + +The remaining modules are exempt rather than unbuilt: the foundation trio (builder, +mesh-controller, distribution — built by the foundation's own path, and distribution cannot build +through the store it provides, ADR 0072/issue 029), and the upstream-image-only modules that run +a stock image and carry no code of their own. + +**Located in:** mesh-catalog (the manifests and Dockerfiles), with the convention enforced by +mesh-controller's builder. Resolved for the mechanical majority; the two structural remainders are +tracked as issue 064 and the route-proxy cross-repo deferral. diff --git a/04-ISSUES/063-the-foundation-ports-are-accepted-on-input-but-not-forwarded/00-report.md b/04-ISSUES/063-the-foundation-ports-are-accepted-on-input-but-not-forwarded/00-report.md new file mode 100644 index 0000000..7890f37 --- /dev/null +++ b/04-ISSUES/063-the-foundation-ports-are-accepted-on-input-but-not-forwarded/00-report.md @@ -0,0 +1,50 @@ +--- +status: resolved +opened: 2026-09-18 +located-in: + - mesh-controller +fixed-by: mesh-controller fix/one-push-is-enough (25e42b3); gated by the built-store-cross-node bed +amended-design: + +--- + +# 063 — The foundation's ports are accepted on input but not forwarded + +## Symptom + +A joined node reaches the mesh's broker once, then — after the broker restarts — can never reach +it again, and so never receives another declaration. Every check between them passes: the port is +open, the broker is up, the overlay is established. + +Observed on the built-store-cross-node bed, intermittently. A joined node's host connected to the +broker, received its first declaration, and worked. When the foundation's own broker was adopted — +which restarts it — the node's link dropped with `CONNECTION_FORCED - Broker shutdown`, then +`connection refused`, then `i/o timeout` for ever. The node was left with no declaration on disk +at all, and every downstream assertion (here, the registry trust) failed as a consequence of a +node that had received nothing. + +The bed passed on the runs where the broker did not happen to restart after the joined node first +connected, which is why three green runs preceded the red one on identical code. + +## Why it matters beyond the instance + +The mesh firewall opens the foundation's ports — the broker above all — in the **input** chain, +from anywhere, so a node can enrol before it has an overlay address. But the broker is a published +container port: a cross-node packet to it is redirected by the runtime and handled in the +**forward** chain, which never carried a rule for the foundation. The first connection survived +only on its conntrack `established` entry; a broker restart dropped the entry, and the next dial +hit the forward chain's default drop. + +This is [issue 047](../047-the-firewall-does-not-cover-published-container-ports/00-report.md)'s +lesson — a firewall with no forward coverage says nothing about container ports — reappearing for +the one port the whole mesh depends on, because the foundation is not a module and was added to +input alone. A rule that works only until the thing it governs restarts is worse than none: it +passes every test written before the restart. + +## Open questions + +- Is the broker the only foundation port that resolves to a container, or should every foundation + port be assumed to be published and forwarded on principle? (The fix forwards all of them, on + that principle.) +- Should a bed assert reachability *after* a deliberate broker restart, so "reachable until it + bounces" can never again read as "reachable"? diff --git a/04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md b/04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md new file mode 100644 index 0000000..e427532 --- /dev/null +++ b/04-ISSUES/064-a-mesh-build-cannot-fetch-a-modules-external-dependencies/00-report.md @@ -0,0 +1,51 @@ +--- +status: located +opened: 2026-09-18 +located-in: + - mesh-catalog + - mesh-controller +fixed-by: +amended-design: +--- + +# 064 — A mesh build cannot fetch a module's external dependencies + +## Symptom + +Three modules cannot be built by the mesh's own builder, each because the build needs an artifact +from a public registry the build environment does not reach: + +- A module whose runtime installs an npm package its code imports (`npm install` in the + Dockerfile) fails with a 404 — the build's npm is pointed at the mesh's own registry, which does + not carry the public package. +- A module whose runtime copies a binary out of a public image (`COPY --from=vendor/tool:latest`) + fails with `pull access denied` — the build's container runtime reaches the mesh's registry, not + the public one. + +apt-based installs in the same Dockerfiles succeed, so the build has ordinary internet: it is npm +and image pulls specifically that are redirected to the mesh's registry. Observed while giving the +catalogue its build sections (hq issue 060): 44 of 70 modules build from their own source with no +external fetch; these three need one and stop. + +## Why it matters beyond the instance + +The delivery design says a module is built by the mesh from a repository and a path (ADR 0069), +compiled against the sdk and tool runtime the base images already carry. That holds for a module +whose only dependencies are the base's — but a module with a third-party runtime dependency, or a +tool binary from a vendor image, has no sanctioned way to bring it into a mesh build. The +workstation build script got away with it by running on a host with public npm and Docker Hub; +retiring that shortcut (the no-fake principle) exposes that the mesh has no answer. + +Left unanswered, "the mesh builds its own catalogue" is true only for modules that happen to have +no external dependency, and which those are is invisible until each is built. + +## Open questions + +- Should a module's third-party npm dependencies be published to the mesh's own registry as part + of building it (a dependency is an artifact like any other), or should the build environment + proxy the public registry read-only? +- Is a vendor binary (`mc`, and others like it) the same question as an npm dependency, or does it + want a distinct answer — a module declaring an external image as a build input the mesh + pre-fetches and pins, the way it pins its bases? +- Should a manifest that declares a build whose Dockerfile fetches from a public registry be + refused at `module add`, so the gap is caught at registration rather than at build?