Compare commits
11
Commits
da41c1cc40
...
a97feeefe5
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a97feeefe5 | ||
|
|
a18550e460 | ||
|
|
f10dce4f9e | ||
|
|
4beb6629db | ||
|
|
d497b37e43 | ||
|
|
183b22997c | ||
|
|
8d67cf63c5 | ||
|
|
75c104c355 | ||
|
|
6a56738d7b | ||
|
|
91a5c63d65 | ||
|
|
545d038198 |
+2
-1
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: canonical
|
||||
updated: 2026-08-23
|
||||
updated: 2026-09-24
|
||||
---
|
||||
|
||||
# The Novox repositories
|
||||
@@ -17,6 +17,7 @@ and a forge address is an operational detail (see [`README`](../README.md)).
|
||||
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
|
||||
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
|
||||
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
|
||||
| `migration` | **Private.** The record of one installation replacing the predecessor mesh with this one: the runbook, a dated log of every step and what it cost, the per-service data procedures, the readiness checks, and the scripts. Private because it is the opposite of this repository in every way that matters — it names machines, addresses, ports and paths, because a procedure that cannot be followed is not one. Where hq asks *what did we decide and why*, that repository answers *what happened on the machines, in what order, and what to do next*. Its `HANDOFF.md` is where somebody picking the work up starts. |
|
||||
|
||||
## What the mesh becomes
|
||||
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
---
|
||||
status: active
|
||||
initiated: 2026-09-24
|
||||
touches:
|
||||
- 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
|
||||
- 02-DECISIONS/0033-the-substrate-is-a-store-and-a-broker.md
|
||||
- 02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md
|
||||
- 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
|
||||
- 02-DECISIONS/0078-the-store-and-broker-are-modules.md
|
||||
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md
|
||||
- 03-DESIGN/00-as-is/03-provisioning.md
|
||||
- 03-DESIGN/01-to-be/07-the-foundation.md
|
||||
- 04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md
|
||||
---
|
||||
|
||||
# 015 — The object store after MinIO: which S3 implementation, and how the data moves
|
||||
|
||||
**The question.** The mesh's object store is MinIO. Its community edition is archived upstream,
|
||||
its server and client images have been deleted from every public registry, and the pinned release
|
||||
is four and a half years old and will never be patched
|
||||
([issue 113](../../04-ISSUES/113-the-object-stores-images-were-withdrawn-upstream/00-report.md)).
|
||||
Which S3-compatible implementation replaces it, and what is the migration track for the data and
|
||||
the provisioning model that sit on top of it?
|
||||
|
||||
**Why now, and why not sooner.** Nothing is on fire: nodes that already hold the images keep
|
||||
running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present
|
||||
image by design. The forcing function is not an outage but a one-way door — **no node that does
|
||||
not already hold the images can ever provision the module again**, so the mesh's ability to stand
|
||||
a node up from its declarations is already broken for this module, and silently.
|
||||
|
||||
**The direction is not a departure from the design; it is the design.** The foundation document
|
||||
already states the commitment:
|
||||
|
||||
> The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object
|
||||
> store, the OCI protocol for the registry. That is what keeps the naming safe rather than a
|
||||
> commitment that cannot be revisited.
|
||||
|
||||
The object store is also **not** a foundation service — ADR 0028 removed it, and it is an
|
||||
ordinary module required through the module graph by whatever wants one. (The "exception that is
|
||||
not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's
|
||||
own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket
|
||||
is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which
|
||||
is the cheapest kind of decision to make and the strongest kind to cite.
|
||||
|
||||
## What the replacement has to carry, measured
|
||||
|
||||
Taken from the module's manifest, its composition, its tool surface, and a search for its
|
||||
consumers across the catalogue — not from assumption.
|
||||
|
||||
| Requirement | Evidence in the module today |
|
||||
|---|---|
|
||||
| S3 API | The protocol every consumer speaks; already the design's stated dependency. |
|
||||
| **OIDC login against the mesh's identity provider** | Six configuration variables are wired and populated in practice — discovery URL, client id, client secret, scopes, display name, redirect — plus a dedicated entrypoint script that blocks startup until the provider answers. This is live, not aspirational. |
|
||||
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
|
||||
| A single-node form | Declared as a flavour, for development and small nodes. |
|
||||
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
|
||||
| A tool surface | Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning. |
|
||||
| A console | Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads. |
|
||||
|
||||
**Consumers, counted:** one application module, one capture module that takes a private bucket per
|
||||
node, one workflow module's tools, and the delivery/rescue internals of the shared library. The
|
||||
surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the
|
||||
OIDC story, not spread across the catalogue.
|
||||
|
||||
## Candidates
|
||||
|
||||
Scoped to **SeaweedFS** as the primary, with the others recorded so the rejection is not
|
||||
rediscovered.
|
||||
|
||||
- **SeaweedFS** — Apache-2.0, Go, twelve-plus years of development, erasure coding, and OIDC
|
||||
support in its S3/STS layer. Chosen to scope because it is the only candidate that plausibly
|
||||
preserves the OIDC requirement above, which is the one requirement that is live and least
|
||||
substitutable.
|
||||
- **Garage** — the lightest to operate and the simplest model, but **no native identity-provider
|
||||
integration**. Adopting it means losing OIDC console login or fronting it with a proxy. A real
|
||||
functional regression against something currently in use.
|
||||
- **RustFS** — markets itself as a binary-level drop-in retaining existing data, buckets and
|
||||
configuration, which would make the data migration close to trivial. Young, and that claim is
|
||||
exactly the kind that must be verified on a copy before it is believed.
|
||||
- **Ceph RGW** — the most capable and the most operationally expensive; disproportionate to a mesh
|
||||
where the object store is an ordinary module, not a platform.
|
||||
|
||||
**The first thing to verify, because the choice turns on it:** how much of SeaweedFS's OIDC story
|
||||
is in the freely licensed build, and whether its shape — IAM/STS token exchange — can actually
|
||||
stand in for a console that redirects a human to an identity provider. If it cannot, the honest
|
||||
finding may be that **no** candidate preserves the current feature set, and the decision becomes
|
||||
which regression to accept. That question is worth answering before any migration work starts.
|
||||
|
||||
## The migration track, in outline
|
||||
|
||||
Data movement is the easy half, and deliberately reversible.
|
||||
|
||||
1. **Stand the replacement up beside the incumbent**, on its own ports and its own provision type.
|
||||
No downtime, nothing removed.
|
||||
2. **Copy bucket by bucket with a neutral tool.** `rclone` rather than the incumbent's own client
|
||||
— the client has been withdrawn upstream too, so building the migration on it would inherit
|
||||
the same dependency this effort exists to remove.
|
||||
3. **Verify per bucket** — object counts and checksums, not a transfer exit code.
|
||||
4. **Repoint consumers through the connection the module already publishes.** Consumers read an
|
||||
API URL from the module's declared connections rather than addressing the store directly, so
|
||||
the cutover surface is that value plus the provisioning and tool handlers.
|
||||
5. **Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
|
||||
rollback until confidence is earned.
|
||||
6. **Retire**, and only then remove the module.
|
||||
|
||||
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
|
||||
handlers**, which are written against MinIO's admin API, and the OIDC wiring.
|
||||
|
||||
## Open questions
|
||||
|
||||
- How much of the OIDC requirement survives, and in which build? See above — this gates the
|
||||
choice.
|
||||
- Does the mesh's bucket provision translate to the candidate's identity model without weakening
|
||||
what ADR 0049 says about a consumer's identity fitting the tightest backend?
|
||||
- Should this effort also answer issue 113's general question — mirroring third-party images into
|
||||
the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with
|
||||
another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
|
||||
- Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking
|
||||
while the product is being chosen, rather than reproducing a shape by default.
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-09-24
|
||||
located-in: []
|
||||
located-in: [mesh-controller internal/inventory, mesh-controller internal/catalogue, mesh-controller cmd/mesh-controller, mesh-catalog modules/dnsmasq]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
+78
@@ -0,0 +1,78 @@
|
||||
# Diagnosis
|
||||
|
||||
## 2026-09-24 — where the mesh's own record ends, and what stands beside it
|
||||
|
||||
Checked the mechanism first. The resolver module's config (`mesh-catalog modules/dnsmasq`) is
|
||||
generated whole from `nodes.conf`, the `dnsmasq.fact-node-zones` fact `mesh-controller` computes:
|
||||
one `address=/<node>.<suffix>/<address>` wildcard per node the mesh has a name for, plus
|
||||
`local=/<suffix>/` — which tells dnsmasq it is *authoritative* for the whole suffix, so anything
|
||||
under it that is not one of those wildcards is refused, not forwarded. That is the exact mechanism
|
||||
the report describes: **the mesh does not lack an upstream to ask, it has told itself there is
|
||||
nothing to ask.**
|
||||
|
||||
Checked whether "forward to the resolver it replaced" (the third open question, ADR 0104's shape)
|
||||
is literal here the way it was for the proxy. It is not, and the difference matters: the proxy's
|
||||
predecessor kept running throughout its migration, a separate process on ports the mesh's proxy did
|
||||
not yet hold. DNS has one process on one port. `mesh-catalog`'s dnsmasq module, on assignment,
|
||||
replaces `/etc/dnsmasq.conf` whole and restarts the one `dnsmasq.service` unit — the same unit the
|
||||
predecessor's own resolver is running as, right now, on this control-node:
|
||||
|
||||
```
|
||||
$ systemctl status dnsmasq
|
||||
● dnsmasq.service ... Active: active (running) since Fri 2026-09-18 ...
|
||||
```
|
||||
|
||||
There is no predecessor process left standing after that assignment for a forward rule to reach.
|
||||
An adapter in ADR 0104's literal shape — a second thing running, written into by a first — does not
|
||||
fit; whatever answers this has to be data carried across the cutover, not a live thing forwarded to
|
||||
across it.
|
||||
|
||||
## What is actually missing, and where it already exists
|
||||
|
||||
The report's first open question — *should a carried peer be nameable, the operator saying which
|
||||
machine a carried address is, before it enrols* — turns out not to need a guess. The predecessor's
|
||||
own resolver config, still live on this control-node, already states it:
|
||||
|
||||
```
|
||||
# /etc/dnsmasq.d/hal-dns.conf — Generated by HAL dnsmasq-app
|
||||
address=/novox.internal/10.10.0.1
|
||||
address=/ace.internal/10.10.0.2
|
||||
address=/g14.internal/10.10.0.4
|
||||
address=/shanks.internal/10.10.0.3
|
||||
```
|
||||
|
||||
Checked against what the mesh itself recorded when it took the tunnel over (`overlay show`):
|
||||
|
||||
```
|
||||
a peer of the tunnel (zTKYP6Sz…) 10.10.0.2 not yet enrolled
|
||||
a peer of the tunnel (XT1l25Z0…) 10.10.0.3 not yet enrolled
|
||||
a peer of the tunnel (mq76YPW0…) 10.10.0.4 not yet enrolled
|
||||
```
|
||||
|
||||
`10.10.0.2`/`.3`/`.4` match `ace`/`shanks`/`g14` exactly, one for one. This is not the operator
|
||||
inventing a fact under time pressure — it is the predecessor's own record, current, and it has been
|
||||
correctly serving these three names for six days without correction. Recording it is transcription,
|
||||
not assertion.
|
||||
|
||||
**located-in, tentatively:** `mesh-controller internal/inventory` (where a carried peer's record
|
||||
lives today — `CarriedPeer`/`TunnelPeer` in `tunnel.go` carry a public key and an address but no
|
||||
name field) and `cmd/mesh-controller` (a command to set it — nothing today lets an operator attach
|
||||
a name to a carried-tunnel-peer record; `node add` is for enrolling nodes, not naming peers, and
|
||||
`node public-domain <name> <d>` is the closest existing shape to model a new verb on). Separately,
|
||||
`mesh-controller internal/catalogue` (`facts.go`'s `nodeZones`) would then need to emit an
|
||||
`address=` wildcard for a *named carried peer* the same way it does for a node — it renders only
|
||||
from the mesh's `addresses` map today, keyed by node name, and does not distinguish a named carried
|
||||
peer from an unnamed one. `mesh-catalog modules/dnsmasq` itself needs no change: it already
|
||||
restarts on the `node-zones` fact and would pick up the new wildcard the moment `internal/catalogue`
|
||||
emits it.
|
||||
|
||||
## What is still open
|
||||
|
||||
- The mechanism for *recording* the name is not designed — a new field on the carried-peer record,
|
||||
a new CLI verb, and what happens if the name later disagrees with what the peer states on
|
||||
enrolling (it should win; nothing says so yet).
|
||||
- Whether this generalises: the predecessor's static hosts file was the answer *here* because it
|
||||
happened to be readable and correct. A migration without one still has the report's harder second
|
||||
option (refuse the resolver until every peer enrols) as its fallback.
|
||||
- Not implemented. This diagnosis is what the fix would touch and why the ADR 0104 framing in the
|
||||
report's third option does not transfer directly — not the fix itself.
|
||||
@@ -0,0 +1,117 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-24
|
||||
located-in: [hal modules/minio]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 113 — The object store's images were withdrawn upstream, and only a node that already holds them can still run it
|
||||
|
||||
## What was observed
|
||||
|
||||
On 2026-09-24, during a service-by-service cutover, the object-store module could not be built on
|
||||
a node that did not already hold its images. Both images the module needs answer an anonymous
|
||||
pull with `401 UNAUTHORIZED`:
|
||||
|
||||
```
|
||||
<registry>/minio/minio 401 <registry>/minio/operator 200
|
||||
<registry>/minio/mc 401 <registry>/minio/console 200
|
||||
```
|
||||
|
||||
The module pins a **tag**, not a digest, and the default is four and a half years old:
|
||||
|
||||
```
|
||||
image: <registry>/minio/minio:${MINIO_VERSION:-RELEASE.2022-01-07T01-53-23Z}
|
||||
```
|
||||
|
||||
The cause is upstream and outside the mesh: the vendor **deleted** the community server and client
|
||||
repositories. It is not an access policy that a credential could answer, and nothing about the
|
||||
mesh's own registry configuration, resolver or trust settings is involved.
|
||||
|
||||
- The vendor removed both repositories from the main public registry on **2026-09-11**. Its API
|
||||
answers `404` for the server repository while a sibling in the same namespace answers `200`.
|
||||
- The secondary registry that the wider ecosystem repointed to as a stopgap **no longer lists
|
||||
them either**. Sixty-eight repositories in that namespace are still public and pull normally;
|
||||
the server and the client are simply absent, and the namespace is now dominated by the vendor's
|
||||
commercially licensed line.
|
||||
- The open-source repository was archived in **February 2026**, and the community edition has been
|
||||
source-only since **October 2025**. No new images are published anywhere.
|
||||
|
||||
## What did not happen, and why it is recorded
|
||||
|
||||
The first reading of this was that the registry had *disabled anonymous pulls for the whole
|
||||
vendor namespace*. That was wrong in a way worth keeping, because the evidence looked conclusive:
|
||||
|
||||
- The anonymous token carries `"actions": []` for the affected repositories and `['pull']` for
|
||||
working ones — a real signal, but it is **also exactly what a repository that does not exist
|
||||
returns**. A deliberately invented repository name in the same namespace produced a
|
||||
byte-identical response. The signal cannot distinguish *revoked* from *absent*.
|
||||
- The token also carries `"$disabled"`, which was read as confirmation. It appears on **every**
|
||||
repository on that registry, including the ones pulling successfully. It describes image
|
||||
**signing**, not access.
|
||||
|
||||
Sibling repositories in the same namespace pulling normally is what rules out a namespace-wide
|
||||
policy, and the registries' own APIs — `404` against `200` — are what establish deletion.
|
||||
|
||||
## The mesh was not blocked, which is the other half
|
||||
|
||||
A node that already holds the images runs the module normally. The node carrying the cutover holds
|
||||
the pinned server image, the client, and the load-balancer image the module composes with, all
|
||||
pulled years ago. Its resolved version variable matches the cached tag exactly.
|
||||
|
||||
This is by design and not by luck. The deploy stage pulls **best-effort** and then asserts only
|
||||
that every image the composition declares **resolves locally**, precisely so that an image which
|
||||
exists on the node but can no longer be fetched does not fail a deploy. The code comment naming
|
||||
the precedent describes this case exactly — *"an old tag pulled years ago and since removed
|
||||
upstream"* — and records that failing on the pull instead had previously made a module
|
||||
undeployable while all of its images sat on the node.
|
||||
|
||||
So the deploy logs a warning and succeeds. Dropping the module from the cutover queue was not
|
||||
necessary.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
The instance is harmless; the standing condition is not.
|
||||
|
||||
1. **No new node can ever provision this module.** Every node that does not already hold the
|
||||
images is permanently unable to obtain them, and the same will be true of any module whose
|
||||
upstream withdraws an image.
|
||||
|
||||
This is the failure mode of a deliberate design choice, which is why it is worth recording
|
||||
rather than patching. The foundation design chooses **references over payload** — *"the bundle
|
||||
names images by digest and the host fetches them"* — on the stated grounds that
|
||||
*"reproducibility comes from pinning the identity of a thing rather than carrying its bytes"*
|
||||
([to-be 07](../../03-DESIGN/01-to-be/07-the-foundation.md)). That reasoning is sound. It holds
|
||||
only while a pinned identity stays **resolvable**, and nothing in the mesh's control guarantees
|
||||
that for an image in somebody else's registry. The passage is about
|
||||
the foundation bundle, and this module is not in it; but the pattern — pin the identity, fetch
|
||||
the bytes on demand — is how every module gets its third-party images, so the exposure is
|
||||
general even though the sentence is local.
|
||||
2. **The pinned release is permanently unpatched.** It is four and a half years old, upstream is
|
||||
archived, and no security fix will ever reach it.
|
||||
3. **Nothing detects this class of failure.** The condition is invisible until a node without the
|
||||
image tries to deploy. The very guard that correctly stops this from breaking existing nodes —
|
||||
assert local resolvability, not the pull — also means a warning is the only trace, and a
|
||||
warning is not a rule. A mesh cannot state that its modules are installable while the only
|
||||
evidence is that they are already installed.
|
||||
|
||||
The third point is the general one, and it is not specific to this vendor: an image pinned by tag
|
||||
against a registry the mesh does not control is a dependency with no guarantee behind it, and the
|
||||
mesh currently learns it has lost one only by trying to use it.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the mesh **hold** the images it depends on — mirroring third-party images into its own
|
||||
registry at adoption, so a module's installability does not depend on an upstream's continued
|
||||
goodwill? That is the fix that generalises. It costs storage and a policy about what to mirror,
|
||||
and it is a deliberate move **away** from references-over-payload for third-party images
|
||||
specifically — so it should be decided as such, not smuggled in as a fix.
|
||||
- Should a module's images be pinned **by digest** rather than by tag? It makes the artifact
|
||||
exact and auditable, but does nothing about withdrawal — a deleted digest is just as gone.
|
||||
- What **checks** that every module in the catalogue is still obtainable from a node that holds
|
||||
nothing? Nothing does today. A periodic cold-pull of the catalogue would have caught this on
|
||||
2026-09-11 rather than thirteen days later, mid-cutover.
|
||||
- For this module specifically: replace the product. The design already says the dependency is on
|
||||
**S3 the protocol, not the product** — see
|
||||
[research 015](../../01-RESEARCH/015-the-object-store-after-minio/00-overview.md).
|
||||
@@ -0,0 +1,110 @@
|
||||
# 113 — Diagnosis
|
||||
|
||||
## 2026-09-24 — the trail
|
||||
|
||||
The symptom arrived already carrying a diagnosis: *the registry has disabled anonymous pulls for
|
||||
the entire vendor namespace.* Everything below was an attempt to confirm that, and it did not
|
||||
survive.
|
||||
|
||||
### Step 1 — the token is not the test
|
||||
|
||||
The reported evidence was the anonymous pull token's contents: `"actions": []` and `"$disabled"`.
|
||||
A token is an intermediate artifact. The test is whether a manifest can actually be fetched with
|
||||
it, so the first step was to request one:
|
||||
|
||||
```
|
||||
GET /v2/minio/minio/manifests/<pinned tag> -> 401 UNAUTHORIZED
|
||||
```
|
||||
|
||||
That confirmed the failure but said nothing about its scope or cause.
|
||||
|
||||
### Step 2 — a control ruled out the stated cause
|
||||
|
||||
The same request flow, in the same minute, against other repositories:
|
||||
|
||||
| repository | token actions | manifest |
|
||||
|---|---|---|
|
||||
| the vendor's server | `[]` | **401** |
|
||||
| the vendor's client | `[]` | **401** |
|
||||
| the vendor's operator | `['pull']` | 200 |
|
||||
| the vendor's console | `['pull']` | 200 |
|
||||
| the vendor's sidecar proxy | `['pull']` | 200 |
|
||||
| an unrelated public project | `['pull']` | 200 |
|
||||
|
||||
**A namespace-wide policy is ruled out.** Two repositories fail; their siblings in the same
|
||||
namespace pull normally.
|
||||
|
||||
### Step 3 — both pieces of the original evidence were red herrings
|
||||
|
||||
- `"$disabled"` is present on **every** repository on that registry, including all of the
|
||||
successful ones above. It belongs to the signing context, not to authorisation. It carries no
|
||||
information about this failure at all.
|
||||
- `"actions": []` with a `401` is **indistinguishable from a repository that does not exist.** A
|
||||
deliberately invented repository name in the vendor's namespace returned a byte-identical
|
||||
response — empty actions, `401`. The signal cannot separate *access revoked* from *not there*,
|
||||
so it cannot support the conclusion it was used for.
|
||||
|
||||
That second point turned the question from *who revoked access* to *is it still there*.
|
||||
|
||||
### Step 4 — the registries' own APIs establish deletion
|
||||
|
||||
Asked directly, rather than through the pull path:
|
||||
|
||||
- **Primary registry:** its repository API answers **404** for the server repository, and **200**
|
||||
for a sibling in the same namespace. The repository is gone, not private.
|
||||
- **Secondary registry:** a listing of the vendor's namespace returns **68 public repositories**.
|
||||
The server and the client are **absent from the list**. Present are the operator, console,
|
||||
sidecar proxy, benchmarking and key-management images — and a large, newer set belonging to the
|
||||
vendor's commercially licensed line.
|
||||
|
||||
### Step 5 — upstream confirms, and dates it
|
||||
|
||||
The vendor deleted the community server and client from the primary registry on **2026-09-11**.
|
||||
This was the last step of a staged withdrawal: free image publishing stopped in **October 2025**,
|
||||
the community console UI was removed mid-2025, and the open-source repository was archived in
|
||||
**February 2026**. The secondary registry was where the ecosystem repointed as a stopgap; it has
|
||||
since lost the two repositories as well.
|
||||
|
||||
**Conclusion: the images were withdrawn, not restricted.** No credential can answer this, because
|
||||
there is nothing left to authenticate against. A different image source is the only remedy.
|
||||
|
||||
## Why the mesh kept working, checked rather than assumed
|
||||
|
||||
The claim that the module "genuinely isn't ready for cutover" was tested and is false for the node
|
||||
in question.
|
||||
|
||||
- The node **holds** the pinned server image, the client, and the load-balancer image the module
|
||||
composes with — all pulled years before the withdrawal.
|
||||
- The module's resolved version variable on that node matches the cached tag **exactly**, so the
|
||||
composition references an image that is present.
|
||||
- The composition declares **no pull policy**, so a present tag is used as-is.
|
||||
- The deploy stage pulls best-effort, then asserts only that every declared image resolves
|
||||
locally. A pull error becomes a logged warning when all images are present.
|
||||
- The start path restarts the unit and does not pull. The one `--pull always` in the tree sits
|
||||
inside a generated guide describing the **superseded** approach, not in the code that writes
|
||||
units.
|
||||
|
||||
So a deploy of this module on that node succeeds today.
|
||||
|
||||
## Claims in the original report that could not be substantiated
|
||||
|
||||
Recorded because they were specific and load-bearing, and acting on them would have wasted time.
|
||||
|
||||
| Claim | Finding |
|
||||
|---|---|
|
||||
| The module's manifest is a `module.json`, pinning the server image by digest at line 83 | There is **no `module.json` anywhere** in the monorepo. The manifest is YAML and pins a **tag**. No digest pin exists. |
|
||||
| The module's own runtime artifact is a placeholder with an all-zero digest | **Zero occurrences** of that image name or of an all-zero digest anywhere in the tree. The module declares no runtime or sidecar artifact. |
|
||||
| A dated readiness document records six modules with placeholder digests | **No such file exists.** |
|
||||
| The server image is cached locally at an older release than the module pins | The cached tag is **exactly** the pinned one, not an older release. |
|
||||
|
||||
None of these change the real finding, which stands: the images are gone upstream.
|
||||
|
||||
## What is located, and what is not
|
||||
|
||||
**Located:** the module in the monorepo's catalogue — it pins, by tag, an image that no longer
|
||||
exists anywhere public.
|
||||
|
||||
**Not located, and deliberately left open:** the general condition. The mesh has no mirror of the
|
||||
third-party images its modules depend on and no check that a module is obtainable by a node
|
||||
holding nothing. That is a design gap rather than a defect in this module, and it is stated as an
|
||||
open question on the report rather than answered here.
|
||||
+73
@@ -0,0 +1,73 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-24
|
||||
located-in: [mesh-controller module.json, mesh-host internal/apply]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 114 — Should the controller run as a container, or as a process the host supervises directly?
|
||||
|
||||
## What was observed
|
||||
|
||||
On the control-node, 2026-09-24, over a long session of operating the mesh through
|
||||
`mesh-controller`'s CLI (build, push, plan, status, module moved). Every mutating step reached the
|
||||
binary the same way: `docker exec mesh-controller /mesh-controller <command>` — because
|
||||
`mesh-controller`'s own manifest declares its one resource as:
|
||||
|
||||
```json
|
||||
{ "id": "server", "type": "container", "name": "mesh-controller", "network": "host", "args": ["serve"] }
|
||||
```
|
||||
|
||||
Two things about that declaration are worth naming together, because neither is a problem on its
|
||||
own and the combination is what raises the question:
|
||||
|
||||
- **`network: host`.** The controller does not use container network isolation, which is the
|
||||
property a `container` resource type usually buys over a `process` one. It runs with the node's
|
||||
own network namespace either way.
|
||||
- **It is the mesh's single point of coordination.** [`03-DESIGN/01-to-be/06-the-controller.md`](../../03-DESIGN/01-to-be/06-the-controller.md)
|
||||
is explicit: "one node runs it, and nothing takes over" — no election, no quorum, no failover;
|
||||
recovery is restore, not failover.
|
||||
|
||||
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) gives the host — the one thing tier 0 requires
|
||||
to be a real system daemon — exactly this reasoning for refusing to run in a container: *"installing
|
||||
the container runtime is a step of the bootstrap, so a host inside a container would need the thing
|
||||
it exists to install."* The controller is one tier up and does not install the runtime, but it
|
||||
shares the profile that argument turns on: something the rest of the mesh's operation depends on,
|
||||
sharing fate with a runtime that is not itself.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
Practically, tonight: every controller interaction was raw shell into a container (`docker exec`),
|
||||
not a first-class surface — no logs command beyond `docker logs`, no `systemctl status`, and a
|
||||
session permission classifier that (correctly) treats arbitrary shell into a container as needing
|
||||
sign-off every time, unlike an ordinary supervised process. That friction is a symptom, not the
|
||||
issue itself.
|
||||
|
||||
The actual question is whether `type: container` is buying the controller anything here besides
|
||||
image-based delivery and a restart policy — both of which [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)'s
|
||||
launcher pattern already describes as buildable directly into the host's own supervision (restart on
|
||||
exit, count consecutive failures, roll back after too many, halt after that), for the host's own
|
||||
unit. If the controller were declared `type: process` instead — still built and versioned through
|
||||
the same delivery pipeline, just executed on the node and supervised by the host the way the host
|
||||
supervises itself — it would stop sharing fate with the container runtime's health (restarts,
|
||||
upgrades, disk pressure evicting containers) for the one piece of software whose absence the rest of
|
||||
the mesh is designed to tolerate but nothing is designed to *want*.
|
||||
|
||||
This is squarely a question, not a claim that today's shape is wrong: [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)
|
||||
already tolerates the controller being down by construction (nodes reconcile from their own
|
||||
last-applied state), which may make the container-runtime coupling moot in practice. Nobody has
|
||||
checked.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Does `mesh-host`'s `process` resource type already support the restart/failure-counting semantics
|
||||
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) describes for the host's own launcher well
|
||||
enough for something this central — or would this need host-side work first?
|
||||
- With `network: host` already in use, what does `type: container` provide the controller today that
|
||||
`type: process` would not?
|
||||
- Is there a real circularity risk — the controller's own health depending on the container runtime
|
||||
it (indirectly, via the host) manages — or does "one node runs it, nothing takes over" already make
|
||||
a controller outage tolerable regardless of which resource type it is?
|
||||
- If the answer is "keep it a container," what does that answer, precisely, that this issue asked —
|
||||
so the next person who notices the same asymmetry finds it answered rather than open again?
|
||||
Reference in New Issue
Block a user