Commit Graph
5 Commits
Author SHA1 Message Date
jschoubben 6bdf7104c8 whole-mesh-novox: the postgres server container is mesh-store, not postgres
The CORE convergence wait hung on a container named 'postgres' that
never exists — the postgres module's server is the adopted-store
container 'mesh-store' (like lavinmq's mesh-broker). Everything else
converged; this was the last phantom-name blocker.
2026-09-20 22:06:26 +02:00
jschoubben 0a05fbb434 whole-mesh-novox goes green: the store superuser, the artifact shape, and the missing CA
The bed had never resolved, then never converged. Fixed, in order:
- loadManifest maps a runtime container's 060 `artifact` to the stocked
  mesh-runtime-<module> image (keyed on the module name), so the push is
  no longer refused by built() — and drops the build section.
- Stale identities renamed: registry->distribution, firewall->nftables.
- step-ca added to the set: the web modules hard-require `route`,
  route-proxy provides it but requires `acme-ca`, and nothing provided
  that — so the whole web stack never resolved. step-ca is the missing CA.
- THE STORE SUPERUSER is delivered via `secret accept` before the push.
  postgres raises mesh-store with POSTGRES_PASSWORD=bootstrap, but
  `module add` minted a random superuser own-secret that did not match,
  so the provisioner could not log in and created NO consumer roles —
  every DB consumer (gitea/keycloak/nextcloud/umami/mailu) failed. This
  was the real cause behind what looked like per-module gaps; keycloak
  and umami converge once it is delivered (ADR 0078, hq phase3).
- mesh() retries through the controller recreating itself during the
  057 cascade (No such exec instance), so a real success is not read as
  a failed push.
- invoicing dropped (private-registry images the lab cannot pull).

Remaining KNOWN_GAPS are genuine catalog/upstream/resource gaps: minio
(stale Docker Hub digest), mssql (Error 945, memory), mailu (config
env), photos (alpine placeholder), nftables (service).
2026-09-20 22:06:26 +02:00
jschoubben 5d6e8fbe7a Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 675facdb0d The beds name images the way a machine would find them
Twenty-eight integration tests each carried their own copy of the same two helpers,
which pointed a manifest and the substrate bundle at whatever the lab's registry had
assigned. They now share two in the harness, and the difference is the point: ours is
rewritten to the ID the machine holds it under, and everything else is left exactly as
written so the machine pulls it.

**The substrate bundle is where the fiction was most load-bearing.** mesh-host's
`examples/substrate-first-node.lock` pins all three of its images at
`192.0.2.250:5000/…`, which is the address the lab's registry served from — it was
written for a target, and the target was the lab. Two of those are ordinary third-party
images and become the digests mesh-catalog's own postgres and lavinmq modules pin, so
the substrate's store and broker are literally the images the mesh runs. mesh-control
exists in no registry at all and becomes the ID the machine was handed. **The bundle
itself should be fixed in mesh-host and this substitution deleted with it.**

Beds that wrote a manifest by hand named an image by repository and let the rewrite
supply a digest. There is nothing to supply one now, so `onTheMachine` refuses an
unpinned reference and hands back the digest the catalogue pins — a bed runs the image
the mesh ships, and a bed that drifts from the catalogue is testing a different
postgres.

Three beds took a third-party image out of the raised list, which no longer contains
one: certificates (pebble), objectstore (minio and its client) and provisioner
(postgres) now name theirs and pull it. builds and mesh publish into the MESH's own
artifact store — the `registry` module's image, on the node, on 5000 — rather than into
scenery the lab raised. That is a different claim, and only one of them exists in
production.

New unit tests cover what a full raise would otherwise be the only way to check: the
routes an egress machine gets (that its gateway is still the path to the rest of the
scenario, that a range with no path is unreachable rather than leaked to the uplink,
that each family gets its own next hop), which machine is handed which of our images,
and the `images:` rule that refuses a third-party entry. The "shipped scenarios are
valid" test now loads every scenario rather than two of them.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-10 23:16:41 +02:00
jschoubben 581fd6da77 Add whole-mesh novox dry-run bed (stage 1 of whole-mesh rehearsal)
Install the real novox server's converted service set together on one node
behind the substrate — the whole-catalogue install this rebuild never ran.
The bed loads each committed module.json from mesh-catalog (no hand-written
manifests), rewrites image refs to the scenario registry's digests, and
remaps the co-located host-port collisions (nextcloud/invoicing/route-proxy
:80, minio/invoicing :9000, gitea/umami :3000).

Proven green: the whole set of 17 modules RESOLVES and applies (191
resources); the CORE 13 converge whole — all five providers (postgres,
redis, minio, mongodb, mssql) plus keycloak, gitea, nextcloud and invoicing
reaching their providers and staying up, plus portainer, verdaccio, registry
and route-proxy.

Reported as escalated gaps (do not gate green): fail2ban (declares
capability intrusion-prevention that no host detector provides, and an
unappliable assignment blocks whole-node resolution), umami/photos/mailu
(catalog manifests do not wire the runtime/app env the images need; photos'
server image is an alpine placeholder), and firewall (nftables.service is a
oneshot that exits, but the module declares state running so mesh-host marks
it failed).

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-07 22:54:10 +02:00