Two-node DB-consumer bed + a scenario disk field

The GREEN multi-node regression bed that proves the DB-consumer gate: substrate/control on
one node, postgres+redis providers and baserow+letta consumers on another, each consumer
getting its own credential and its own mesh-named database across the overlay. Requires the
mesh-control provider-seal-key fix and the mesh-catalog db-name fix.

Includes a general lab capability: a machine 'disk' field sizing the VM root disk (a broad
install exhausts the pool default and the host fails mid-apply with 'no space left on
device'). The bed sets 60GiB.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
2026-09-06 13:48:26 +02:00
parent b66b330dba
commit a9ecce25cb
5 changed files with 635 additions and 1 deletions
+62
View File
@@ -0,0 +1,62 @@
# The DB-consumer chain a single node cannot host, proved across two machines.
#
# The app-postgres provider and the mesh's own substrate store both want host port 5432, so they
# cannot share a machine — the collision that blocked this chain single-node. Here the substrate
# (store, broker, control) lives on `anchor` and NOTHING else; `laptop` runs the whole chain —
# postgres and redis PROVIDERS plus the baserow and letta CONSUMERS that require them. Both
# machines sit on one shared segment and enrol into the one mesh; only enrolment crosses to anchor,
# over the underlay both machines already share. Provider and consumers are co-located on laptop, so
# no cross-node module comms and no overlay are needed — and the 5432-vs-substrate conflict is gone
# because the substrate store is on the OTHER node.
scenario: two-node-db
segments:
hosting:
kind: public
cidr: [192.0.2.0/24]
machines:
# The substrate ONLY: store, broker, control — three containers. Four gigabytes is plenty for a
# node that hosts no modules; the thrash the two-nodes bed warns of comes from stacking eleven
# containers on a node, which this one never does.
anchor:
at: { segment: hosting, address: [192.0.2.10] }
inbound: allow
memory: 4GiB
cpus: 4
# The whole DB-consumer chain: postgres + redis providers, each a server and a broker-bound
# runtime, plus the baserow and letta consumer services and their tools runtimes — a dozen
# containers, two of them memory-hungry app servers (the Baserow all-in-one and the Letta server).
# At the 2GiB the two-nodes bed gives this machine it would thrash — its own anchor comment says
# so — and convergence would present as "the mesh hangs". Six gigabytes gives it room.
laptop:
at: { segment: hosting, address: [192.0.2.20] }
inbound: allow
memory: 6GiB
cpus: 4
# The runtime images (280MB–600MB each) plus the service images — two of them heavy app images
# (Baserow ~1.5GB, Letta ~1.8GB) — are pulled from the scenario's own registry by digest, so the
# same bytes land on this node twice. The pool default root disk exhausts mid-apply ("no space
# left on device"); sixty gigabytes holds the whole chain.
disk: 60GiB
images:
# The first-node substrate: store, broker, control. postgres:17-alpine doubles as postgres's own
# service image.
- postgres:17-alpine
- cloudamqp/lavinmq:latest
- mesh-control:development
# The module service images.
- redis:7-alpine
- baserow/baserow:latest
- letta/letta:latest
# The per-module runtimes, built by scripts/build-module-runtime.sh and stocked here. Each carries
# its module's provisioner, so no separate mesh-provision-* image is listed — the runtime is the
# provisioner (ADR 0048).
- mesh-runtime-postgres:development
- mesh-runtime-redis:development
- mesh-runtime-baserow:development
- mesh-runtime-letta:development
place:
all: [host, runtime]