Files
mesh-lab/scenarios/two-nodes.yml
T
jschoubben 4a343a2652 A machine is as big as the scenario says, and may reach the world
Three changes, found by one failing test.

The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.

So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.

`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.

And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
2026-09-01 21:42:18 +02:00

60 lines
2.7 KiB
YAML

# Two machines, one mesh.
#
# The first raises everything from the bundle its host carries and joins the mesh it made. The
# second is an ordinary node: it has a host and nothing else, and a person carries it a token.
#
# This is the first scenario where the mesh is a mesh. Everything before it proved a machine could
# talk to a control plane on its own loopback, which proves less than it looks.
scenario: two-nodes
segments:
hosting:
kind: public
cidr: [192.0.2.0/24]
machines:
anchor:
at: { segment: hosting, address: [192.0.2.10] }
inbound: allow
# The whole substrate, the registry, the builder, an adopted workload and the modules under
# test all land here — eleven containers before the forge arrives. At the 1GiB default this
# machine thrashes, and it presents as "the mesh hangs": every exec slows from 15s to 105s
# and the forge test fails on a status poll that is merely queued behind page-outs.
memory: 4GiB
cpus: 4
laptop:
at: { segment: hosting, address: [192.0.2.20] }
inbound: allow
memory: 2GiB
images:
- postgres:17-alpine
- cloudamqp/lavinmq:latest
- mesh-control:development
# So a module can mirror one into a registry of the mesh's own. The scenario's registry serves
# what the mesh's registry is built from — the same chicken-and-egg the bootstrap has, resolved
# the same way.
- registry:2
# A real third-party workload, for adopting one the way the conversion will. Its database is
# the substrate's postgres image rather than its own: what is under test is the mesh delivering
# a module, not which postgres it delivers.
- ghcr.io/umami-software/umami:postgresql-latest
# And the builder, because it is a module the mesh assigns rather than a program somebody
# starts by hand — which is the only way its credential can be one the mesh delivered.
- mesh-builder:development
# And the provisioner, which is what makes a sealed credential true on a machine — the mesh
# discarded the plaintext and cannot tell a database to start accepting it.
- mesh-provision-postgres:development
# And the proxy, which is what turns a route grant into traffic actually arriving.
- mesh-route-proxy:development
# And the object store's provisioner, so the module describing it can be planned. Without it
# that module still names an image nothing serves, and planning it is refused — correctly.
- mesh-provision-objectstore:development
# And a forge, so one of the real module descriptions can be started rather than only planned.
# It is the first of them to run: it needs a database from another module, a credential it did
# not choose, and a connection string it could not have written itself.
- gitea/gitea:1.22
place:
all: [host, runtime]