The control plane's image is FROM scratch and runs as 65534, and docker cp keeps
the mode a file had outside — openssl writes a private key 0600 root-owned, so
the copy landed unreadable, secret accept failed with permission denied, and the
CA crash-looped on a root it never got. Chowning it inside the container is not
available: there is no shell in there to do it with.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The bed set no node a `public-domain` and assigned no `acme-ca` provider, so it
was testing a mesh the design no longer describes — and going green while doing
it, which is the worse half.
**No public domain means no route.** A module now contributes a `label` and
nothing else; the mesh joins it to the node's public domain, and a label with no
domain to join composes to nothing at all. Every routed module on this bed was
therefore unreachable by name, silently, and no assertion noticed. novox now
carries `novox.incus` and ace `zurag.incus` — `.incus`, because this repository's
beds name nothing routable. The workstations carry none, which is also the design
being exercised: a node that does not face outward has no public domain.
**No acme-ca provider means no proxy.** route-proxy requires one, so without a
provider it is unresolvable and takes every routed module with it. step-ca is
assigned on the anchor, at mesh scope, and given an operator root — made with
openssl on the anchor and handed over through the real `secret accept` path,
because the mesh cannot invent a PEM and the random bytes it makes for an
own-secret nobody supplied would leave the CA crash-looping on a root key that is
not a key.
**What is asserted is the half that is decided and cheap**: that each routed
module's name composes to `<label>.<public-domain>` — read from the proxy's own
received-routes file, the mesh's answer on the machine rather than this test's
arithmetic checked against itself — with `@` composing to the bare domain, and
that the proxy answers for one of them over HTTP.
**What is NOT asserted is issuance.** Whether route-proxy obtains a certificate
from step-ca over ACME depends on mesh-control fixes landing as this is written,
and a bed that gated on them would report somebody else's in-flight work as its
own failure. step-ca is listed as a reported gap for the same reason.
The substrate apply also retries up to three times. `raise` now refuses to return
until every machine can fetch a manifest from the scenario registry, so the first
attempt should be the only one; a pull is simply the one step here that can fail
for a reason that goes away by itself, and the cost of not retrying was a whole
raise left as a bare shell.
Typechecks; not run end-to-end — see the ADR 0056 section for what is expected to
fail until the issuance path is fixed.
Rewrite the flat three-node whole-mesh-full (separate anchor, one public segment)
into production's real shape: two segments and one access point. novox sits on
the routable `hosting` segment and IS the anchor — it runs the substrate, its own
service set, the overlay hub and public ingress; there is no separate anchor node.
ace, shanks and g14 sit on the household `home` segment behind a NAT gateway,
reachable from outside only through what they dial out to.
The bed drives, and verifies, the thing the flat beds never could: the WireGuard
overlay forming ACROSS the access point — a home node dialling novox's public hub
endpoint out through the gateway's masquerade, the handshake completing through the
NAT, the keepalive holding the hole open. Phase A proves it (handshake state + a
ping over the overlay) before any heavy module lands; Phase B converges both server
sets. With MESH_LAB_KEEP the instance is raised under a fixed id and left standing.
Collapsing the substrate onto novox exposed real facts the separate-anchor beds
never hit, fixed here:
- the substrate bundle advertises the broker at 192.0.2.10 (the old anchor); a
token carries that verbatim as the endpoint a node dials, so with the substrate
on novox it must be novox's own public address. Rewritten at apply (the cert is
fingerprint-pinned, not hostname-checked, so only the address needs correcting).
- the two provider host-port collisions with the co-located substrate: postgres
5432 vs the store's 127.0.0.1:5432, lavinmq 5672 vs the broker's 127.0.0.1:5672.
Both provider host publishes are remapped off the substrate's ports.
And a lab limitation this first large-union bed exposed: the image registry VM took
the profile's default `dir` pool and a ~10GiB root, which the ~28GiB union of both
server sets overflows ("no space left on device"). raiseRegistry now places the
registry on the scenario's copy-on-write pool with a sized (default 80GiB, thin)
root disk, MESH_LAB_REGISTRY_DISK overridable.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Re-runs the capstone from main after the dry-run fixes merged.
fail2ban: added to the novox set. The capability fix (intrusion-prevention ->
firewall) makes it HOSTABLE — it is now assigned, not refused — which is the
gate. Its service reaching active is a host concern the offline lab cannot meet
(the VM ships nftables but not fail2ban, and the isolated segment has no route to
the package mirror, so pacman cannot fetch it), so fail2ban joins GAPS_NOVOX: its
failed package resource is tolerated like firewall's oneshot nftables.service.
7 credential sidecars: before the push, a FAKE app credential is delivered for
each (plex/bazarr/ombi/home-assistant/nzbget/qbittorrent on ace, umami on novox)
through the real operator path — `secret accept <node> <module> <name> --from`.
The bed asserts each sidecar advances PAST its old "no credential" crash (it reads
the delivered value); app-auth failure against the real app with a bogus value is
expected and not gated.
Result: SUITE_EXIT=0. Both node-plans converge on one substrate (novox 13/13
core, ace 17/17 core), fail2ban hostable, all 7 sidecars past their crash.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Combine the novox (17-module) and ace (24-module) sets on ONE substrate and
prove both node-plans converge together. anchor runs the substrate only;
novox and ace each run their own self-contained set (own postgres/redis), so
nothing crosses a node boundary except enrolment and the shared broker/store.
The four modules both nodes run (postgres, redis, mssql, portainer) are added
once and assigned to each node, each getting its own per-node broker account.
An overlay is placed across all three nodes.
Proven green: both nodes converge together on the one substrate. ace reaches
applied+current with all 17 of its CORE up (and letta too this run); novox
reaches all 13 CORE up with its only failed resource the known firewall.load
oneshot gap. The two node-plans share one broker without collision — distinct
novox-<mod> and ace-<mod> accounts for the modules both run. No new cross-node
bug (overlay/DNS/identity/port) surfaced; ports are per-VM and the sets are
node-self-contained. Tolerates the same nine credential-sidecar gaps and
firewall's nftables.service oneshot documented in the per-server beds.
Resource envelope: 3 VMs (anchor 4GiB, novox 16GiB, ace 18GiB) + registry
scenery, ~79 union images (~35GB) stocked to one registry VM and pulled
concurrently by both nodes; fit within 125GiB host RAM and the 180GiB lab pool.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Install the real ace server's converted service set (24 modules) together on
one node behind the substrate — sibling of the whole-mesh-novox bed, the
media/home-automation half. Loads each committed module.json from
mesh-catalog, rewrites image refs to the scenario registry's digests, remaps
the co-located host-port collisions (qbittorrent/searxng/unifi :8080,
nzbget/unifi :6789), and pre-creates the ADR-0051 operator-owned media
library dirs under /services/media so the media stack's `accesses` resolve.
Proven green: the whole 24-module set RESOLVES and applies (214 resources,
node applied+current) — the ADR-0051 shared-dir `accesses` mechanism works
cleanly across eight co-accessing media modules. The CORE 17 converge whole:
postgres/redis/mssql, sonarr/radarr/lidarr/jackett/tautulli/bookshelf,
mosquitto/influxdb/grafana/baserow/nodered/searxng/unifi/portainer.
Reported as escalated gaps (do not gate green): six tool-runtime sidecars
crash-loop because the committed manifest does not wire the app credential
they need (plex MESH_PLEX_TOKEN, bazarr MESH_BAZARR_API_KEY, nzbget
MESH_NZBGET_URL/PASSWORD, qbittorrent MESH_QBITTORRENT_URL/PASSWORD, ombi
MESH_OMBI_API_KEY, home-assistant MESH_HOMEASSISTANT_TOKEN) — the umami/photos
class from novox; each server is up, only the sidecar is down. sonarr/radarr/
lidarr/jackett/tautulli self-configure from the app's config file and their
runtimes come up. letta's app has a first-boot postgres migration race
(pgvector the deeper blocker, per two-node-db).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Install the real novox server's converted service set together on one node
behind the substrate — the whole-catalogue install this rebuild never ran.
The bed loads each committed module.json from mesh-catalog (no hand-written
manifests), rewrites image refs to the scenario registry's digests, and
remaps the co-located host-port collisions (nextcloud/invoicing/route-proxy
:80, minio/invoicing :9000, gitea/umami :3000).
Proven green: the whole set of 17 modules RESOLVES and applies (191
resources); the CORE 13 converge whole — all five providers (postgres,
redis, minio, mongodb, mssql) plus keycloak, gitea, nextcloud and invoicing
reaching their providers and staying up, plus portainer, verdaccio, registry
and route-proxy.
Reported as escalated gaps (do not gate green): fail2ban (declares
capability intrusion-prevention that no host detector provides, and an
unappliable assignment blocks whole-node resolution), umami/photos/mailu
(catalog manifests do not wire the runtime/app env the images need; photos'
server image is an alpine placeholder), and firewall (nftables.service is a
oneshot that exits, but the module declares state running so mesh-host marks
it failed).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A single-node VM bed: an ollama provider and a local-model consumer are
assigned; the resolver answers the consumer's model-access with the local node
(no licence demanded), the consumer's openai.env is templated with the served
endpoint, and a request to it reaches the running model server. Proves the
node-answer of model-access end to end (ollama on host network, keyless).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A single-node VM bed: the operator sets an API key on an openai licence, the
mesh seals it to the consumer, the host unseals and mounts it, and the consumer
writes it as OPENAI_API_KEY (env + Codex auth.json). Asserts the written key
equals the one set — the other shape ADR 0050 defines, and the ADR 0054 branch
where a static-key vendor records no usage. Green on the first run.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A VM bed: a postgres provider and the model-usage consumer on one node, the
substrate on the other. A usage event injected into the mesh is upserted into
model-usage's provisioned store, asserted at both grains, latest-per-key, and
in the clear.
Also, in build-module-runtime.sh, add migrate/index.ts and pg.d.ts to the
compiled entrypoint set so a module may carry a run-once entry and an ambient
type declaration (model-usage uses the latter for the pg driver).
The bed surfaced and drove several fixes elsewhere: a short module slug for the
S3-key identity bound (ADR 0049), host-network containers getting the mesh's
names (mesh-control), and injecting the event from a publisher rather than the
pure-consumer store (its account has no publish right by design).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Two harness fixes the green end-to-end run needed:
- build-module-runtime.sh installs a module's non-@novox runtime deps under
/app/modules/<module>/node_modules, so a module can carry a private dependency
(the anthropic-manager seals with tweetnacl-sealedbox-js). The shared tree still
answers @novox/* and common packages. A no-op for modules that declare none.
- stageIntoControl chmods the manager's 0600 adopt/refresh outputs to 0644 on the
anchor host before docker cp, so the distroless mesh-control (non-root, no chmod)
can read the staged file. What is staged is a sealed box or the access token,
never a cleartext refresh token.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The bed follows the reworked flow: the manager module seals the refresh token to the node's
PUBLIC key, the HOST unseals it and mounts the cleartext at the manager's bound path, and the
refresh reads that cleartext -- no fake node key pair is mounted any more, the host uses its
own real sealing key.
- the manager is a model-access holder deployed first, so its bound facts (carrying the node
public key) are delivered; the consumer is added only once an access token exists to seal.
- adopt reads the node public key from the bound facts; the test asserts the host mounts the
cleartext refresh token for the manager, and that it reaches nowhere on the consuming node.
- the refresh_grant assertion reads { sealed, manager_key }.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A lab bed for Phase C of model-access (ADR 0050), OAuth endpoint stubbed.
It drives the real runtime images through the whole flow: the manager
seals a refresh token at rest and opens it on the manager node alone,
mesh-control is handed only the access token and an opaque re-sealed
envelope via licence submit-refresh, and the consumer writes an
access-token-only credential. Asserts the refresh token -- original and
rotated -- is nowhere on the consuming node and only ciphertext in the
control plane's database.
build-module-runtime.sh also compiles adopt/refresh/apply/usage
entrypoints. Stubbed and flagged: the vendor endpoint, the manager node's
private key (mounted; a host capability to deliver it does not exist
today), and the submit transport (the test invokes the CLI on the
manager's output).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Two-node: substrate broker on anchor, lavinmq provider + amqp-ping consumer on laptop. The
consumer gets its scoped vhost+user, connects, and round-trips a message. Requires the
mesh-control require-only-mint fix. Diagnostic removed now it's green. SUITE_EXIT=0.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The lavinmq AMQP provider comes up and serves, but its receives file has given:[] — the
consumer amqp-ping (requires amqp, contributes nothing, as a parameterless provision like
redis-cache takes no per-consumer payload) is never minted a credential, so the provisioner
creates no vhost. Diagnostic in the test dumps the empty grants file + the provisioner log.
This is the resolve/plan mint path (grantsFor -> SecretsFrom), ADR 0048 territory, and it
likely affects redis-cache consumers the same way. Preserved for a focused fix; not merged.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A lavinmq provider and an amqp-ping consumer ride laptop while the
substrate's own broker owns 5672 on anchor — the twin of two-node-db.
lavinmq is the mesh's control broker AND a user-facing capability, so a
provider must publish 5672 for its consumers and cannot share a node
with the control broker that already owns it; the split unblocks the
chain single-node.
The bed proves, layered: the run-once bootstrap computed the admin hash
and wrote the broker config before the broker started (ADR 0052); the
service and both runtimes are up and stable; each module got its scoped
broker account on the substrate broker; the provisioner created the
consumer's vhost AND user, both named for the derived login; and the
consumer connected to that vhost with the mesh-minted password and
round-tripped a message. The consumer uses ${bound:amqp:as} for user
and vhost both, and the provider's serves carries the port.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A scenario and integration test assign route-proxy (provider) and hello-web
(consumer) on one node, then assert a request to the consumer's name -- sent to
the proxy -- is forwarded to the workload and returns its answer, and that
unassigning the consumer withdraws the route so the same request stops working
(the proxy replaces its table rather than merging). Modeled on
mesh-grant-end-to-end and schedule-tick: module add, assign, one push, settled,
with no module issue (route-proxy needs no scoped account).
build-route-proxy-image.sh compiles the Go proxy from
mesh-control/examples/route-proxy into mesh-route-proxy:development for the
scenario to stock. This bed proves route-forwarding over plain HTTP;
public-ACME TLS is proven separately by certificates.test.ts against a real
ACME server.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The bed that proves the scheduled-container primitive end to end. schedtest
is the thinnest carrier of ADR 0053: one container marked
schedule: "* * * * *" that appends a timestamp to a mounted data dir each
time the host fires it -- no service, no listener, no provisioner, no
runtime, no tools, no events.
The three claims it proves, from the ADR's "How each claim is checked":
installing the schedule leaves the node current WITHOUT a run (baseline
captured right after settled, the deliberate inversion of run-once); the
container fires on its cadence (a line beyond the baseline within ~150s);
and it recurs (a second line on the next minute -- cadence, not a one-shot).
schedtest serves and consumes nothing and carries no runtime, so it is not
issued a broker account: module add -> assign -> one push is the whole
sequence, no module issue. The tick image is a bare alpine served by the
scenario's registry by digest.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The GREEN multi-node regression bed that proves the DB-consumer gate: substrate/control on
one node, postgres+redis providers and baserow+letta consumers on another, each consumer
getting its own credential and its own mesh-named database across the overlay. Requires the
mesh-control provider-seal-key fix and the mesh-catalog db-name fix.
Includes a general lab capability: a machine 'disk' field sizing the VM root disk (a broad
install exhausts the pool default and the host fails mid-apply with 'no space left on
device'). The bed sets 60GiB.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Sixth green regression bed. Same shape as tools-gitlab: a runtime-only module comes
up under the mesh, serves its full tool surface with no valid credentials (the Servarr
lesson), stays up, binds its serve queues, and gets its scoped broker account.
SUITE_EXIT=0.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Prove gitlab — the exemplar tools-only, outbound-only external-SaaS
integration — installs: a scenario assigning gitlab to one node, and a test
asserting the mesh-runtime-gitlab container comes up and stays up, logs
[mesh-tools] serving 23 tool(s), binds its serve queues on the broker, and
gets its own scoped account — all with NO valid GitLab token, the case the
Servarr lesson is about.
No gitlab arm is needed in build-module-runtime.sh: gitlab speaks HTTP and
needs no extra CLI in the image.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A bed that assigns mosquitto and asserts the run-once step seeded dynsec
before the broker: the bootstrap ran to completion (not left running), the
seed is on disk owned by the broker's uid, the broker is up and stable
(it crash-loops against an unseeded store, so a stable broker is the proof),
and the node reached current. On top, the seeded admin authenticates over
MQTT and the provisioner grants a scoped client a consumer connects with.
build-module-runtime.sh gains a mosquitto arm (install mosquitto_ctrl from
the mosquitto package — it is not in mosquitto-clients on bookworm, and a
musl binary from eclipse-mosquitto would not load) and compiles the module's
bootstrap/index.ts entrypoint.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
sonarr and radarr both access /services/media/downloads — the exact duplicate path
the resolver refused before novox/hq ADR 0051 (04-ISSUES/036, 012). Each now declares
it as an `access`, not a `directory` resource, so the pair co-resolves and one push
configures both. The operator provides the shared media dirs before apply (the host
refuses an absent access); the bed creates them after enrol and before the push.
Proves: the push is not refused, the node converges once, both modules' server and
runtime containers are up, and both server containers mount the same operator-owned
spool.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Adds the mongodb runtime CLI (mongosh) to build-module-runtime.sh, a
catalogue-apps scenario, and its install test. Proven so far: the ADR-0054 slug
applies and the mesh accepts the push (mongodb consumer identity mesh_anchor_mongo
fits). NOT green: the node applies but never reaches 'current' within 1200s — a
persistent reconcile divergence (applied-but-never-current, no crash), likely a
module declaring a resource its container mutates (issue-011 class). Needs live
VM inspection to name the module. Not merged.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Raises the first-node substrate and assigns postgres, minio, redis and
plex to one anchor in a single push, proving they resolve and come up
together on one node. postgres and minio each get a consumer that
connects with a real granted credential.
redis follows the corrected provider contract (ADR 0048, issue 032): its
runtime reconciles the contributions the mesh delivers at MESH_RECEIVES
and creates each consumer's ACL user with the mesh-minted password,
sealing nothing — no MESH_SEAL_KEY, no *.grant.json/*.credential path.
The provisioning proof authenticates as the consumer with the mesh's
password (PONG), matching the green provider-uses-mesh-credential bed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Issue 010 fixed: bucketuser declares slug `bkt`, so its identity mesh_anchor_bkt (15)
fits an S3 access key where mesh_anchor_bucketuser (22) did not. Unskips the test.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Mirrors the postgres bed for minio: a provider (runtime carries mc) + a consumer
requiring s3-bucket, proving the consumer reaches its bucket with the access key and
secret the mesh delivered. It surfaced a real limit: the mesh derives `as` =
mesh_<node>_<module> (22 chars), and an S3 access key is capped at 20, so minio refuses
the service account. The test is correct and skipped pending 04-ISSUES/010, not worked
around.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Assigns a postgres provider (its runtime carries psql) and a module that requires
postgres-database; the mesh mints one password, postgres's provisioner creates a role
and database under the mesh's login with it, and the consumer connects to its database
with the delivered credential (a password-checked connection) — select 1. Nothing placed
by the test. The postgres half of the per-backend provider proof (ADR 0052/0053).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Assigns a redis provider and a module that requires redis-cache; the mesh mints one
password, seals a copy to each end, writes redis its contributions and the consumer
its bound file, and the host unseals each side. redis's provisioner creates the ACL
user under the mesh's login with the mesh's password, and the consumer's delivered
credential authenticates (PONG). Nothing is placed by the test — the provider/consumer
contract (ADR 0053) working as one thing, no shared key anywhere.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
build-module-runtime.sh adds psql to the postgres image and mc to the minio image
(their clients shell out to those). provider-on-backend-network asserts redis's
runtime, on the backend's private network, binds the broker via NAT and provisions
a consumer with the mesh's credential — the shape the committed provider manifests use.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Assigns redis as a provider, puts the contributions and unsealed password the mesh
would deliver in its receives path, and authenticates as the consumer with the mesh's
password — PONG proves the login was created with exactly that password (a
self-generated one answers WRONGPASS), with MESH_SEAL_KEY set nowhere.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Assigns grafana configured by settings, changes the token, pushes again, and
asserts the container was replaced (new id) and the rendered config carries the new
value. Builds the host from source, since the behaviour under test is the host's.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
assigned-sonarr proves the Servarr detection path: the runtime discovers its API
key from the app's config.xml and serves its tools. assigned-grafana proves the
settings path: the operator states URL and token as settings, the mesh merges them
into the module's config file, and the runtime serves from that with nothing in the
manifest. Both green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
build-module-runtime.sh generalises the audit-logger runtime image to any module
(mesh-tools + sdk + the module's dist, entrypoints for tools/events/provisioner).
Two scenarios and two tests: assigned-plex proves a tools+events module serves its
tools over a mesh-issued scoped account; assigned-redis proves a provider's runtime
serves tools AND runs its provisioner in the same broker-bound process, provisioning
a grant and emitting its lifecycle event. Both green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
test/integration/assigned-audit.test.ts raises a node into a mesh, assigns it
the audit-logger through the control plane, and asserts the mesh delivered a
scoped amqps account (not the broker's own), the host ran the container, and an
emitted event reached the trail — the delivered credential authenticating is
the proof. scenarios/audit-node.yml is the lean single-node bed that stocks the
runtime image. Passes 1/1 against the real lab.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Adds test/integration/events.test.ts: raises first-node (which raises the
broker as tier-1 substrate), runs the runtime+audit-logger against that
broker, emits a module and a node event, and asserts they reach the trail
with their metadata read back from ADR 0047 headers (a pure body), plus that
the durable per-consumer queue and mesh.events.dead exchange exist on the
raised broker — asked of the broker, not assumed.
scripts/build-runtime-image.sh builds the self-contained runtime image
(mesh-tools + vendored sdk + audit-logger) it runs, saved to a tar for
MESH_LAB_RUNTIME. The events path itself is verified; the incus raise is the
part a lab run exercises.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
With caught-up finally an equality, the last red test turned out to be
telling the truth about something real: declarations queue, the machine
applies them one at a time at half a minute each, and by the twenty-
fifth test it is minutes behind the latest push. 240 seconds was not a
generous bound on one apply — it was an accidental bound on the whole
backlog.
Doubled rather than tuned, and the real remedy filed instead: a machine
asked to be five successive things should become the last one, which
is a decision about the link rather than about this timeout
(novox/hq 04-ISSUES/031).
The timestamp comparison lost the race between one test's closing push
and the next test's opening one: the old apply's report landed newer
than the new send and settled() passed for a declaration the machine
had not read. The report names its declaration now, the mesh says
whether it is the current one, and this reads the answer instead of
inferring it.
The provisioner makes the role and then the database; a poll that
waited for the first and checked the second once was racing the gap
between two statements, and lost it once, eighteen seconds into a run.
The cache test's provisioner could not resolve the store's name, which
usually means the store's container never registered it — and the
diagnostics showed only the provisioner's side of that conversation. A
grant that never arrives now prints the container states, the store's
log and the provisioner's, so the next failure names the half that
actually fell over.
Eleven modules planned as one set on one machine, which is what a real
node looks like. Two are left out by name rather than silently: the
mesh under test already runs a module called registry and an adopted
workload called umami, and adding the catalogue's manifests would
replace the records of things that are live and assigned — the adopted
umami would suddenly require a database it never asked for.
settled() returned the moment a declaration was current, because the
sent digest is recorded at send — so both new tests asserted on a
machine still applying, and found containers not created yet and
bindings not written. The eternal-waiting fault had been standing in
front of this gap the whole time; fixing it is what let the tests get
far enough to fall in.
Caught up now means the machine's own last report is newer than what
was sent to it — two timestamps the mesh recorded itself, read from the
`reported` section status now carries. The residual latency between
"reported" and "every container answers" stays with the tests' own
polls, where it always was.
A consumer contributes a key prefix and gets an ACL user; the test is
that the grant means exactly what the manifest said, in both
directions: its own keys usable, anyone else's refused by the store
itself, and the flush a tenant must never have refused with them.
Waited for through the store rather than through logs: the user list,
asked with the password the host wrote into the server's own conf file
on the machine — nothing invented, both ends reading what the mesh
delivered.
And the forge is asked on the port the mesh assigned, not the one the
module declared. The old curl aimed at 3000, which was right until
ADR 0038 moved the machine side — a latent break that would have fired
on the first run to get past the settling that used to fail first.
The scenario stocks redis and its provisioner, and the rebuild builds
the provisioner image with the others.
A segment named "uplink" is refused. The lab claims that name for the
NAT bridge behind `egress: true`, and a scenario wearing it first would
have its egress machines silently attached to an isolated bridge — a
declared key doing nothing, which is the fault this repo exists to
refuse, in the repo that refuses it.
settled() parses inside the try. A truncated status from a struggling
machine was the one shape of bad answer that still threw out of the
wait, and the likeliest moment for one is exactly the machine the poll
is watching. Malformed now counts as "could not ask", like the exec
that times out.
And a sentence on the uplink's UseDNS saying its inertness is
load-bearing: it matters only where systemd-resolved runs, and on a
machine whose modules own resolv.conf the uplink must not outvote the
resolver a scenario is testing.
Three changes, found by one failing test.
The forge failed three runs in a row as "status hangs", and it was
diagnosed twice as contention — real defects, fixed, and not the cause.
The heartbeats told the truth in the end: every exec on anchor crawled
from 15s to 105s, because eleven containers plus a database pull were
running in a 1GiB machine. Starvation presents as whatever you were
doing when the page-outs start, which is why it wore two other bugs'
clothes first.
So machine size is now the scenario's to declare — memory and cpus per
machine, default unchanged. The anchor that carries the whole substrate
is bigger than the laptop that joins it, and the comment on the
scenario says why in terms of what lands there.
`egress: true` gives a machine one extra interface on a lab-supplied
NAT network, addressed by DHCP because the one address a scenario has
no business choosing is on the host's side of the fence. Declared
per machine and off by default: a closed scenario stays the rule
(novox/hq ADR 0016), and the exception exists because a first node
fetches its images before any mesh can serve them — which is now the
tested path (04-ISSUES/029), and a lab that can never reach upstream
cannot prove the bootstrap it exists to prove. The uplink route is
metric-4096, so it never shadows a route the scenario declared. A
detached machine declaring egress is refused, not ignored.
And settled() treats a poll that threw as a poll that missed. An exec
timeout at minute four of a wait is "could not ask", not a verdict on
the machine.
A mesh that has just bootstrapped cannot build the module that gives it
an artifact store: building publishes to the store, and the builder will
not start without one (novox/hq 04-ISSUES/029).
This test built it and passed, because the scenario's registry was
already standing to receive the push — which is precisely why a real
first mesh would have hit this and the lab never did. A stand-in for
Docker Hub was quietly also standing in for the thing under test.
So the module now names its image by digest, the way the bundle names
the three a first node starts from, and is added as a manifest rather
than built. That is the only path open to a real first mesh, so it is
the path this walks.
Its skip on MESH_LAB_BUILDER goes with it. Nothing in the test needs a
builder any more, and a skip that names a thing the test does not use
sends the next person to look in the wrong place.
`settled` used `must`, so a failed exec ended the wait as though the
machine had reported a failure. It had reported nothing: the control
plane is a container on the node being polled, and while that node
applies a declaration an exec into it can lose its stdout fifo to
containerd. The run then blamed the mesh for a question that missed.
Could not ask and asked, and the answer was bad are different facts, and
only the second is the machine's. A failed poll now keeps the reason and
tries again; the timeout reports whichever came last, so a control plane
that is genuinely unreachable still fails the test — with the reason
rather than with a stack trace.
Every five seconds rather than every two. Each poll is an exec into a
container on a machine that is busy applying, and thirty times a minute
was competing with the apply rather than observing it.
The forge test read `status` the instant `push` returned and concluded
the machine was fine. It was describing the apply before this one.
`push` sends and returns — it prints "sent N resource(s)" and the
machine applies afterwards. So every assertion made immediately after
one is racing it, and this race lost quietly: no failure reported, and
a container that did not exist yet read as a container that would never
exist.
`settled` asks the mesh, in its own terms: a node is caught up when it
is neither waiting for what it was sent nor wrong about what it applied
— the two questions `status` already answers, read as JSON so a test is
not parsing a report written for a person. A machine reporting a failure
ends the wait immediately rather than at the timeout, because it will
not become right by being waited for.
The container assertion now also prints the plan. A container missing
because the mesh never asked for it and one missing because the machine
could not make it are one sentence and two entirely different faults,
and the plan is what separates them.
The forge test pushed and then waited for a database login. When the
containers were never created at all, it reported "no login was created"
— true, and silent about why. Two hundred and thirty seconds spent
proving something downstream of the actual failure.
A push being accepted and an apply having worked are different facts,
and this test depends on the second. It now reads what the machine says
about itself, and whether a container exists, before it starts waiting —
and carries the host's own log into the failure either way.
The suite otherwise passed 24 of 25 on this run, which is the first time
the forge reached a clean attempt with nothing upstream blocking it.
A scenario is a closed address space: two raised from the same
declaration hold the same addresses and never meet, which is what lets
two run at once and why the lab talks to machines through the
hypervisor rather than over IP. Reaching in from outside breaks that, so
it is opt-in, one scenario at a time, and reversible.
`connect` takes an address on the scenario's public link and writes a
resolver rule answering everything under each machine's name.
`disconnect` gives both back. `connected` says what is true right now,
for somebody who cannot remember.
It refuses rather than guessing when more than one scenario is standing
— the failure being avoided is not an error but one scenario's traffic
arriving in another. It also refuses when a machine's name is already
answered here for something real, because connecting would point that
name at the lab, and the damage would land on the real thing.
Names answer with the segment address rather than the overlay one.
Inside the mesh a name gives a machine's private address; from here that
would need this workstation on the overlay, which is a much larger door.
The segment address reaches the same machine and the same ports, which
is what opening a board in a browser actually needs.
Proven against a live two-node scenario: registry.internal:5000/v2/
answered 200 from this workstation, and so did a wildcard name under the
same machine. Disconnect put the address back, stopped answering, and
left the real mesh's own names alone.
One thing measured rather than assumed: it restarts dnsmasq instead of
reloading it. A reload is SIGHUP, which re-reads the hosts file and
clears the cache but not the configuration — the rule was written, the
reload reported success, and nothing resolved. The daemon's start time
was nine days old afterwards.