The vault bed grows into a create-once file and pushes again; the genesis bed
probes the machine from the workstation for the whole install and asserts the
store's port never answers while the bus's does.
Applying the packet filter restarts the container runtime a few seconds
after the installer's last push returns; a command racing that window dies
with 'No such exec instance'. Wait for the node to report applied and
current, and retry that error like the recreate it is.
V5: the template's password is refused by the store, the operator key and the
export sit beside the bundle at 0600, the vault keeps the export, and a person
with the key recovers the superuser off the mesh and opens the store with it.
V2 dials the broker with the administrator password genesis made.
Design 13's three logins, for a secret that had no owner before (novox/hq
ADR 0085): the delivered password authenticates against the real redis, the
one `rotate secret` delivers authenticates, and the one rotated away is
refused. Plus the owner's half: the vault's ledger names the holder and the
fingerprint, notices the rotation, and answers over the mesh by fingerprint,
never by value. Runs the catalogue's own manifests.
The CORE convergence wait hung on a container named 'postgres' that
never exists — the postgres module's server is the adopted-store
container 'mesh-store' (like lavinmq's mesh-broker). Everything else
converged; this was the last phantom-name blocker.
The bed had never resolved, then never converged. Fixed, in order:
- loadManifest maps a runtime container's 060 `artifact` to the stocked
mesh-runtime-<module> image (keyed on the module name), so the push is
no longer refused by built() — and drops the build section.
- Stale identities renamed: registry->distribution, firewall->nftables.
- step-ca added to the set: the web modules hard-require `route`,
route-proxy provides it but requires `acme-ca`, and nothing provided
that — so the whole web stack never resolved. step-ca is the missing CA.
- THE STORE SUPERUSER is delivered via `secret accept` before the push.
postgres raises mesh-store with POSTGRES_PASSWORD=bootstrap, but
`module add` minted a random superuser own-secret that did not match,
so the provisioner could not log in and created NO consumer roles —
every DB consumer (gitea/keycloak/nextcloud/umami/mailu) failed. This
was the real cause behind what looked like per-module gaps; keycloak
and umami converge once it is delivered (ADR 0078, hq phase3).
- mesh() retries through the controller recreating itself during the
057 cascade (No such exec instance), so a real success is not read as
a failed push.
- invoicing dropped (private-registry images the lab cannot pull).
Remaining KNOWN_GAPS are genuine catalog/upstream/resource gaps: minio
(stale Docker Hub digest), mssql (Error 945, memory), mailu (config
env), photos (alpine placeholder), nftables (service).
- The RestartCount==0 assertion cannot discriminate the 058 fix: the
consumer is built after its broker is already up, so patient and
exit-on-unreachable code both connect first-try; and restart-on
recreates reset the count. Downgraded to an honest liveness check and
the comment now points at the mesh-tools unit test as the real proof.
- The trust-failure dump ran base64 -d over declared.json, which is JSON
(not base64), so it always reported 'no trust' — removed; the adjacent
python check that decodes the inner declaration field is kept.
- The 063 comment claimed the vhost is re-listed after the restart; the
code only runs a TCP probe. Comment corrected to what the code proves,
and the docker-proxy-vs-DNAT coupling is noted.
A broker that is reachable only until its conntrack entry drops passes
every test written before it restarts. The bed now restarts mesh-broker
after adoption and asserts the joined node can still reach 5671 — the
forward rule, not a surviving entry, carrying the connection.
The provider's workaround re-push is gone: pushing the consumer's node
must cascade to the provider (issue 057), and the vhost assertion is
what says so. The joined consumer must also show zero container-runtime
restarts (issue 058): a runtime whose broker is not up yet waits for it
in-process, so overlay-after-container ordering produces no churn.
Run 11 failed with node2's daemon.json never written and nothing to say
whether the declaration lacked the trust or never applied. The dump now
answers that, and the push output is printed so a compose that refused
is visible in the run log.
The networking module delivers the registry trust, so the bed pushes both machines after
assigning it and waits for each runtime to actually hold the trust (file present AND the
daemon reloaded) before the first build pushes to anchor.internal:5000.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ready() asked a gone instance for its snapshots before judge could say 'no longer
standing'; a stale warm.json from any earlier scenario made every warm run fail in
milliseconds. The question is now asked only of an instance that still exists.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
MESH_LAB_WARM=1 snapshots the post-genesis, both-nodes-enrolled state and restores it in
seconds on later runs — refused, not silently rebuilt, when the commits have moved
(src/warm.ts, the mechanism mesh.test.ts already uses and this session had ignored).
Fresh stays the default.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
tsconfig.test.json existed precisely so a test that does not compile cannot silently be a
test that never ran — and four beds did not compile: two returned strings from test bodies,
two predate GenesisOptions gaining sdkSource, one passed a nullable host binary. All clean;
the gate is now part of launching any bed.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The review found a race: a container that settled in the window's last seconds could be
re-inspected past the deadline and failed as 'never stopped restarting'. A boolean now says
what happened.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
letta is not mesh-buildable (hq issue 060) — the store's cross-node proof stays with the
stocked bed until a DB consumer gains a build section; amqp-ping carries the no-fake proof
alone, and the node2 delivery is named as the honest red gate for issues 042/048 (no
registry account, no registry trust). Also: overlay sites match genesis (hosting), the
manifest is read locally instead of a swallowed docker-exec, before() gets the one-node
budget, a node2 failure appends the host log (a failed pull never reaches container logs),
and the foundation's survival plus the consumer's steadiness are asserted.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The review found a race: a container that settled in the window's last seconds could be
re-inspected past the deadline and failed as 'never stopped restarting'. A boolean now says
what happened.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Run 3 showed the Phase-3 installer already adopts postgres (superuser included), nftables
and the catalogue at genesis — re-registering them was redundant. Only lavinmq and the
joined node's consumers are the bed's to add. Issuance is now asserted per module and each
module is pushed as it lands, mirroring the one-node bringUp.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The scenario names no images and the bed rewrites nothing: the installer raises anchor
(building the control plane), the mesh's own builder builds base, postgres, lavinmq and
the joined node's consumers from the forge and pins every digest itself, the committed
manifests are registered verbatim, and node2 proves both foundation halves cross-node.
Being iterated toward green (run 3 in flight); banked so nothing is lost.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0048 (2026-09-05) settled that a provider is handed the credential the mesh minted —
sealed to the provider node, unsealed by the host into a 0600 file — and removed the
symmetric seal from the SDK entirely; hq issue 032 records it resolved. The lab-only
MESH_SEAL_KEY injections were tombstones read by nothing: two-node-db went green on the
superuser delivery, not the seal key. Removed, and provider-uses-mesh-credential's
citations corrected from ADR 0053 (a scheduled step) to ADR 0048.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The complete recipe, found across five runs: adopt BOTH the store (postgres) and broker
(lavinmq) on the control-node — the broker's `listens` is what opens 5671 in the firewall
for cross-node bus access; deliver the store's genesis superuser via `secret accept` (else
the module mints a random one that cannot log in to the running store); inject the provider
seal key (open hq issue 022 workaround); push the provider node again after the remote
consumers (issue 057); and tolerate provisioner-runtime startup churn — a cross-node
provisioner exits until the overlay tunnel is up, then settles. baserow and letta on the
joined node get their databases from the one foundation store over the overlay.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The bed assigned a separate app-postgres, which ADR 0079 now refuses. Converted to
adopt the foundation store on anchor and have baserow/letta consume it cross-node over
the overlay, with the 057 push-ordering. The store DB path reaches its asserts, but the
run is blocked by issue 058: redis's host-networked provisioner cannot reach the broker
across nodes (bridge consumers can). Committed as WIP until 058 is fixed.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A provision secret is minted as a side-effect of composing the CONSUMER's plan, and the
provider's grant list is a pure read of secrets already issued from it. So a cross-node
consumer's grant exists only after its node is pushed, and the provider's provisioner mints
the vhost only when the provider node is composed again. Push anchor once more after node2,
and the bed passes: amqp-ping on node2 reaches mesh-broker on anchor over the overlay, its
binding names anchor.internal, and its vhost is minted. Proves both halves of issue 055.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
anchor raises the foundation and adopts lavinmq; node2 joins and runs amqp-ping,
which requires amqp and provides nothing. Asserts the grant names anchor.internal,
a vhost is minted on the far broker, and the consumer stays up. Currently RED: it
caught two real gaps — the broker's amqps port not in the firewall (fixed in
mesh-catalog) and the module broker URL using the public address not the overlay
(issue 055, needs a controller fix). Goes green when 055 is fixed.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
S1 upgrades the store: a spec change recreates mesh-store (the server the control
plane reads from), and asserts the data on the named volume survives and the
pool reconnects — the stated window. It also asserts postgres/lavinmq are now
source-tracked modules the mesh can report behind (3.4), the question that could
not form before adoption. S2 does the same for the broker, the harder case: the
push that upgrades it travels over it, so it proves the mesh reconnects to the
bus it just replaced.
Issue 051 (WBS 3.3, 3.4).
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The V3 networking check asserted a `lavinmq` docker network exists, from the
two-server world. The module now adopts the foundation's broker rather than
raising its own on a private network (issue 051, WBS 3.2), so only the
consumer's own network remains.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Installs the shipped nox-mesh-host launcher and unit in the machine and lets the
installer's --host-service start and enable it, instead of --host-in-background
which cannot survive a reboot. E2 now proves the mesh comes back on its own.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx