Commit Graph
100 Commits
Author SHA1 Message Date
jschoubben 26177d81be Six beds retired: their coverage lives in the catalogue beds and the whole-mesh beds now
The four sidecar beds (grafana, plex, sonarr, redis) proved a sidecar comes up and
serves tools with the module's server cut away; the catalogue beds and the whole-mesh
beds prove the modules whole. The minio and postgres grant beds proved a grant
mechanism with a second store beside the foundation's; the grant bed and the vault bed
prove it against the catalogue. Ten conversions become six deletions (novox/hq
04-ISSUES/074).
2026-09-21 21:53:11 +02:00
jschoubben 543ccb7380 Merge pull request 'Multiple fixes: the vault bed keeps two secrets, the confluence bed asks through the control plane, the runtime build installs its SDK and speaks' (#44) from multiple-fixes into main 2026-09-21 21:05:12 +02:00
jschoubben 85dc827660 The confluence bed asks a tool through the control plane, and the tool answers
No account in the mesh but the control plane's may create a reply queue and publish
to a module's request key (novox/hq 04-ISSUES/049, ADR 0095).
2026-09-21 20:34:03 +02:00
jschoubben 2f11a29c4d The runtime build installs the module's SDK itself and shows its compiler's output
A module never built on this workstation had no node_modules, tsc failed on the
first import behind /dev/null, and the suite reported a build that said nothing.
2026-09-21 20:31:47 +02:00
jschoubben 64c081d025 The vault bed installs a consumer that keeps two secrets, and both are delivered and rotated
Two files with two values, two holders in the vault's ledger — the identity with the
local name after it — and one rotate moves both (novox/hq ADR 0094).
2026-09-21 20:29:49 +02:00
jschoubben 467416728e Merge pull request 'The run rebuilds the module runtimes its beds stock (075); two mechanism beds read the catalogue's redis (074)' (#43) from feat/mesh-tests-and-runtimes into main 2026-09-21 19:36:44 +02:00
jschoubben e4c924a85e The grant bed's consumer takes a slug: its derived identity was one character over what a backend keeps 2026-09-21 19:32:36 +02:00
jschoubben 1b2443690d Two mechanism beds install the catalogue's redis with the vault beside it
A module's name is its tool namespace and its broker scope, so a fixture running
the module's runtime cannot carry another name (novox/hq ADR 0093). The grant and
backend-network beds now read the catalogue's redis and install mesh-vault, which
provides the secret it requires; the redis-node scenario stocks the vault's runtime.
The remaining declared copies say why they stand (novox/hq 04-ISSUES/074).
2026-09-21 19:30:34 +02:00
jschoubben d7959001dd The run rebuilds the module runtimes its beds stock
A per-module bed stocks mesh-runtime-<module>:development from the workstation's
image store, built by hand by a script the suite never called; six were two weeks
older than the manifests they served. For the beds named, every runtime their
scenarios stock is compared against the module's source and the tool runtime and
SDK it is built on, and rebuilt where older, missing or uncommitted; a failed
build stops the suite (novox/hq 04-ISSUES/075).
2026-09-21 19:26:59 +02:00
jschoubben be58bd96f6 Merge pull request 'Genesis reads the registry's and the builder's manifests from the catalogue, not the control plane's (hq issue 072)' (#42) from feat/one-controller-manifest into main 2026-09-21 19:23:56 +02:00
jschoubben 6083b9310a Genesis reads the registry's and the builder's manifests from the catalogue, not the control plane's
The control plane's manifest comes out of the build the installer runs (novox/hq ADR
0069, 04-ISSUES/072); the catalogue no longer holds a copy, and a bed that demanded
one would stop a genesis that is about to succeed.
2026-09-21 19:23:44 +02:00
jschoubben c8c3028f38 Merge pull request 'Beds read the catalogue: one loader, eight beds converted, the rest declared (hq issue 073)' (#41) from feat/beds-read-the-catalogue into main 2026-09-21 19:23:15 +02:00
jschoubben ec23e9f4cd The catalogue is named or absent, the scanner reads both key orders, the loader has unit tests
Review findings: a sibling-path fallback read a catalogue the receipt never claimed;
a manifest literal naming its version first slipped the fence; the shared loader
thirteen beds install through had no test short of a lab run.
2026-09-21 19:21:54 +02:00
jschoubben 8c37328ba3 The catalogue is recognised by the registry's manifest, not the control plane's
The control plane's manifest is leaving the catalogue (ADR 0069, issue 072); a marker
that named it would make every catalogue-reading bed skip the day it goes.
2026-09-21 15:18:46 +02:00
jschoubben e596758db9 The five beds that read the catalogue read it through the harness
adopted-store-cross-node, two-node-db and the three whole-mesh beds each carried a
private loader; they drifted. The ace loader never resolved a runtime artifact, so a
module the mesh builds travelled unresolved; whole-mesh-full still asked for
'registry' and 'firewall', which the catalogue names distribution and nftables, and
swallowed the miss as NOT ASSIGNED. One loader now (novox/hq 04-ISSUES/073).
2026-09-21 14:33:02 +02:00
jschoubben 2456b2f533 Beds read the catalogue: a shared loader, eight beds converted, the rest declared
catalogueModule() in the harness reads a module's manifest from the catalogue and
rewrites only what the lab must: the build section goes, each artifact becomes the
image the machine holds, images are pinned, and a bed may declare a host-port remap
or a lab-local address. confluence, gitlab, openai-consumer, audit-logger, ollama,
local-model-consumer, model-usage, mosquitto, anthropic-manager and
anthropic-consumer now install the catalogue's manifest. A unit test refuses any
inline copy naming a catalogue module unless the bed is declared with its reason;
the declared list is the debt (novox/hq 04-ISSUES/073).
2026-09-21 14:27:49 +02:00
jschoubben f2d29491b2 The receipt claims the catalogue the beds read
A bed installs a catalogue module by reading its manifest from MESH_LAB_CATALOG at
run time, so a receipt naming no catalogue commit cannot say whether a catalogue
change was ever proven. Claimed under either spelling of the variable; not built,
because a manifest is read, not compiled (novox/hq 04-ISSUES/073).
2026-09-21 14:14:54 +02:00
jschoubben 43bf78bd63 Merge pull request 'Beds for the seed file and the foundation's filter; beds deliver secrets as files; a bundle-raised bed derives the anchor's filter' (#40) from feat/migration-blockers into main 2026-09-21 13:48:57 +02:00
jschoubben b0e7cf96d1 The apps bed's mongodb names its secrets' owner, as the catalogue's does, and says what the server said when the consumer cannot reach it 2026-09-21 13:43:42 +02:00
jschoubben e085e31f95 A bed that raises the foundation from the bundle derives the anchor's filter before it relies on the hub
The base ruleset (ADR 0088) admits ssh, the bus and the registry and nothing else until the mesh
derives one, and the mesh derives one only where the filter module is assigned — which genesis
does and these beds did not. Without it the hub's WireGuard port stayed closed, no joined node's
tunnel formed, and every module dialling the anchor by its overlay name timed out fetching the
broker's certificate; the model-usage bed showed it as a login that failed for a role never made.
2026-09-21 13:30:26 +02:00
jschoubben af14f1c909 The model-usage bed can be left standing, and its account names the filters 2026-09-21 13:18:23 +02:00
jschoubben ba49f571dc The model-usage bed says what the machine knows when the login fails 2026-09-21 13:12:19 +02:00
jschoubben 67a1f6a012 The harness pins minio from quay.io, as the catalogue does (docker.io denies anonymous pulls) 2026-09-21 12:58:35 +02:00
jschoubben b7316d40fe The three beds' inline postgres reads its superuser from a file, as the catalogue's does 2026-09-21 12:55:35 +02:00
jschoubben 97d71503ee Three beds deliver the secrets they copy from the catalogue as files (issue 073) 2026-09-21 12:32:00 +02:00
jschoubben 599d41eb42 Beds for the seed file and for the foundation's filter
The vault bed grows into a create-once file and pushes again; the genesis bed
probes the machine from the workstation for the whole install and asserts the
store's port never answers while the bus's does.
2026-09-21 12:11:52 +02:00
jschoubben 7e2e97f056 Merge pull request 'Beds for the vault and for genesis's root secrets' (#39) from feat/secrets-vault into main 2026-09-21 10:03:22 +02:00
jschoubben 7fad006e56 A push that said 'told' is a push that sent, whatever became of the exec afterwards 2026-09-21 01:47:26 +02:00
jschoubben fb72db73bc The vault bed recovers redis's vault-provided secret from the export 2026-09-21 00:36:16 +02:00
jschoubben 960fa3607f The genesis bed waits for the node to settle before asking it anything
Applying the packet filter restarts the container runtime a few seconds
after the installer's last push returns; a command racing that window dies
with 'No such exec instance'. Wait for the node to report applied and
current, and retry that error like the recreate it is.
2026-09-21 00:33:23 +02:00
jschoubben 9da2d01ca0 The genesis bed checks the root secrets: made, sealed to the operator key, recoverable
V5: the template's password is refused by the store, the operator key and the
export sit beside the bundle at 0600, the vault keeps the export, and a person
with the key recovers the superuser off the mesh and opens the store with it.
V2 dials the broker with the administrator password genesis made.
2026-09-21 00:12:55 +02:00
jschoubben 639175ec4a A bed for the vault: redis's password as a secret it provides, rotated
Design 13's three logins, for a secret that had no owner before (novox/hq
ADR 0085): the delivered password authenticates against the real redis, the
one `rotate secret` delivers authenticates, and the one rotated away is
refused. Plus the owner's half: the vault's ledger names the holder and the
fingerprint, notices the rotation, and answers over the mesh by fingerprint,
never by value. Runs the catalogue's own manifests.
2026-09-21 00:01:50 +02:00
jschoubben a43e6d2927 Merge pull request 'Land the whole-mesh-novox bed (green) and the minio quay.io fix' (#38) from fix/minio-and-whole-mesh-novox-bed into main 2026-09-20 22:08:50 +02:00
jschoubben 6bdf7104c8 whole-mesh-novox: the postgres server container is mesh-store, not postgres
The CORE convergence wait hung on a container named 'postgres' that
never exists — the postgres module's server is the adopted-store
container 'mesh-store' (like lavinmq's mesh-broker). Everything else
converged; this was the last phantom-name blocker.
2026-09-20 22:06:26 +02:00
jschoubben 0a05fbb434 whole-mesh-novox goes green: the store superuser, the artifact shape, and the missing CA
The bed had never resolved, then never converged. Fixed, in order:
- loadManifest maps a runtime container's 060 `artifact` to the stocked
  mesh-runtime-<module> image (keyed on the module name), so the push is
  no longer refused by built() — and drops the build section.
- Stale identities renamed: registry->distribution, firewall->nftables.
- step-ca added to the set: the web modules hard-require `route`,
  route-proxy provides it but requires `acme-ca`, and nothing provided
  that — so the whole web stack never resolved. step-ca is the missing CA.
- THE STORE SUPERUSER is delivered via `secret accept` before the push.
  postgres raises mesh-store with POSTGRES_PASSWORD=bootstrap, but
  `module add` minted a random superuser own-secret that did not match,
  so the provisioner could not log in and created NO consumer roles —
  every DB consumer (gitea/keycloak/nextcloud/umami/mailu) failed. This
  was the real cause behind what looked like per-module gaps; keycloak
  and umami converge once it is delivered (ADR 0078, hq phase3).
- mesh() retries through the controller recreating itself during the
  057 cascade (No such exec instance), so a real success is not read as
  a failed push.
- invoicing dropped (private-registry images the lab cannot pull).

Remaining KNOWN_GAPS are genuine catalog/upstream/resource gaps: minio
(stale Docker Hub digest), mssql (Error 945, memory), mailu (config
env), photos (alpine placeholder), nftables (service).
2026-09-20 22:06:26 +02:00
jschoubben a57eba9de0 build-module-runtime: pull minio mc from quay.io (docker.io denies anonymous pulls) 2026-09-20 22:06:26 +02:00
jschoubben 3c3fa04949 Merge pull request 'The bed's 058 claim, base64 dump, and 063 comment made honest (review)' (#36) from bed/honest-058-and-diagnostics into main 2026-09-20 13:42:00 +02:00
jschoubben 064a7a0934 The bed's 058 claim, base64 dump, and 063 comment are made honest (review)
- The RestartCount==0 assertion cannot discriminate the 058 fix: the
  consumer is built after its broker is already up, so patient and
  exit-on-unreachable code both connect first-try; and restart-on
  recreates reset the count. Downgraded to an honest liveness check and
  the comment now points at the mesh-tools unit test as the real proof.
- The trust-failure dump ran base64 -d over declared.json, which is JSON
  (not base64), so it always reported 'no trust' — removed; the adjacent
  python check that decodes the inner declaration field is kept.
- The 063 comment claimed the vhost is re-listed after the restart; the
  code only runs a TCP probe. Comment corrected to what the code proves,
  and the docker-proxy-vs-DNAT coupling is noted.
2026-09-20 13:32:14 +02:00
jschoubben 27665501dc Merge pull request 'The bed enforces one-push provisioning, a patient runtime, and a broker restart (057/058/063)' (#35) from bed/one-push-and-a-patient-runtime into main 2026-09-20 12:53:18 +02:00
jschoubben 68133c2c56 The bed restarts the broker and reconnects across the overlay (issue 063)
A broker that is reachable only until its conntrack entry drops passes
every test written before it restarts. The bed now restarts mesh-broker
after adoption and asserts the joined node can still reach 5671 — the
forward rule, not a surviving entry, carrying the connection.
2026-09-18 02:21:48 +02:00
jschoubben 38a7180b8b The bed enforces one-push provisioning and a patient runtime (057/058)
The provider's workaround re-push is gone: pushing the consumer's node
must cascade to the provider (issue 057), and the vhost assertion is
what says so. The joined consumer must also show zero container-runtime
restarts (issue 058): a runtime whose broker is not up yet waits for it
in-process, so overlay-after-container ordering produces no churn.
2026-09-18 02:07:59 +02:00
jschoubben 85c8ad8131 Merge pull request 'The no-fake multi-node gate: built-store-cross-node' (#34) from bed/built-store-cross-node into main 2026-09-18 01:00:49 +02:00
jschoubben cb353f9881 The trust wait dumps mesh-host's log and the declaration on failure
Run 11 failed with node2's daemon.json never written and nothing to say
whether the declaration lacked the trust or never applied. The dump now
answers that, and the push output is printed so a compose that refused
is visible in the run log.
2026-09-18 00:15:16 +02:00
jschoubben ca263d2a8b The trust lands before anything builds
The networking module delivers the registry trust, so the bed pushes both machines after
assigning it and waits for each runtime to actually hold the trust (file present AND the
daemon reloaded) before the first build pushes to anchor.internal:5000.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:16:26 +02:00
jschoubben a17a2e854b The bed rides the warm fix 2026-09-17 23:08:19 +02:00
jschoubben 1e510787c1 Merge pull request 'warm refuses a departed instance instead of crashing into it' (#33) from fix/warm-refuses-the-departed into main 2026-09-17 23:08:04 +02:00
jschoubben 8374f4dc19 warm refuses a departed instance instead of crashing into it
ready() asked a gone instance for its snapshots before judge could say 'no longer
standing'; a stale warm.json from any earlier scenario made every warm run fail in
milliseconds. The question is now asked only of an instance that still exists.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 23:07:50 +02:00
jschoubben 5ded8e2a8e The bed keeps its genesis warm
MESH_LAB_WARM=1 snapshots the post-genesis, both-nodes-enrolled state and restores it in
seconds on later runs — refused, not silently rebuilt, when the commits have moved
(src/warm.ts, the mechanism mesh.test.ts already uses and this session had ignored).
Fresh stays the default.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:55:00 +02:00
jschoubben b6d2a0602a Merge pull request 'The settle check records settling instead of inferring it from the clock' (#32) from fix/settle-check-edge into main 2026-09-17 22:54:24 +02:00
jschoubben 832c287ccc The bed passes the gate under exactOptionalPropertyTypes
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:54:09 +02:00
jschoubben 31969e20e3 The bed rides the typecheck-gated main 2026-09-17 22:54:06 +02:00
jschoubben 90e84279f7 Merge pull request 'Every test passes the typecheck gate' (#31) from fix/tests-pass-the-typecheck-gate into main 2026-09-17 22:53:53 +02:00
jschoubben 494f73369a Every test passes the typecheck gate
tsconfig.test.json existed precisely so a test that does not compile cannot silently be a
test that never ran — and four beds did not compile: two returned strings from test bodies,
two predate GenesisOptions gaining sdkSource, one passed a nullable host binary. All clean;
the gate is now part of launching any bed.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:53:34 +02:00
jschoubben 27b5f8d092 The bed passes the typecheck gate
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:52:31 +02:00
jschoubben b77d03a1f8 readFileSync is imported, not assumed
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:48:34 +02:00
jschoubben 48f8fe6168 The settle check records settling instead of inferring it from the clock
The review found a race: a container that settled in the window's last seconds could be
re-inspected past the deadline and failed as 'never stopped restarting'. A boolean now says
what happened.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:40:16 +02:00
jschoubben 158fe5c327 Apply the bed review: one buildable consumer, honest gates, tighter plumbing
letta is not mesh-buildable (hq issue 060) — the store's cross-node proof stays with the
stocked bed until a DB consumer gains a build section; amqp-ping carries the no-fake proof
alone, and the node2 delivery is named as the honest red gate for issues 042/048 (no
registry account, no registry trust). Also: overlay sites match genesis (hosting), the
manifest is read locally instead of a swallowed docker-exec, before() gets the one-node
budget, a node2 failure appends the host log (a failed pull never reaches container logs),
and the foundation's survival plus the consumer's steadiness are asserted.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:39:49 +02:00
jschoubben 428f13bfae The settle check records settling instead of inferring it from the clock
The review found a race: a container that settled in the window's last seconds could be
re-inspected past the deadline and failed as 'never stopped restarting'. A boolean now says
what happened.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:39:48 +02:00
jschoubben d3a6dc9784 The bed adopts only what genesis leaves it: the broker
Run 3 showed the Phase-3 installer already adopts postgres (superuser included), nftables
and the catalogue at genesis — re-registering them was redundant. Only lavinmq and the
joined node's consumers are the bed's to add. Issuance is now asserted per module and each
module is pushed as it lands, mirroring the one-node bringUp.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:23:13 +02:00
jschoubben bfd70e9e57 The no-fake multi-node bed: two machines, everything built by the mesh itself
The scenario names no images and the bed rewrites nothing: the installer raises anchor
(building the control plane), the mesh's own builder builds base, postgres, lavinmq and
the joined node's consumers from the forge and pins every digest itself, the committed
manifests are registered verbatim, and node2 proves both foundation halves cross-node.
Being iterated toward green (run 3 in flight); banked so nothing is lost.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:15:34 +02:00
jschoubben e3d96d9fa9 Merge pull request 'Remove the inert MESH_SEAL_KEY; cite ADR 0048 correctly' (#30) from cleanup/seal-key-tombstones into main 2026-09-17 22:02:45 +02:00
jschoubben 2438883a5f Remove the inert MESH_SEAL_KEY, and cite the ADR that retired it correctly
ADR 0048 (2026-09-05) settled that a provider is handed the credential the mesh minted —
sealed to the provider node, unsealed by the host into a 0600 file — and removed the
symmetric seal from the SDK entirely; hq issue 032 records it resolved. The lab-only
MESH_SEAL_KEY injections were tombstones read by nothing: two-node-db went green on the
superuser delivery, not the seal key. Removed, and provider-uses-mesh-credential's
citations corrected from ADR 0053 (a scheduled step) to ADR 0048.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:02:31 +02:00
jschoubben 1e4713c311 Merge pull request 'two-node-db on the one-store model: store cross-node proven end-to-end' (#29) from multi-node/one-store-beds into main 2026-09-17 09:54:43 +02:00
jschoubben 0e382f1db7 two-node-db is GREEN on the one-store model — store cross-node proven end-to-end
The complete recipe, found across five runs: adopt BOTH the store (postgres) and broker
(lavinmq) on the control-node — the broker's `listens` is what opens 5671 in the firewall
for cross-node bus access; deliver the store's genesis superuser via `secret accept` (else
the module mints a random one that cannot log in to the running store); inject the provider
seal key (open hq issue 022 workaround); push the provider node again after the remote
consumers (issue 057); and tolerate provisioner-runtime startup churn — a cross-node
provisioner exits until the overlay tunnel is up, then settles. baserow and letta on the
joined node get their databases from the one foundation store over the overlay.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 09:53:35 +02:00
jschoubben d8540bec42 Convert two-node-db to the one-store model (WIP: blocked on issue 058)
The bed assigned a separate app-postgres, which ADR 0079 now refuses. Converted to
adopt the foundation store on anchor and have baserow/letta consume it cross-node over
the overlay, with the 057 push-ordering. The store DB path reaches its asserts, but the
run is blocked by issue 058: redis's host-networked provisioner cannot reach the broker
across nodes (bridge consumers can). Committed as WIP until 058 is fixed.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 09:00:33 +02:00
jschoubben a8afe5ce53 Merge pull request 'A two-node bed: a joined node opens the adopted broker over the overlay' (#28) from multi-node/cross-node-bed into main 2026-09-17 01:44:53 +02:00
jschoubben b6bf31d4e6 The cross-node bed is green: push the provider node after adding the remote consumer
A provision secret is minted as a side-effect of composing the CONSUMER's plan, and the
provider's grant list is a pure read of secrets already issued from it. So a cross-node
consumer's grant exists only after its node is pushed, and the provider's provisioner mints
the vhost only when the provider node is composed again. Push anchor once more after node2,
and the bed passes: amqp-ping on node2 reaches mesh-broker on anchor over the overlay, its
binding names anchor.internal, and its vhost is minted. Proves both halves of issue 055.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:34:22 +02:00
jschoubben 155800bc9a A two-node bed: a joined node opening the adopted broker over the overlay (055)
anchor raises the foundation and adopts lavinmq; node2 joins and runs amqp-ping,
which requires amqp and provides nothing. Asserts the grant names anchor.internal,
a vhost is minted on the far broker, and the consumer stays up. Currently RED: it
caught two real gaps — the broker's amqps port not in the firewall (fixed in
mesh-catalog) and the module broker URL using the public address not the overlay
(issue 055, needs a controller fix). Goes green when 055 is fixed.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:01:54 +02:00
jschoubben 03d036cbeb Merge pull request 'One-node bed: SDK-by-version, rename, and the store/broker upgrade proofs' (#27) from feat/a-bed-that-hands-over-nothing into main 2026-09-16 23:25:34 +02:00
jschoubben 440e2653b2 Phase 3.3/3.4: prove the store and broker upgrade in place, through the window
S1 upgrades the store: a spec change recreates mesh-store (the server the control
plane reads from), and asserts the data on the named volume survives and the
pool reconnects — the stated window. It also asserts postgres/lavinmq are now
source-tracked modules the mesh can report behind (3.4), the question that could
not form before adoption. S2 does the same for the broker, the harder case: the
push that upgrades it travels over it, so it proves the mesh reconnects to the
bus it just replaced.

Issue 051 (WBS 3.3, 3.4).

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 22:55:56 +02:00
jschoubben 70c1545161 Phase 3.2: lavinmq declares no network — it adopts mesh-broker
The V3 networking check asserted a `lavinmq` docker network exists, from the
two-server world. The module now adopts the foundation's broker rather than
raising its own on a private network (issue 051, WBS 3.2), so only the
consumer's own network remains.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 21:33:16 +02:00
jschoubben 5d6e8fbe7a Rename mesh-control -> mesh-controller, substrate -> foundation
One name per thing, per the HQ glossary: the module/container/image/binary/repo
becomes mesh-controller, the seat the-controller, and the store+broker pair the
foundation (embedded base bundles, default template and example lock renamed with
their go:embed directives). No behaviour change — a pure vocabulary rename.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 18:40:40 +02:00
jschoubben 49b80d8516 The one-node bed supervises the host as a service, so reboot is a real check
Installs the shipped nox-mesh-host launcher and unit in the machine and lets the
installer's --host-service start and enable it, instead of --host-in-background
which cannot survive a reboot. E2 now proves the mesh comes back on its own.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 16:20:44 +02:00
jschoubben f57e05e75e The one-node bed places at the site it operates, and tolerates the CP recreating
Genesis places the anchor at the same site the test re-places it at, so that step
is a no-op rather than a change that recreates the control plane. And mesh() —
which runs commands inside the control-plane container — retries a transient
"container not running", because the control plane is a live mesh-managed
container the mesh recreates when its declaration changes (e.g. its first
.internal add-host). Also passes --sdk-source/--tools-ref for the SDK build.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 14:10:26 +02:00
jschoubben eff91742e8 The bed publishes the SDK before the base build
genesis passes --sdk-source so the installer raises the registry and publishes
the SDK ahead of the base, which now resolves it by version.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 10:27:37 +02:00
jschoubben 61abeb7f6b The lab stocks the full catalogue, not the bootstrap three
Phase two of the installer reads each module's manifest from --catalog, which is
documented as a checkout of the catalogue repository. Genesis stocked it with
only the three modules the pivot needs, so step 14 failed reading postgres's
manifest — a file nobody had put there.

The installer code is right: --catalog is meant to be a full checkout. The lab
was the shortcut. It now copies the whole modules tree once (tar, push, extract)
rather than three files, and still checks the bootstrap three are present so a
missing one fails at preparation rather than at step 8. Production's equivalent is
an operator with a full checkout, or the installer cloning the repo.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 01:01:19 +02:00
jschoubben 09a22de78f The lab answers the installer's questions the unattended way
The installer asks where a human must choose, and this bed has no human — so
every choice arrives as a flag, and a required choice with no flag is the
installer refusing, which is the behaviour rather than a lab problem.

Both choices have one option today, so the flags are redundant on purpose: the
day a second filter exists this bed keeps working instead of refusing, and
choosing becomes a thing it visibly does.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 21:57:26 +02:00
jschoubben e451a9191a The lab names the modules as the catalogue does
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 20:51:00 +02:00
jschoubben fce554d150 A run-once step leaves nothing to ask, and V4 asked anyway
V4's first run reported a working mesh as broken: it asserted that lavinmq's
run-once bootstrap container existed, and the host removes an exited run-once
container on purpose — so a later apply is not confused by a stopped one, keeping
the record that it ran in its own store instead.

So the check asserted the opposite of correct behaviour. The step had run; it is
why the broker came up configured.

Counted as unverified now rather than assumed good. What would verify a step is
the host's own record of having run it, and this walks the machine rather than
the host — so the honest answer is that this check says nothing about steps, and
it now says so.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 01:35:42 +02:00
jschoubben 1425f5e7d4 Verify every resource kind, not the one I kept looking at
Every check in this file asked about containers. A container is one resource kind
out of ten — directory, file, user, network, access, archive, service, package,
container, action — and a module is far more often the others: the firewall is a
package and a service, the mesh's names are a file, a run-once step is an action
or a container that exits. Asking only about containers is how a module with no
container at all went unnoticed.

V4 takes the declaration the machine was actually sent and verifies each resource
in it, by kind, on the machine. Nothing is hand-picked — whatever the installed
modules declared is what gets checked.

And it reports which kinds were never exercised, rather than counting their
absence as success. A vocabulary this test never sees is a vocabulary this test
says nothing about, and saying so is the difference between a passing run and a
meaningful one.

The service check accepts a one-shot that has done its work and reports inactive,
which is the reading that made the firewall module look broken on every machine
for months.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-15 00:46:05 +02:00
jschoubben 3ebf0bb38c Containers coming back is not the mesh coming back
E2 asserted that every container was running after a reboot and stopped there.
The runtime restarts containers by itself; what makes a machine part of a mesh is
an agent listening for what it should be. A machine whose containers returned and
whose agent did not looks healthy and cannot be told anything.

The installer is explicit that a host started the way the lab starts it does not
survive a reboot, so this may now fail — and if it does, it is the packaging gap
the design already records under what is not yet true, not a fault in the mesh.
Better a named failure than a pass that means less than it appears to.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 23:44:24 +02:00
jschoubben 9b8b21ac17 E1 moves the module's source for real
Reading the branch head and telling the mesh the source had moved there named the
commit it had just built, so the mesh correctly answered that everything was
current. Naming a different commit would not work either: staleness compares
artifacts, not commits, deliberately, so that editing a comment in a shared base
does not rebuild everything standing on it to arrive back where it started.

So the step makes a real change and pushes it, and asserts the module comes back
on a DIFFERENT artifact than it had. A test that writes to a branch is worth
knowing about; the alternative is proving the loop by telling the mesh something
untrue.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 23:36:05 +02:00
jschoubben 475fd9a088 Register the firewall before assigning it
'no module of that name: firewall'. The networking family is computed by the
control plane, so the mesh knows those exist without anyone saying so; the
firewall is an ordinary catalogue module and has to be added like any other.

That is a second way for a module to be absent, and a less obvious one than
being present and placed nowhere — the mesh does not hold it at all, so nothing
can even report it unassigned.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:52:26 +02:00
jschoubben f132a4bcbd The packet filter is a module too, and it was assigned to nothing
V3 asked whether the machine's networking is what the modules asked for and found
no mesh firewall table at all. The firewall is a module — it claims the
packet-filter seat, installs the filter and loads the rules — and like networking
before it, it had never been assigned to anything.

So every rule the mesh generates from module listen declarations had never been
applied to any machine in this test. Not open by accident: a mesh where that
whole generation has never run.

Assigned separately from networking because they answer different questions. One
is how machines reach each other; the other is what may reach this one.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:45:52 +02:00
jschoubben 6b482ee7dc Assigning networking is not being on the network
N1 assigned the module and stopped, and the mesh wrote a names file with no names
in it. That is correct behaviour, not a bug: a node with no address on the
network has no name, because a name resolving to nothing is worse than no name —
a connection to an address that does not answer hangs, where a name that does not
resolve fails at once and says so.

The missing act is placement. Assigning installs the module that answers how
machines reach each other; placing says where this machine is on the resulting
network. The four-machine test did both and this one did neither, which is how
the distinction stayed invisible.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:26:04 +02:00
jschoubben d0d56782a0 Split describing the mesh from the catalogue holding it
Two claims were bundled in one step: that the control plane can describe the mesh
correctly, and that the catalogue holds a complete record of what was built. The
first passes; the second is novox/hq issue 050. Bundled, one open fault stopped
three later steps from ever being attempted, which is exactly the information the
run existed to produce.

The catalogue is now measured against what the control plane ordered rather than
against a list written in the test — a catalogue cannot know what it was never
told, so it has to be compared with something that does.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:17:06 +02:00
jschoubben 3690e17cf4 Ask the catalogue as the only thing that is allowed to
Invoking a tool from inside the module's own container is refused: its account is
scoped to what it emits and consumes, and a tool call needs a reply queue. Filed
as novox/hq issue 049 — the account is right, the request is reasonable, and
nothing can make it.

Until that is decided the caller is the substrate's bootstrap admin over the
broker's loopback, reached by joining its network namespace the way genesis
reaches a substrate container. Recorded in the step as the workaround it is,
rather than left looking like how a mesh is meant to be asked a question.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:08:52 +02:00
jschoubben 572aae2eb4 A plan with codes, stated up front and reported against
Steps were identified by their own sentences, so 'which one failed' meant reading
prose, and rewording a step silently made it a different step with no history.
Each now carries a stable code: R for raising the mesh, P for it being able to
produce, U for something being used on it, V for verifying what it says about
itself, E for enduring — a change following on its own, and coming back after the
machine stops.

The plan is data, printed before anything is attempted, so a reader knows what
the run intends to establish rather than inferring it from what happens to be
printed. The run ends with a table and a JSON report, and distinguishes SKIP from
FAIL: a step whose dependency failed was never asked, which is not the same as a
step that was asked and said no.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 22:00:13 +02:00
jschoubben 2f470f27b8 Genesis makes six claims; assert them separately
Genesis is twelve steps and was reported as one line, so a failure said nothing
about which claim broke and a pass was one tick standing in for six things being
true: the substrate up, the control plane built rather than handed over, the
pivot finished, the registry serving what was published into it, the machine
enrolled with an agent actually running, and the builder installed as a module.

Each is asked of the machine rather than read from the installer's own output.
The installer saying it published an image and the registry serving one are
different facts, and only the second matters.

And the catalogue is invoked by its real entrypoint. 'mesh-tools' is not on PATH
in the runtime image; the image runs 'node dist/main.js', and the invoke mode is
missing from the header comment that says there are three modes.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:58:58 +02:00
jschoubben 7641bb2059 A one-node mesh, and twelve things that have to be true of it
The common case, and the one that was never tested as a whole. What existed
asked whether four machines converged; it never asked whether ONE machine ends
up holding a mesh.

The order was wrong too. Three machines were enrolled second, into a mesh that
could not yet produce a single module, and that was reported as though something
had been shown. 17-raising-a-mesh is explicit: genesis ends with a mesh that
RUNS, and what remains after the core modules are built is "adding machines".
So the core comes first and machines arrive last — here, not at all, because a
second node is only meaningful once the first is complete.

Three things were missing entirely and nothing complained, because nothing asked:
the mesh never built its own catalogue, never had a store of its own for that
catalogue to use, and never rebuilt its own control plane through the module
path.

And four checks that were absent rather than failing:

  - it can describe itself — status, module list, plan --json, and the
    catalogue's five tools ASKED rather than observed. A container being up was
    being read as the catalogue working, which is the same error as matching a
    container by substring and finding the wrong one.
  - its networking is what the modules asked for — default closed, ssh open,
    declared ports open, .internal names written, module networks present. Left
    out altogether, which is hard to defend given the firewall work this week.
  - a change to a module's source reaches the machine on its own. The capability
    the migration depends on.
  - it comes back after a reboot. Never once tested; the lab had no way to
    restart a machine, because nothing had ever needed one.

Machines are named by role now — anchor, home-server, workstation, laptop — not
after the operator's own nodes, which made test output and real state hard to
tell apart.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:52:24 +02:00
jschoubben 9146f30859 Name the container, and issue the account
Waiting for "a container whose name contains lavinmq" was satisfied by the
broker — lavinmq, up and healthy — while the thing under test, the module's own
runtime mesh-lavinmq, crash-looped beside it. The step went green and the fault
was found by reading docker ps by hand. Containers are named exactly now, and a
failure prints that container's own last words.

And lavinmq gets a broker account, which it was never issued. Without one the
mesh still fills the secret the module declares it owns, with a generated value,
so the runtime starts, fails to parse a password as a credential document, and
loops on a JSON syntax error that mentions no missing account.

The two module-issue calls written .catch(() => {}) are not. That pattern has now
hidden three separate faults in this file.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:20:57 +02:00
jschoubben c3d5ec1d55 Build each repository from its own ref, and build the provider
lavinmq needs building now, so the step that assigns it builds it first.

And the ref is per repository rather than one value for all of them: a change
under test lives in one repository, and building the others from that branch
would prove it against itself.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 21:05:28 +02:00
jschoubben f04911a763 Give the module the broker it asks for, rather than a module that asks for nothing
The mesh refused to place amqp-ping: nothing provides amqp. That refusal is
right. The substrate raises a broker, but as a bundle resource — plumbing, not a
module the mesh has a record of — so it offers nothing to anything, and a module
wanting a broker wants one in the graph.

lavinmq is that module and needs no building, its image being upstream, so this
is a register and an assign. The alternative was to pick a module with no
requires, which would have passed by testing less.

Also: the control plane's image has no /tmp to copy a manifest into. Root does.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:56:35 +02:00
jschoubben 4ef0a19053 A manifest the control plane can open, and stop swallowing the failure when it cannot
mesh-control runs in a container, so a manifest pushed to the machine is not a
file it can read; `module add` said so plainly and it was briefly taken for a
missing manifest. It is copied the last step of the way now.

The base's registration was doing this too, and its failure was swallowed by a
bare catch on the reasoning that the module might already be known. The step
passed regardless — a base with nothing to stand on builds whether or not the
mesh holds a record of it — and the fault surfaced one step later, where the
record was needed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:46:11 +02:00
jschoubben babf08b9f8 Raise machines that are somebody, and a bed that hands over nothing
Every machine in a bed is a clone of one base image, so all of them booted with
the same /etc/machine-id. systemd's DHCP client derives its client identifier
from that file and dnsmasq keys leases on the identifier rather than the MAC, so
four machines with four distinct MACs were handed one address and the host kept
one ARP entry for it. Whichever machine last answered an ARP request received
everybody's replies.

This is the fault behind every run lost to "flaky lab DNS": resolution that works
two times in three, pulls that succeed on a retry, and one machine out of four
being fine while the rest have no path at all. It survived an earlier diagnosis
that blamed resolver ordering, because reordering resolvers on a machine that has
just won the ARP race looks exactly like a fix.

Each machine is now given its own machine-id before the uplink lease is asked
for, and a check after addresses are applied refuses to go on if two machines
took the same one — the positive control this never had, since the fault is
invisible where it happens and unrecognisable where it surfaces.

The egress check also now demands five consecutive lookups rather than one. A
single answer is what let a machine resolving one query in three pass and then
die twenty minutes later inside a pull.

And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The
four-machine bed loads thirty-four of the mesh's own images onto its machines
from the workstation because it does not build them, which is a shape no real
installation has and the same fiction the lab removed when it deleted its own
registry. This scenario names no images at all. The machines pull what is public,
the installer builds the control plane, and the mesh builds the rest — including,
last and deliberately, a module on a machine that did not build it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-14 20:41:55 +02:00
jschoubben 7fa1f1f0bb Merge pull request 'One archive per machine, not one per image' (#26) from fix/the-lab-ships-runtimes-not-toolchains into main 2026-09-14 19:58:20 +02:00
jschoubben 5d1f7762c2 One archive per machine, not one per image
The images a bed stocks are almost entirely the same bytes: the base they share
is 227 MB and a module's own code is a few. Exported one at a time that base is
written, pushed and loaded once per image — for this bed, the same 227 MB crossed
thirty-odd times, and the whole set measured 9.7 GB.

Measured on eight of them: 2.24 GB as separate archives, 0.29 GB as one. 87 per
cent less, and it improves with the count.

This is the slowest thing a raise does, and the four-machine bed has been
exceeding its own ninety-minute limit while still copying — so it was failing on
the clock rather than on anything it was testing.

Also stops shipping a compiler in every module image. The image runs compiled
code and never compiles any; tsc runs on the workstation. Worth 26 MB an image,
which is small beside the above but was pure waste.
2026-09-14 19:58:01 +02:00
jschoubben 8fca04f560 Merge pull request 'Configure the resolver rather than fight it' (#25) from fix/configure-the-resolver-rather-than-fight-it into main 2026-09-14 18:45:58 +02:00
jschoubben fa5e5cb912 Configure the resolver rather than fight it
The first attempt wrote /etc/resolv.conf. On these images that is a symlink
owned by systemd-resolved, so the file is either reverted or the link is broken
— found by reading a running machine instead of assuming the change had worked.

The real shape shows in resolvectl: the machine has sensible global fallbacks,
and the link carrying the default route has exactly one server, the uplink
gateway. resolved will not reach a global fallback while the link has a server
of its own, so one unanswered packet is one failed lookup. Three runs have died
that way, each long after the egress check passed.

The uplink stays first, so the modelled path is still what is used and still
what the check proves. Verified on a live machine: three servers on the link,
uplink first, resolution intact.
2026-09-14 18:45:40 +02:00
jschoubben eed5647146 Merge pull request 'One dropped lookup should not cost a two-hour run' (#24) from fix/one-dropped-lookup-should-not-cost-a-run into main 2026-09-14 18:23:49 +02:00