--- effort: 002-local-mesh updated: 2026-08-22 --- # A mesh that runs locally — current state and obstacles Evidence for Phase 0. Every claim here is either a file location or something measured on 2026-08-22; where a claim was checked and found false, that is recorded too. --- ## 1. What exists today, and why neither is a mesh ### `test/pipeline/` — the only containerised HAL node, and it is dead This harness builds `@hal/sdk` and the `hal/brain` daemon into a `node:23-alpine` image, compiles `hal/coordinator`'s tools against it, and runs `brainstem.js` with `HAL_MODE=daemon` against a containerised postgres and LavinMQ (`test/pipeline/Dockerfile`, `test/pipeline/docker-compose.yml`). It proves the valuable thing: **a node runtime needs no systemd, no `/services/`, and no nvm** — environment variables and a seeded `nodes` row are sufficient. It also cannot build. Measured 2026-08-22: ``` Step 7/16 : RUN npm run build -w modules/hal/sdk npm error No workspaces found: npm error --workspace=modules/hal/sdk ``` The npm workspace it depends on was removed on 2026-06-04 (`21ef4a4e`, "kill npm workspace"), and the root `package.json` now carries a comment explaining why it will not come back. The harness has not been touched since before that commit. So the repository's only end-to-end pipeline test has been unrunnable for two and a half months and nothing reported it — which is the same shape as the faults Phase 0 exists to catch. **A test nobody runs is indistinguishable from a test that passes.** ### `test/dev-mesh/` — a work plane, not a control plane Four `runtime-*` services, a noxflow API, a meshboard and a shared postgres + LavinMQ on one network (`test/dev-mesh/docker-compose.yml`). It is genuinely useful and the topology is a good precedent: one network, per-service `HAL_NODE`, `EXECUTOR_STUB=1` so dispatch and run lifecycle are real while the model call is faked. It is the wrong layer for Phase 0, on three counts: | | dev-mesh does | Phase 0 needs | |---|---|---| | state | restores `pg_dump`s (`test/dev-mesh/init.sh`) | schema built by **migrations**, or fixture 3 is untestable by construction | | code | bind-mounts pre-built `dist/` from the host | artifacts **built and delivered** by the pipeline | | scope | noxflow runtime, API, meshboard | `hal/coordinator`, `hal/meshware`, MinIO, provisioning | No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises what the mesh *runs*; Phase 0 needs what the mesh *is*. ### `dev_up` — confirmed host-coupled The claim in the work breakdown holds. `startDevProvider()` starts providers through the host's user systemd instance and reads credentials from a host path (`modules/hal/developer/tools/dev-env.ts:146-178`): ```ts const unitName = `hal-module@${provider}.service`; execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" }); const creds = resolveRunningProviderCreds(provider); // reads /services//.env ``` It borrows the host. There is no mesh to stand up. --- ## 2. What is already portable - **Node identity is one environment variable.** `resolveNodeName()` is `process.env.HAL_NODE || hostname()`, lowercased (`modules/hal/sdk/src/mesh-config.ts:215`). No file, no registration handshake, no host coupling. - **Registration is a plain upsert.** `mesh_node_register` inserts into `nodes` (`modules/hal/mesh/tools/index.ts:490-540`); mandatory fields are `name` and `user_name`. - **The daemons are AMQP loops.** `hal/meshware` consumes `cmd.feature.{build,install,configure,start}` (`modules/hal/meshware/daemon/src/cerebellum.ts:780-857`); `hal/coordinator` consumes `event.gitea.push` and drives the cascade. Both are configured entirely by environment. - **`hal/brain` in daemon mode already runs in a plain node image**, per `test/pipeline/`. --- ## 3. The obstacles, located Four couplings stand between the daemons and a container. All are in code, none are mysterious. | # | coupling | location | |---|---|---| | 1 | meshware restarts itself via the host: `execFileSync("systemctl", ["--user","restart","hal-meshware.service"])` | `modules/hal/meshware/daemon/src/cerebellum.ts:826` | | 2 | meshware restarts the broker via a hardcoded host path: `execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"])` | `cerebellum.ts:843-844` | | 3 | **a module service *is* a systemd unit** — `hal-module@.service` runs `docker compose --project-directory /services/%i` | `modules/hal/meshware/systemd/hal-module@.service:10-11` | | 4 | env generation writes `homedir()`-relative host paths — `~/.config/hal/env`, `~/.config/hal/modules/*.env`, `/services/*/.env` | `modules/hal/sdk/src/env-generator.ts:188-223`, `:538-568`, `:584-606` | 1, 2 and 4 are small: a guard and a configurable root. **3 is the architectural one**, and it is the first open question below — in a container, "systemd unit that runs docker compose" has no natural translation. The bootstrap scripts add their own host assumptions — nvm for node resolution, `sudo mkdir -p /services && chown`, and a `~/dotfiles` repository that is not in this repo (`install.d/adopt.sh:165-170`) — but Phase 0 does not have to run them. A container image can be built the way `test/pipeline/Dockerfile` builds one, bypassing the bootstrap path entirely. That is a deliberate divergence to record: **the local mesh would not test the bootstrap**, only the running mesh. --- ## 4. External services, and whether they can be local | service | used for | local substitute | |---|---|---| | PostgreSQL | registry + pipeline state | yes — already the pattern in both harnesses | | LavinMQ | all coordinator↔meshware messaging | yes — `cloudamqp/lavinmq`, already used | | MinIO | per-feature artifact tarballs, `modules/{name}/{version}/{feature}.tar.gz` (`modules/hal/sdk/src/artifact-manager.ts:21,65`) | yes — `minio/minio`, needs only `REGISTRY_MINIO_*` | | Gitea | source, push webhook, **and** the `@hal/*` npm registry | container exists, but heavy; see open question 2 | | Docker registry | `docker push` for modules declaring `docker:` (`build-executor.ts:195-231`) | `registry:2`, or avoid modules that need it | | Traefik | writes routing config at install; does not gate success | omit | Only Gitea is awkward, and only because it carries two roles at once. --- ## 5. The fixtures Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they differ sharply in cost. ### Fixture B — a provider deploy rotates a shared credential without fanning out **Cheapest, best understood, and the root cause is still open.** Documented at `troubleshooting/provision-adoption-rotates-live-credential`. The mechanism is two lines: ```ts async provision(project, username, options) { const password = generatePassword(); // ALWAYS a fresh password ``` and, in the adoption branch of `provisionDatabase`, an `ALTER ROLE ... WITH PASSWORD` that rotates the live secret while updating only the provider's own `mesh_provisions` row (`modules/postgres/tools/index.ts`). Every consumer sharing that role keeps a stale password and fails permanently; the provider recovers alone, and that asymmetry is the tell. Reproduction needs one provider node and two consumer nodes — the minimum interesting mesh. It re-fires on every provider deploy, so it does not need to be provoked, only observed. Confirmation is a timestamp comparison, not a hash: the stored values are `enc:v1:` with a random IV, so identical plaintexts hash differently. Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors each; the provider failed for 22 minutes and recovered by itself. ### Fixture C — a migration ships nothing while the pipeline reports success Three independent silent-skip points, any one of which produces it: 1. **Not-applicable and succeeded are the same status.** `MigrationsHandler.detect()` is `existsSync(join(moduleDir, "migrations"))` against the freshly cloned workspace (`modules/hal/sdk/src/feature-handlers/migrations.ts:32-41`). If the directory is not there — uncommitted, ignored, or a `module_path` that does not line up — the build reports `success` (`cerebellum.ts:634-638`). 2. **Upload failure is a warning.** `uploadFeatureArtifact()` returns `{sha256: ""}` with a `warn` when none of the listed files exist (`artifact-manager.ts:41-44`), and callers catch and warn rather than throw (`cerebellum.ts:672-673`). The symmetric download failure on the target node is also non-fatal (`artifact-manager.ts:93-98`, `cerebellum.ts:426-431`, commented "Non-fatal: feature may work without artifact"). 3. **Missing packaged files are skipped in silence.** `addFileEntries()` does `if (!existsSync(abs)) continue;` (`module-builder-core.ts:134-135`) — an explicit `package:` entry that is not on disk simply is not in the tarball. This is the same class as `troubleshooting/flavors-never-packaged` (`flavors/` was never staged, so a node kept its first copy forever) and is catalogued in `troubleshooting/deploy-reports-transport-not-effect`, which counted **0 of 119 modules implementing `verify`** — the stage that exists, is dispatched, has a working handler, and would turn every one of these into a red pipeline. ### Fixture A — a credential rotation does not reach a running session The most valuable and the most work: it needs a *running session* to rotate underneath, which means the local mesh must be able to start one. The mechanism is understood — an `EnvironmentFile` is read once and `process.env` is a snapshot — and it is why credentials became files that get re-read. Defer it behind B and C. **Recommendation:** B first. It needs three nodes and a provision, no build and no session, and it is the one whose root cause is still open — so reproducing it locally has value beyond proving the harness. --- ## 6. A claim that was checked and found false The survey behind this document suspected that the live AMQP pipeline never writes `deployments` or `node_modules.installed_version`, because the code that does so sits in `installer-core.ts:1045-1126` on what is commented as the "CLI path". Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours: | node | deploys | newest | |---|---|---| | 1 | 24 | 17:26:17 | | 2 | 27 | 17:22:22 | | 3 | 47 | 17:25:54 | | 4 | 26 | 17:22:23 | The record is live on all four nodes and minutes old. **The claim is refuted.** Recorded because the reasoning was plausible and someone will retrace it. --- ## 7. Questions ### Settled (jochen, 2026-08-22) **Phase 0 is a development environment, not a fixture rig.** It is built as a supported surface to work in daily, and it is what `dev_up` should become. The definition of done — "demonstrated in the local mesh" — is therefore meant literally. **The trigger is a real Gitea container.** The webhook relay is part of what is under test, including the changed-file enrichment that exists because the webhook truncates at 20 commits (`modules/hal/gitea/tools/index.ts:136-195`). A synthetic `event.gitea.push` would skip it. Gitea also carries the `@hal/*` npm registry, so the local mesh needs it twice over either way. ### Open — blocking 1. **How does a node supervise a module service?** Today a module service is a systemd unit running `docker compose` against `/services/` (`modules/hal/meshware/systemd/hal-module@.service:10-11`), which has no direct translation inside a container. The question raised in response is the better one, and is broader than Phase 0: **what would it cost to stop using systemd altogether and have the mesh supervise its own services?** That is a mesh-level architectural question, not a local-mesh implementation detail — if the answer is that the mesh should own supervision, Phase 0 should not build a container-only workaround first. Under investigation; findings will land in [`003-service-supervision`](../003-service-supervision/) and, if it goes ahead, a decision record of its own. 2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable root keeps one code path; container-specific targets keep the host paths untouched. This decides whether env generation gets a seam or a conditional. 3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap scripts, and it will run one OS where the real mesh is heterogeneous by design (`00-META/context.md`). Stating the divergence up front is what stops "it works locally" from becoming its own class of silent failure. Now sharper, because a development environment people use daily is trusted far more than a rig — and drifting from production costs correspondingly more. --- ## References - [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — the decision this phase unblocks - [`03-DESIGN/00-work-breakdown.md`](../../03-DESIGN/01-to-be/00-work-breakdown.md) — Phase 0 tasks and checkpoint - `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open - `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped green, and the unused `verify` stage - `troubleshooting/flavors-never-packaged` — fixture C, previously seen in the wild