Files
hq/01-RESEARCH/002-local-mesh/analysis.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

264 lines
13 KiB
Markdown

---
effort: 002-local-mesh
updated: 2026-08-22
---
# A mesh that runs locally — current state and obstacles
Evidence for Phase 0. Every claim here is either a file location or something measured on
2026-08-22; where a claim was checked and found false, that is recorded too.
---
## 1. What exists today, and why neither is a mesh
### `test/pipeline/` — the only containerised HAL node, and it is dead
This harness builds `@hal/sdk` and the `hal/brain` daemon into a `node:23-alpine` image,
compiles `hal/coordinator`'s tools against it, and runs `brainstem.js` with
`HAL_MODE=daemon` against a containerised postgres and LavinMQ
(`test/pipeline/Dockerfile`, `test/pipeline/docker-compose.yml`). It proves the valuable
thing: **a node runtime needs no systemd, no `/services/`, and no nvm** — environment
variables and a seeded `nodes` row are sufficient.
It also cannot build. Measured 2026-08-22:
```
Step 7/16 : RUN npm run build -w modules/hal/sdk
npm error No workspaces found:
npm error --workspace=modules/hal/sdk
```
The npm workspace it depends on was removed on 2026-06-04 (`21ef4a4e`, "kill npm
workspace"), and the root `package.json` now carries a comment explaining why it will not
come back. The harness has not been touched since before that commit.
So the repository's only end-to-end pipeline test has been unrunnable for two and a half
months and nothing reported it — which is the same shape as the faults Phase 0 exists to
catch. **A test nobody runs is indistinguishable from a test that passes.**
### `test/dev-mesh/` — a work plane, not a control plane
Four `runtime-*` services, a noxflow API, a meshboard and a shared postgres + LavinMQ on
one network (`test/dev-mesh/docker-compose.yml`). It is genuinely useful and the topology
is a good precedent: one network, per-service `HAL_NODE`, `EXECUTOR_STUB=1` so dispatch and
run lifecycle are real while the model call is faked.
It is the wrong layer for Phase 0, on three counts:
| | dev-mesh does | Phase 0 needs |
|---|---|---|
| state | restores `pg_dump`s (`test/dev-mesh/init.sh`) | schema built by **migrations**, or fixture 3 is untestable by construction |
| code | bind-mounts pre-built `dist/` from the host | artifacts **built and delivered** by the pipeline |
| scope | noxflow runtime, API, meshboard | `hal/coordinator`, `hal/meshware`, MinIO, provisioning |
No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises
what the mesh *runs*; Phase 0 needs what the mesh *is*.
### `dev_up` — confirmed host-coupled
The claim in the work breakdown holds. `startDevProvider()` starts providers through the
host's user systemd instance and reads credentials from a host path
(`modules/hal/developer/tools/dev-env.ts:146-178`):
```ts
const unitName = `hal-module@${provider}.service`;
execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" });
const creds = resolveRunningProviderCreds(provider); // reads /services/<provider>/.env
```
It borrows the host. There is no mesh to stand up.
---
## 2. What is already portable
- **Node identity is one environment variable.** `resolveNodeName()` is
`process.env.HAL_NODE || hostname()`, lowercased (`modules/hal/sdk/src/mesh-config.ts:215`).
No file, no registration handshake, no host coupling.
- **Registration is a plain upsert.** `mesh_node_register` inserts into `nodes`
(`modules/hal/mesh/tools/index.ts:490-540`); mandatory fields are `name` and `user_name`.
- **The daemons are AMQP loops.** `hal/meshware` consumes `cmd.feature.{build,install,configure,start}`
(`modules/hal/meshware/daemon/src/cerebellum.ts:780-857`); `hal/coordinator` consumes
`event.gitea.push` and drives the cascade. Both are configured entirely by environment.
- **`hal/brain` in daemon mode already runs in a plain node image**, per `test/pipeline/`.
---
## 3. The obstacles, located
Four couplings stand between the daemons and a container. All are in code, none are
mysterious.
| # | coupling | location |
|---|---|---|
| 1 | meshware restarts itself via the host: `execFileSync("systemctl", ["--user","restart","hal-meshware.service"])` | `modules/hal/meshware/daemon/src/cerebellum.ts:826` |
| 2 | meshware restarts the broker via a hardcoded host path: `execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"])` | `cerebellum.ts:843-844` |
| 3 | **a module service *is* a systemd unit** — `hal-module@.service` runs `docker compose --project-directory /services/%i` | `modules/hal/meshware/systemd/hal-module@.service:10-11` |
| 4 | env generation writes `homedir()`-relative host paths — `~/.config/hal/env`, `~/.config/hal/modules/*.env`, `/services/*/.env` | `modules/hal/sdk/src/env-generator.ts:188-223`, `:538-568`, `:584-606` |
1, 2 and 4 are small: a guard and a configurable root. **3 is the architectural one**, and
it is the first open question below — in a container, "systemd unit that runs docker
compose" has no natural translation.
The bootstrap scripts add their own host assumptions — nvm for node resolution, `sudo
mkdir -p /services && chown`, and a `~/dotfiles` repository that is not in this repo
(`install.d/adopt.sh:165-170`) — but Phase 0 does not have to run them. A container image
can be built the way `test/pipeline/Dockerfile` builds one, bypassing the bootstrap path
entirely. That is a deliberate divergence to record: **the local mesh would not test the
bootstrap**, only the running mesh.
---
## 4. External services, and whether they can be local
| service | used for | local substitute |
|---|---|---|
| PostgreSQL | registry + pipeline state | yes — already the pattern in both harnesses |
| LavinMQ | all coordinator↔meshware messaging | yes — `cloudamqp/lavinmq`, already used |
| MinIO | per-feature artifact tarballs, `modules/{name}/{version}/{feature}.tar.gz` (`modules/hal/sdk/src/artifact-manager.ts:21,65`) | yes — `minio/minio`, needs only `REGISTRY_MINIO_*` |
| Gitea | source, push webhook, **and** the `@hal/*` npm registry | container exists, but heavy; see open question 2 |
| Docker registry | `docker push` for modules declaring `docker:` (`build-executor.ts:195-231`) | `registry:2`, or avoid modules that need it |
| Traefik | writes routing config at install; does not gate success | omit |
Only Gitea is awkward, and only because it carries two roles at once.
---
## 5. The fixtures
Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they
differ sharply in cost.
### Fixture B — a provider deploy rotates a shared credential without fanning out
**Cheapest, best understood, and the root cause is still open.** Documented at
`troubleshooting/provision-adoption-rotates-live-credential`. The mechanism is two lines:
```ts
async provision(project, username, options) {
const password = generatePassword(); // ALWAYS a fresh password
```
and, in the adoption branch of `provisionDatabase`, an `ALTER ROLE ... WITH PASSWORD` that
rotates the live secret while updating only the provider's own `mesh_provisions` row
(`modules/postgres/tools/index.ts`). Every consumer sharing that role keeps a stale
password and fails permanently; the provider recovers alone, and that asymmetry is the
tell.
Reproduction needs one provider node and two consumer nodes — the minimum interesting
mesh. It re-fires on every provider deploy, so it does not need to be provoked, only
observed. Confirmation is a timestamp comparison, not a hash: the stored values are
`enc:v1:` with a random IV, so identical plaintexts hash differently.
Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors
each; the provider failed for 22 minutes and recovered by itself.
### Fixture C — a migration ships nothing while the pipeline reports success
Three independent silent-skip points, any one of which produces it:
1. **Not-applicable and succeeded are the same status.** `MigrationsHandler.detect()` is
`existsSync(join(moduleDir, "migrations"))` against the freshly cloned workspace
(`modules/hal/sdk/src/feature-handlers/migrations.ts:32-41`). If the directory is not
there — uncommitted, ignored, or a `module_path` that does not line up — the build
reports `success` (`cerebellum.ts:634-638`).
2. **Upload failure is a warning.** `uploadFeatureArtifact()` returns `{sha256: ""}` with a
`warn` when none of the listed files exist (`artifact-manager.ts:41-44`), and callers
catch and warn rather than throw (`cerebellum.ts:672-673`). The symmetric download
failure on the target node is also non-fatal (`artifact-manager.ts:93-98`,
`cerebellum.ts:426-431`, commented "Non-fatal: feature may work without artifact").
3. **Missing packaged files are skipped in silence.** `addFileEntries()` does
`if (!existsSync(abs)) continue;` (`module-builder-core.ts:134-135`) — an explicit
`package:` entry that is not on disk simply is not in the tarball.
This is the same class as `troubleshooting/flavors-never-packaged` (`flavors/` was never
staged, so a node kept its first copy forever) and is catalogued in
`troubleshooting/deploy-reports-transport-not-effect`, which counted **0 of 119 modules
implementing `verify`** — the stage that exists, is dispatched, has a working handler, and
would turn every one of these into a red pipeline.
### Fixture A — a credential rotation does not reach a running session
The most valuable and the most work: it needs a *running session* to rotate underneath,
which means the local mesh must be able to start one. The mechanism is understood — an
`EnvironmentFile` is read once and `process.env` is a snapshot — and it is why credentials
became files that get re-read. Defer it behind B and C.
**Recommendation:** B first. It needs three nodes and a provision, no build and no session,
and it is the one whose root cause is still open — so reproducing it locally has value
beyond proving the harness.
---
## 6. A claim that was checked and found false
The survey behind this document suspected that the live AMQP pipeline never writes
`deployments` or `node_modules.installed_version`, because the code that does so sits in
`installer-core.ts:1045-1126` on what is commented as the "CLI path".
Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours:
| node | deploys | newest |
|---|---|---|
| 1 | 24 | 17:26:17 |
| 2 | 27 | 17:22:22 |
| 3 | 47 | 17:25:54 |
| 4 | 26 | 17:22:23 |
The record is live on all four nodes and minutes old. **The claim is refuted.** Recorded
because the reasoning was plausible and someone will retrace it.
---
## 7. Questions
### Settled (jochen, 2026-08-22)
**Phase 0 is a development environment, not a fixture rig.** It is built as a supported
surface to work in daily, and it is what `dev_up` should become. The definition of done —
"demonstrated in the local mesh" — is therefore meant literally.
**The trigger is a real Gitea container.** The webhook relay is part of what is under test,
including the changed-file enrichment that exists because the webhook truncates at 20
commits (`modules/hal/gitea/tools/index.ts:136-195`). A synthetic `event.gitea.push` would
skip it. Gitea also carries the `@hal/*` npm registry, so the local mesh needs it twice
over either way.
### Open — blocking
1. **How does a node supervise a module service?** Today a module service is a systemd unit
running `docker compose` against `/services/<module>`
(`modules/hal/meshware/systemd/hal-module@.service:10-11`), which has no direct
translation inside a container. The question raised in response is the better one, and
is broader than Phase 0: **what would it cost to stop using systemd altogether and have
the mesh supervise its own services?** That is a mesh-level architectural question, not
a local-mesh implementation detail — if the answer is that the mesh should own
supervision, Phase 0 should not build a container-only workaround first. Under
investigation; findings will land in [`003-service-supervision`](../003-service-supervision/)
and, if it goes ahead, a decision record of its own.
2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable
root keeps one code path; container-specific targets keep the host paths untouched. This
decides whether env generation gets a seam or a conditional.
3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap
scripts, and it will run one OS where the real mesh is heterogeneous by design
(`00-META/context.md`). Stating the divergence up front is what stops "it works
locally" from becoming its own class of silent failure. Now sharper, because a
development environment people use daily is trusted far more than a rig — and drifting
from production costs correspondingly more.
---
## References
- [`02-DECISIONS/0001`](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this
phase unblocks
- [`03-DESIGN/00-work-breakdown.md`](../../03-DESIGN/01-to-be/00-work-breakdown.md) — Phase 0 tasks
and checkpoint
- `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open
- `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped
green, and the unused `verify` stage
- `troubleshooting/flavors-never-packaged` — fixture C, previously seen in the wild