Files
hq/01-RESEARCH/002-local-mesh/analysis.md
T
jschoubben 702efca6bb Base layer: the mesh as it is, under the mesh as it should be
HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
2026-08-23 03:08:26 +02:00

13 KiB

effort, updated
effort updated
002-local-mesh 2026-08-22

A mesh that runs locally — current state and obstacles

Evidence for Phase 0. Every claim here is either a file location or something measured on 2026-08-22; where a claim was checked and found false, that is recorded too.


1. What exists today, and why neither is a mesh

test/pipeline/ — the only containerised HAL node, and it is dead

This harness builds @hal/sdk and the hal/brain daemon into a node:23-alpine image, compiles hal/coordinator's tools against it, and runs brainstem.js with HAL_MODE=daemon against a containerised postgres and LavinMQ (test/pipeline/Dockerfile, test/pipeline/docker-compose.yml). It proves the valuable thing: a node runtime needs no systemd, no /services/, and no nvm — environment variables and a seeded nodes row are sufficient.

It also cannot build. Measured 2026-08-22:

Step 7/16 : RUN npm run build -w modules/hal/sdk
npm error No workspaces found:
npm error   --workspace=modules/hal/sdk

The npm workspace it depends on was removed on 2026-06-04 (21ef4a4e, "kill npm workspace"), and the root package.json now carries a comment explaining why it will not come back. The harness has not been touched since before that commit.

So the repository's only end-to-end pipeline test has been unrunnable for two and a half months and nothing reported it — which is the same shape as the faults Phase 0 exists to catch. A test nobody runs is indistinguishable from a test that passes.

test/dev-mesh/ — a work plane, not a control plane

Four runtime-* services, a noxflow API, a meshboard and a shared postgres + LavinMQ on one network (test/dev-mesh/docker-compose.yml). It is genuinely useful and the topology is a good precedent: one network, per-service HAL_NODE, EXECUTOR_STUB=1 so dispatch and run lifecycle are real while the model call is faked.

It is the wrong layer for Phase 0, on three counts:

dev-mesh does Phase 0 needs
state restores pg_dumps (test/dev-mesh/init.sh) schema built by migrations, or fixture 3 is untestable by construction
code bind-mounts pre-built dist/ from the host artifacts built and delivered by the pipeline
scope noxflow runtime, API, meshboard hal/coordinator, hal/meshware, MinIO, provisioning

No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises what the mesh runs; Phase 0 needs what the mesh is.

dev_up — confirmed host-coupled

The claim in the work breakdown holds. startDevProvider() starts providers through the host's user systemd instance and reads credentials from a host path (modules/hal/developer/tools/dev-env.ts:146-178):

const unitName = `hal-module@${provider}.service`;
execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" });
const creds = resolveRunningProviderCreds(provider);  // reads /services/<provider>/.env

It borrows the host. There is no mesh to stand up.


2. What is already portable

  • Node identity is one environment variable. resolveNodeName() is process.env.HAL_NODE || hostname(), lowercased (modules/hal/sdk/src/mesh-config.ts:215). No file, no registration handshake, no host coupling.
  • Registration is a plain upsert. mesh_node_register inserts into nodes (modules/hal/mesh/tools/index.ts:490-540); mandatory fields are name and user_name.
  • The daemons are AMQP loops. hal/meshware consumes cmd.feature.{build,install,configure,start} (modules/hal/meshware/daemon/src/cerebellum.ts:780-857); hal/coordinator consumes event.gitea.push and drives the cascade. Both are configured entirely by environment.
  • hal/brain in daemon mode already runs in a plain node image, per test/pipeline/.

3. The obstacles, located

Four couplings stand between the daemons and a container. All are in code, none are mysterious.

# coupling location
1 meshware restarts itself via the host: execFileSync("systemctl", ["--user","restart","hal-meshware.service"]) modules/hal/meshware/daemon/src/cerebellum.ts:826
2 meshware restarts the broker via a hardcoded host path: execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"]) cerebellum.ts:843-844
3 a module service is a systemd unit — hal-module@.service runs docker compose --project-directory /services/%i modules/hal/meshware/systemd/hal-module@.service:10-11
4 env generation writes homedir()-relative host paths — ~/.config/hal/env, ~/.config/hal/modules/*.env, /services/*/.env modules/hal/sdk/src/env-generator.ts:188-223, :538-568, :584-606

1, 2 and 4 are small: a guard and a configurable root. 3 is the architectural one, and it is the first open question below — in a container, "systemd unit that runs docker compose" has no natural translation.

The bootstrap scripts add their own host assumptions — nvm for node resolution, sudo mkdir -p /services && chown, and a ~/dotfiles repository that is not in this repo (install.d/adopt.sh:165-170) — but Phase 0 does not have to run them. A container image can be built the way test/pipeline/Dockerfile builds one, bypassing the bootstrap path entirely. That is a deliberate divergence to record: the local mesh would not test the bootstrap, only the running mesh.


4. External services, and whether they can be local

service used for local substitute
PostgreSQL registry + pipeline state yes — already the pattern in both harnesses
LavinMQ all coordinator↔meshware messaging yes — cloudamqp/lavinmq, already used
MinIO per-feature artifact tarballs, modules/{name}/{version}/{feature}.tar.gz (modules/hal/sdk/src/artifact-manager.ts:21,65) yes — minio/minio, needs only REGISTRY_MINIO_*
Gitea source, push webhook, and the @hal/* npm registry container exists, but heavy; see open question 2
Docker registry docker push for modules declaring docker: (build-executor.ts:195-231) registry:2, or avoid modules that need it
Traefik writes routing config at install; does not gate success omit

Only Gitea is awkward, and only because it carries two roles at once.


5. The fixtures

Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they differ sharply in cost.

Fixture B — a provider deploy rotates a shared credential without fanning out

Cheapest, best understood, and the root cause is still open. Documented at troubleshooting/provision-adoption-rotates-live-credential. The mechanism is two lines:

async provision(project, username, options) {
  const password = generatePassword();   // ALWAYS a fresh password

and, in the adoption branch of provisionDatabase, an ALTER ROLE ... WITH PASSWORD that rotates the live secret while updating only the provider's own mesh_provisions row (modules/postgres/tools/index.ts). Every consumer sharing that role keeps a stale password and fails permanently; the provider recovers alone, and that asymmetry is the tell.

Reproduction needs one provider node and two consumer nodes — the minimum interesting mesh. It re-fires on every provider deploy, so it does not need to be provoked, only observed. Confirmation is a timestamp comparison, not a hash: the stored values are enc:v1: with a random IV, so identical plaintexts hash differently.

Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors each; the provider failed for 22 minutes and recovered by itself.

Fixture C — a migration ships nothing while the pipeline reports success

Three independent silent-skip points, any one of which produces it:

  1. Not-applicable and succeeded are the same status. MigrationsHandler.detect() is existsSync(join(moduleDir, "migrations")) against the freshly cloned workspace (modules/hal/sdk/src/feature-handlers/migrations.ts:32-41). If the directory is not there — uncommitted, ignored, or a module_path that does not line up — the build reports success (cerebellum.ts:634-638).
  2. Upload failure is a warning. uploadFeatureArtifact() returns {sha256: ""} with a warn when none of the listed files exist (artifact-manager.ts:41-44), and callers catch and warn rather than throw (cerebellum.ts:672-673). The symmetric download failure on the target node is also non-fatal (artifact-manager.ts:93-98, cerebellum.ts:426-431, commented "Non-fatal: feature may work without artifact").
  3. Missing packaged files are skipped in silence. addFileEntries() does if (!existsSync(abs)) continue; (module-builder-core.ts:134-135) — an explicit package: entry that is not on disk simply is not in the tarball.

This is the same class as troubleshooting/flavors-never-packaged (flavors/ was never staged, so a node kept its first copy forever) and is catalogued in troubleshooting/deploy-reports-transport-not-effect, which counted 0 of 119 modules implementing verify — the stage that exists, is dispatched, has a working handler, and would turn every one of these into a red pipeline.

Fixture A — a credential rotation does not reach a running session

The most valuable and the most work: it needs a running session to rotate underneath, which means the local mesh must be able to start one. The mechanism is understood — an EnvironmentFile is read once and process.env is a snapshot — and it is why credentials became files that get re-read. Defer it behind B and C.

Recommendation: B first. It needs three nodes and a provision, no build and no session, and it is the one whose root cause is still open — so reproducing it locally has value beyond proving the harness.


6. A claim that was checked and found false

The survey behind this document suspected that the live AMQP pipeline never writes deployments or node_modules.installed_version, because the code that does so sits in installer-core.ts:1045-1126 on what is commented as the "CLI path".

Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours:

node deploys newest
1 24 17:26:17
2 27 17:22:22
3 47 17:25:54
4 26 17:22:23

The record is live on all four nodes and minutes old. The claim is refuted. Recorded because the reasoning was plausible and someone will retrace it.


7. Questions

Settled (jochen, 2026-08-22)

Phase 0 is a development environment, not a fixture rig. It is built as a supported surface to work in daily, and it is what dev_up should become. The definition of done — "demonstrated in the local mesh" — is therefore meant literally.

The trigger is a real Gitea container. The webhook relay is part of what is under test, including the changed-file enrichment that exists because the webhook truncates at 20 commits (modules/hal/gitea/tools/index.ts:136-195). A synthetic event.gitea.push would skip it. Gitea also carries the @hal/* npm registry, so the local mesh needs it twice over either way.

Open — blocking

  1. How does a node supervise a module service? Today a module service is a systemd unit running docker compose against /services/<module> (modules/hal/meshware/systemd/hal-module@.service:10-11), which has no direct translation inside a container. The question raised in response is the better one, and is broader than Phase 0: what would it cost to stop using systemd altogether and have the mesh supervise its own services? That is a mesh-level architectural question, not a local-mesh implementation detail — if the answer is that the mesh should own supervision, Phase 0 should not build a container-only workaround first. Under investigation; findings will land in 003-service-supervision and, if it goes ahead, a decision record of its own.

  2. What replaces ~/.config/hal/env and /services/ inside a container? A configurable root keeps one code path; container-specific targets keep the host paths untouched. This decides whether env generation gets a seam or a conditional.

  3. How faithful must the local mesh be to be trusted? It will not run the bootstrap scripts, and it will run one OS where the real mesh is heterogeneous by design (00-GENESIS/context.md). Stating the divergence up front is what stops "it works locally" from becoming its own class of silent failure. Now sharper, because a development environment people use daily is trusted far more than a rig — and drifting from production costs correspondingly more.


References

  • adr/0001 — the decision this phase unblocks
  • 02-DESIGN/00-work-breakdown.md — Phase 0 tasks and checkpoint
  • troubleshooting/provision-adoption-rotates-live-credential — fixture B, root cause open
  • troubleshooting/deploy-reports-transport-not-effect — the nine defects that shipped green, and the unused verify stage
  • troubleshooting/flavors-never-packaged — fixture C, previously seen in the wild