HQ — the mesh's own documentation

What the mesh is, what it is becoming, and why. Implementation lives in the
code repositories; the reasoning lives here.

  00-GENESIS   mission, engineering context, effect, and the rules that hold
  01-RESEARCH  investigations, before they harden into design
  02-DESIGN    the authoritative specification
  adr          numbered decisions — what was chosen, and what was rejected
  DECISIONS.md the ledger: every decision, in the order it was taken

Written for a reader who is not its author and has no access to the mesh it
describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918;
nodes are named by role.

Single initial commit by intent. The prior history came from a private
repository and carried operational detail — a routable address identified as a
VPN hub, real domain names, a hosting provider — which sanitising a tip commit
would not have removed from the log.
This commit is contained in:
2026-08-22 22:01:32 +02:00
commit cf9357e8e9
22 changed files with 2376 additions and 0 deletions
+258
View File
@@ -0,0 +1,258 @@
# A mesh that runs locally — current state and obstacles
Evidence for Phase 0. Every claim here is either a file location or something measured on
2026-08-22; where a claim was checked and found false, that is recorded too.
---
## 1. What exists today, and why neither is a mesh
### `test/pipeline/` — the only containerised HAL node, and it is dead
This harness builds `@hal/sdk` and the `hal/brain` daemon into a `node:23-alpine` image,
compiles `hal/coordinator`'s tools against it, and runs `brainstem.js` with
`HAL_MODE=daemon` against a containerised postgres and LavinMQ
(`test/pipeline/Dockerfile`, `test/pipeline/docker-compose.yml`). It proves the valuable
thing: **a node runtime needs no systemd, no `/services/`, and no nvm** — environment
variables and a seeded `nodes` row are sufficient.
It also cannot build. Measured 2026-08-22:
```
Step 7/16 : RUN npm run build -w modules/hal/sdk
npm error No workspaces found:
npm error --workspace=modules/hal/sdk
```
The npm workspace it depends on was removed on 2026-06-04 (`21ef4a4e`, "kill npm
workspace"), and the root `package.json` now carries a comment explaining why it will not
come back. The harness has not been touched since before that commit.
So the repository's only end-to-end pipeline test has been unrunnable for two and a half
months and nothing reported it — which is the same shape as the faults Phase 0 exists to
catch. **A test nobody runs is indistinguishable from a test that passes.**
### `test/dev-mesh/` — a work plane, not a control plane
Four `runtime-*` services, a noxflow API, a meshboard and a shared postgres + LavinMQ on
one network (`test/dev-mesh/docker-compose.yml`). It is genuinely useful and the topology
is a good precedent: one network, per-service `HAL_NODE`, `EXECUTOR_STUB=1` so dispatch and
run lifecycle are real while the model call is faked.
It is the wrong layer for Phase 0, on three counts:
| | dev-mesh does | Phase 0 needs |
|---|---|---|
| state | restores `pg_dump`s (`test/dev-mesh/init.sh`) | schema built by **migrations**, or fixture 3 is untestable by construction |
| code | bind-mounts pre-built `dist/` from the host | artifacts **built and delivered** by the pipeline |
| scope | noxflow runtime, API, meshboard | `hal/coordinator`, `hal/meshware`, MinIO, provisioning |
No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises
what the mesh *runs*; Phase 0 needs what the mesh *is*.
### `dev_up` — confirmed host-coupled
The claim in the work breakdown holds. `startDevProvider()` starts providers through the
host's user systemd instance and reads credentials from a host path
(`modules/hal/developer/tools/dev-env.ts:146-178`):
```ts
const unitName = `hal-module@${provider}.service`;
execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" });
const creds = resolveRunningProviderCreds(provider); // reads /services/<provider>/.env
```
It borrows the host. There is no mesh to stand up.
---
## 2. What is already portable
- **Node identity is one environment variable.** `resolveNodeName()` is
`process.env.HAL_NODE || hostname()`, lowercased (`modules/hal/sdk/src/mesh-config.ts:215`).
No file, no registration handshake, no host coupling.
- **Registration is a plain upsert.** `mesh_node_register` inserts into `nodes`
(`modules/hal/mesh/tools/index.ts:490-540`); mandatory fields are `name` and `user_name`.
- **The daemons are AMQP loops.** `hal/meshware` consumes `cmd.feature.{build,install,configure,start}`
(`modules/hal/meshware/daemon/src/cerebellum.ts:780-857`); `hal/coordinator` consumes
`event.gitea.push` and drives the cascade. Both are configured entirely by environment.
- **`hal/brain` in daemon mode already runs in a plain node image**, per `test/pipeline/`.
---
## 3. The obstacles, located
Four couplings stand between the daemons and a container. All are in code, none are
mysterious.
| # | coupling | location |
|---|---|---|
| 1 | meshware restarts itself via the host: `execFileSync("systemctl", ["--user","restart","hal-meshware.service"])` | `modules/hal/meshware/daemon/src/cerebellum.ts:826` |
| 2 | meshware restarts the broker via a hardcoded host path: `execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"])` | `cerebellum.ts:843-844` |
| 3 | **a module service *is* a systemd unit** — `hal-module@.service` runs `docker compose --project-directory /services/%i` | `modules/hal/meshware/systemd/hal-module@.service:10-11` |
| 4 | env generation writes `homedir()`-relative host paths — `~/.config/hal/env`, `~/.config/hal/modules/*.env`, `/services/*/.env` | `modules/hal/sdk/src/env-generator.ts:188-223`, `:538-568`, `:584-606` |
1, 2 and 4 are small: a guard and a configurable root. **3 is the architectural one**, and
it is the first open question below — in a container, "systemd unit that runs docker
compose" has no natural translation.
The bootstrap scripts add their own host assumptions — nvm for node resolution, `sudo
mkdir -p /services && chown`, and a `~/dotfiles` repository that is not in this repo
(`install.d/adopt.sh:165-170`) — but Phase 0 does not have to run them. A container image
can be built the way `test/pipeline/Dockerfile` builds one, bypassing the bootstrap path
entirely. That is a deliberate divergence to record: **the local mesh would not test the
bootstrap**, only the running mesh.
---
## 4. External services, and whether they can be local
| service | used for | local substitute |
|---|---|---|
| PostgreSQL | registry + pipeline state | yes — already the pattern in both harnesses |
| LavinMQ | all coordinator↔meshware messaging | yes — `cloudamqp/lavinmq`, already used |
| MinIO | per-feature artifact tarballs, `modules/{name}/{version}/{feature}.tar.gz` (`modules/hal/sdk/src/artifact-manager.ts:21,65`) | yes — `minio/minio`, needs only `REGISTRY_MINIO_*` |
| Gitea | source, push webhook, **and** the `@hal/*` npm registry | container exists, but heavy; see open question 2 |
| Docker registry | `docker push` for modules declaring `docker:` (`build-executor.ts:195-231`) | `registry:2`, or avoid modules that need it |
| Traefik | writes routing config at install; does not gate success | omit |
Only Gitea is awkward, and only because it carries two roles at once.
---
## 5. The fixtures
Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they
differ sharply in cost.
### Fixture B — a provider deploy rotates a shared credential without fanning out
**Cheapest, best understood, and the root cause is still open.** Documented at
`troubleshooting/provision-adoption-rotates-live-credential`. The mechanism is two lines:
```ts
async provision(project, username, options) {
const password = generatePassword(); // ALWAYS a fresh password
```
and, in the adoption branch of `provisionDatabase`, an `ALTER ROLE ... WITH PASSWORD` that
rotates the live secret while updating only the provider's own `mesh_provisions` row
(`modules/postgres/tools/index.ts`). Every consumer sharing that role keeps a stale
password and fails permanently; the provider recovers alone, and that asymmetry is the
tell.
Reproduction needs one provider node and two consumer nodes — the minimum interesting
mesh. It re-fires on every provider deploy, so it does not need to be provoked, only
observed. Confirmation is a timestamp comparison, not a hash: the stored values are
`enc:v1:` with a random IV, so identical plaintexts hash differently.
Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors
each; the provider failed for 22 minutes and recovered by itself.
### Fixture C — a migration ships nothing while the pipeline reports success
Three independent silent-skip points, any one of which produces it:
1. **Not-applicable and succeeded are the same status.** `MigrationsHandler.detect()` is
`existsSync(join(moduleDir, "migrations"))` against the freshly cloned workspace
(`modules/hal/sdk/src/feature-handlers/migrations.ts:32-41`). If the directory is not
there — uncommitted, ignored, or a `module_path` that does not line up — the build
reports `success` (`cerebellum.ts:634-638`).
2. **Upload failure is a warning.** `uploadFeatureArtifact()` returns `{sha256: ""}` with a
`warn` when none of the listed files exist (`artifact-manager.ts:41-44`), and callers
catch and warn rather than throw (`cerebellum.ts:672-673`). The symmetric download
failure on the target node is also non-fatal (`artifact-manager.ts:93-98`,
`cerebellum.ts:426-431`, commented "Non-fatal: feature may work without artifact").
3. **Missing packaged files are skipped in silence.** `addFileEntries()` does
`if (!existsSync(abs)) continue;` (`module-builder-core.ts:134-135`) — an explicit
`package:` entry that is not on disk simply is not in the tarball.
This is the same class as `troubleshooting/flavors-never-packaged` (`flavors/` was never
staged, so a node kept its first copy forever) and is catalogued in
`troubleshooting/deploy-reports-transport-not-effect`, which counted **0 of 119 modules
implementing `verify`** — the stage that exists, is dispatched, has a working handler, and
would turn every one of these into a red pipeline.
### Fixture A — a credential rotation does not reach a running session
The most valuable and the most work: it needs a *running session* to rotate underneath,
which means the local mesh must be able to start one. The mechanism is understood — an
`EnvironmentFile` is read once and `process.env` is a snapshot — and it is why credentials
became files that get re-read. Defer it behind B and C.
**Recommendation:** B first. It needs three nodes and a provision, no build and no session,
and it is the one whose root cause is still open — so reproducing it locally has value
beyond proving the harness.
---
## 6. A claim that was checked and found false
The survey behind this document suspected that the live AMQP pipeline never writes
`deployments` or `node_modules.installed_version`, because the code that does so sits in
`installer-core.ts:1045-1126` on what is commented as the "CLI path".
Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours:
| node | deploys | newest |
|---|---|---|
| 1 | 24 | 17:26:17 |
| 2 | 27 | 17:22:22 |
| 3 | 47 | 17:25:54 |
| 4 | 26 | 17:22:23 |
The record is live on all four nodes and minutes old. **The claim is refuted.** Recorded
because the reasoning was plausible and someone will retrace it.
---
## 7. Questions
### Settled (jochen, 2026-08-22)
**Phase 0 is a development environment, not a fixture rig.** It is built as a supported
surface to work in daily, and it is what `dev_up` should become. The definition of done —
"demonstrated in the local mesh" — is therefore meant literally.
**The trigger is a real Gitea container.** The webhook relay is part of what is under test,
including the changed-file enrichment that exists because the webhook truncates at 20
commits (`modules/hal/gitea/tools/index.ts:136-195`). A synthetic `event.gitea.push` would
skip it. Gitea also carries the `@hal/*` npm registry, so the local mesh needs it twice
over either way.
### Open — blocking
1. **How does a node supervise a module service?** Today a module service is a systemd unit
running `docker compose` against `/services/<module>`
(`modules/hal/meshware/systemd/hal-module@.service:10-11`), which has no direct
translation inside a container. The question raised in response is the better one, and
is broader than Phase 0: **what would it cost to stop using systemd altogether and have
the mesh supervise its own services?** That is a mesh-level architectural question, not
a local-mesh implementation detail — if the answer is that the mesh should own
supervision, Phase 0 should not build a container-only workaround first. Under
investigation; findings will land in [`003-service-supervision`](../003-service-supervision/)
and, if it goes ahead, ADR 0002.
2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable
root keeps one code path; container-specific targets keep the host paths untouched. This
decides whether env generation gets a seam or a conditional.
3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap
scripts, and it will run one OS where the real mesh is heterogeneous by design
(`00-GENESIS/context.md`). Stating the divergence up front is what stops "it works
locally" from becoming its own class of silent failure. Now sharper, because a
development environment people use daily is trusted far more than a rig — and drifting
from production costs correspondingly more.
---
## References
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this
phase unblocks
- [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) — Phase 0 tasks
and checkpoint
- `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open
- `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped
green, and the unused `verify` stage
- `troubleshooting/flavors-never-packaged` — fixture C, previously seen in the wild
+53
View File
@@ -0,0 +1,53 @@
# 002 — A mesh that runs locally
- **Status:** GRADUATED — the design is [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md)
- **Initiated by:** jochen, 2026-08-22
- **Areas touched:** `install.d/`, `hal/meshware`, `hal/coordinator`, `hal/brain`,
`hal/developer` (`dev_up`), `hal/sdk` (env generation, feature handlers, artifact
manager), `test/pipeline/`, `test/dev-mesh/`, the provisioning path in
`modules/postgres/`.
## Summary
Phase 0 of [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) requires
a mesh that comes up in containers, runs its own pipeline, and reproduces known faults on
demand. Nothing else in the decomposition starts until it exists, because every fault the
decomposition addresses was found in production — there was nowhere else to find it.
This effort establishes what already runs in a container, what is welded to the host, and
what it would take to close the gap. It does **not** choose an approach: the central
question — how a containerised node executes a module service, when a module service is
defined today as a systemd unit shelling to `docker compose` in `/services/` — is not
answered by ADR 0001 and is recorded below rather than decided.
## What was established
- The two existing container harnesses are **neither of them a mesh**, and one of them has
not been able to build since 2026-06-04.
- Node identity is already portable — a single environment variable, no host handshake.
- Four concrete host couplings block a containerised node, all with known locations.
- All three Phase 0 fixtures are reproducible; one of them is documented in the knowledge
base with an open root cause and is the cheapest place to start.
Detail and evidence in [`analysis.md`](analysis.md).
## Questions — all settled 2026-08-22
**Phase 0 is a development environment**, not a fixture rig — it is what the host-borrowing
dev tooling becomes.
**The trigger is a real source-forge container**, because the webhook relay is part of what
is under test.
**A node is a system container**, promotable to a virtual machine per node. This answered
the question the effort was stuck on, and dissolved it rather than solving it: against a
real node with a real init, the four host couplings catalogued in `analysis.md` §3 are not
couplings — they are how a node works. They were obstacles only to a node modelled as an
application container.
Consequently [`003-service-supervision`](../003-service-supervision/) **no longer blocks
Phase 0**. It remains a live architecture question, on its own timeline.
The remaining questions in `analysis.md` — what replaces host paths in a container, and how
faithful the lab must be — are answered in the design: nothing replaces them, because the
paths are real; and the divergences are enumerated rather than discovered.