HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log.
This commit is contained in:
@@ -0,0 +1,237 @@
|
||||
# Current state → ideal state
|
||||
|
||||
## 1. The count is a symptom
|
||||
|
||||
124 modules. The number is not the problem; the reason for it is. Modules are split
|
||||
because splitting is the **only granularity lever the platform offers**:
|
||||
|
||||
| symptom | module | why it exists |
|
||||
|---|---|---|
|
||||
| one daemon must run on one node | `hal/claude-licences` | no per-node feature opt-in |
|
||||
| a library must not drag a 450 MB dep | `hal/claude` | no way to expose two npm packages |
|
||||
| a domain needs a daemon *and* a library | `hal/claude-code` + `hal/claude` | one feature of each type per module |
|
||||
| a handful of SDK verbs need a home | `infra` (10 lines) | tools must belong to *some* module |
|
||||
|
||||
A feature is a **handler type**, discovered by directory presence, singleton per module
|
||||
(`FEATURE_HANDLERS` in `modules/hal/sdk/src/feature-handlers/index.ts`). So "one more
|
||||
deployable unit" always means "one more module".
|
||||
|
||||
Splitting also fragments domains. Claude licence logic sits in three modules —
|
||||
resolution in `hal/claude`, health in `hal/claude-code`, refresh in
|
||||
`hal/claude-licences` — so following one token means reading three.
|
||||
|
||||
## 2. Two things are conflated in one repo
|
||||
|
||||
`noxflow` is 45 tables. Its original purpose — a HAL-native Jira — is 7 of them:
|
||||
|
||||
| domain | tables |
|
||||
|---|---:|
|
||||
| tasks (the original concept) | 7 |
|
||||
| agents / HR | 13 |
|
||||
| knowledge (a Confluence) | 8 |
|
||||
| meetings | 7 |
|
||||
| scheduling / ops | 2 |
|
||||
| other | 8 |
|
||||
|
||||
Agents alone are nearly double the concept the module was built for. This is the direct
|
||||
cause of concrete faults: per-agent Claude credentials had to be written by the noxflow
|
||||
runtime purely because `agents.claude_account` lives in noxflow's database, while node
|
||||
identity — the same kind of fact — lives in the mesh registry.
|
||||
|
||||
Separately, 91 of the 124 modules are third-party software (`plex`, `sonarr`,
|
||||
`firefox`, `postgres`). These **run on the mesh; they are not part of it**. They belong
|
||||
outside this repository. The pipeline and provisioning system are deliberately generic,
|
||||
so HAL's own modules dogfood exactly what external modules use — that property is what
|
||||
makes the separation safe.
|
||||
|
||||
## 3. Knowledge is split three ways
|
||||
|
||||
| store | owner | size | search |
|
||||
|---|---|---:|---|
|
||||
| `mesh_docs` (+ revisions) | `hal/hippocampus`, in the mesh DB | 144 docs, 73 touched in 30d | yes |
|
||||
| `knowledge_*` (8 tables) | `noxflow/runtime`, in the noxflow DB | 20 pages | yes |
|
||||
| repo markdown | git | `VISION.md`, `FLAVOR-PROPOSAL.md`, `docs/` | **no** |
|
||||
|
||||
The librarian is split across two of them: `librarian_file_document` writes via
|
||||
hippocampus, `librarian_ask` and `librarian_retrieve` read via noxflow. Filing and
|
||||
retrieval, different modules, different databases.
|
||||
|
||||
Both DB stores have revisions; only noxflow has a review workflow (`_proposals`,
|
||||
`_promotion_requests`). Neither has file access. Git has file access and history but no
|
||||
search integration. **Every piece needed already exists — none of them together.**
|
||||
|
||||
## 4. `hal/developer` and the missing local mesh
|
||||
|
||||
`hal/developer` (≈2,600 lines) exposes `dev_stage`, `dev_diff`, `dev_validate`,
|
||||
`dev_typecheck`, `dev_verify_flavor`, `dev_bootstrap`, `dev_build`, `dev_up`, `dev_down`,
|
||||
`dev_env`. It is the tooling that both reflects and enforces the module model, so it
|
||||
moves with any change to that model.
|
||||
|
||||
`dev_up` builds a dev environment **for a single module**: resolves `requires:`, starts
|
||||
providers "via systemctl (same as production)", provisions workspace-namespaced
|
||||
databases. It borrows the host node's services.
|
||||
|
||||
**There is no way to run a mesh locally.** Consequently anything that only manifests
|
||||
*between* nodes — credential rotation reaching a running session, provision credentials
|
||||
fanning out, cascade ordering, artifact packaging — is discoverable only in production.
|
||||
Every fault fixed on 2026-08-22 was of that kind.
|
||||
|
||||
A containerised mesh (N nodes, a broker, a registry, a coordinator) is therefore not a
|
||||
convenience. It is the precondition for a refactor of this size being verifiable at all.
|
||||
|
||||
## 5. `hal/sdk` is where domains go to hide
|
||||
|
||||
155 files, 34,636 lines. It contains code from essentially every bounded context:
|
||||
|
||||
| file | domain it belongs to |
|
||||
|---|---|
|
||||
| `installer-core.ts` (1,495) | runtime |
|
||||
| `tools/provisions.ts` (1,311) | provisioning |
|
||||
| `tools/deploy-service.ts` (1,081) | deploy |
|
||||
| `workflow-engine.ts` (970), `task-commands.ts` (588) | **work** — noxflow's domain, inside HAL's SDK |
|
||||
| `artifact-manager.ts` (948), `build-executor.ts` (760), `feature-handlers/*` | delivery |
|
||||
| `env-generator.ts` (944) | config |
|
||||
| `amqp-client.ts` (899) | comms |
|
||||
| `meshware.ts` (799) | runtime |
|
||||
| `module-registry.ts` (743) | mesh |
|
||||
| `claude-credentials.ts` | **ai** |
|
||||
|
||||
Nothing here was misplaced carelessly. Each landed in the SDK because the SDK is the one
|
||||
package every module may import — so putting a shared type there is the way to avoid a
|
||||
circular dependency between two modules that both need it.
|
||||
|
||||
The result is that the dependency graph is trivially acyclic and completely
|
||||
uninformative: everything depends on one 34k-line package, so a change anywhere in it
|
||||
rebuilds everything. On 2026-08-22 a three-line fix in `feature-handlers/migrations.ts`
|
||||
produced a 1,445-job, 26-minute mesh-wide cascade.
|
||||
|
||||
The last row is from that same day, and is the clearest illustration: per-agent Claude
|
||||
config-directory helpers were added to `@hal/sdk/claude-credentials.ts` because it was
|
||||
the only place both `hal/claude` and the noxflow runtime could import from. The
|
||||
alternative — `hal/ai` owning it and noxflow depending on `hal/ai` — was unavailable
|
||||
without answering the ownership question this effort exists to answer.
|
||||
|
||||
**Extraction is therefore blocked on the decomposition, not the reverse.** A domain
|
||||
module can only own its shared code once the domain has a module. The SDK should retain
|
||||
only what is genuinely cross-cutting: transport, manifest types, logging, process
|
||||
helpers.
|
||||
|
||||
## 6. The crux: two notions of "agent"
|
||||
|
||||
The original vision was **a node is an agent with thinking abilities**. Later, noxflow
|
||||
implemented actual agents — employees with skills, workspaces and tasks. Both survive,
|
||||
and they collide.
|
||||
|
||||
The collision is visible in the data. There are **two agent rows per node**:
|
||||
|
||||
| agent | display name | skills | bound to |
|
||||
|---|---|---|---|
|
||||
| one named after each node | the node's name | `{deploy,verify,operate}` | its node |
|
||||
| `hal-<node>`, one per node | "HAL (\<node\>)" | `{}` | its node |
|
||||
|
||||
One carries the **work**; the other carries only **identity** — no skills at all, existing
|
||||
to hold a licence for HAL's own sessions (`resolveHalModuleEnv`). The node-as-agent idea
|
||||
was implemented twice, half each.
|
||||
|
||||
The same split appears in the schema: `nodes.hal_claude_account` and
|
||||
`agents.claude_account` are the same fact in two tables in two databases. What was
|
||||
documented as "three licence touchpoints" (NODE, HAL, AGENT) is **one concept modelled
|
||||
three times**.
|
||||
|
||||
### Resolution
|
||||
|
||||
**Nodes and agents are decoupled.** The agents module owns agents; agents run on nodes.
|
||||
A node is a place where an agent can run — that is the whole of the relationship. There
|
||||
is no resident agent, no node-owned identity, no ownership in either direction.
|
||||
|
||||
Agents named after a node still exist, but they are **ordinary agents defined through the
|
||||
agents module** that happen to hold infra skills and be named after the node they usually run
|
||||
on. Nothing about them is structural.
|
||||
|
||||
Two consequences follow:
|
||||
|
||||
1. **`nodes.node_license` and `nodes.hal_claude_account` should not exist.** A node does
|
||||
not authenticate to a model provider — agents do. Both collapse into
|
||||
`agents.claude_account`, and what was documented as three licence touchpoints becomes
|
||||
one.
|
||||
2. **The pair of agent rows per node merges.** `hal-<node>` exists only to hold a licence
|
||||
for HAL's own sessions; once licences belong to agents and any module may employ an
|
||||
agent, it has no reason to be separate from `<node>`.
|
||||
|
||||
What remains unaccounted for is the operator's own interactive session:
|
||||
`~/.claude/.credentials.json` on a node serves a **human at a prompt**, not an agent.
|
||||
Under this decoupling that is not a node property either — it is the operator's
|
||||
credential on whichever machine they are sitting at. Whether the operator is modelled as
|
||||
a persona (soulstream treats humans and agents as peers with identical credentials) is an
|
||||
open question, and the last remaining reason `nodes` carries a licence column today.
|
||||
|
||||
## 7. The record, not "meetings"
|
||||
|
||||
soulstream models work as **one signed log** in which humans and agents apply identical
|
||||
changes, with work arriving by **mention**. Memory and decisions live in that single
|
||||
auditable stream — no hidden state between components. Its stated economics,
|
||||
*"cheap awareness and expensive thinking"*, is the same principle as the budget guard:
|
||||
constant monitoring, selective deep processing.
|
||||
|
||||
Under that model "meetings" is not a domain. A meeting is a **thread in the record**.
|
||||
Which means six of today's modules and table groups are one thing seen from different
|
||||
angles:
|
||||
|
||||
`hal/axon` (DMs, rooms) · `hal/cortex` · `hal/synapse` · `hal/notifications` ·
|
||||
`conversations` · `meetings`
|
||||
|
||||
A notification, in particular, is just a mention not yet read — which is why it needed
|
||||
its own table, its own id column, and its own failure mode.
|
||||
|
||||
## 8. Ideal state
|
||||
|
||||
The mesh repository contains only the mesh, decomposed by domain rather than by
|
||||
deployment accident:
|
||||
|
||||
| context | what it is | absorbs today's |
|
||||
|---|---|---|
|
||||
| `hal/infra` | **the room** — nodes, registry, provisioning, pipeline, node runtime | `hal/mesh`, `hal/brain`, `hal/coordinator`, `hal/meshware`, `hal/developer`, `hal/bootstrap` |
|
||||
| `hal/agents` | **the name and the thinking** — Agent owns identity, licence, runs, memory, thoughts | noxflow agents (13 tables), `hal/thoughts` |
|
||||
| `hal/stream` | **the record** — mentions, threads, meetings, notifications | `hal/axon`, `hal/cortex`, `hal/synapse`, `hal/notifications`, meetings (7), conversations |
|
||||
| `hal/knowledge` | documents, spaces, revisions, review, search | `hal/hippocampus` + noxflow `knowledge_*` |
|
||||
| `hal/work` | tasks, workflows, boards — the original noxflow | noxflow tasks (7 tables) |
|
||||
| `hal/ai` | provider integration, flavored per vendor | `hal/claude*`, `claude-code` |
|
||||
| `hal/config` | env and config distribution, secrets, PKI | `hal/env-sync`, `hal/config-sync`, `hal/secrets`, `mesh-ca` |
|
||||
| `hal/observability` | health, dashboards, logs | `hal/health`, `hal/meshboard` |
|
||||
|
||||
Eight contexts instead of 33 platform modules, with the catalogue elsewhere.
|
||||
|
||||
The naming corrects itself in the process: what is called "brain" today is
|
||||
infrastructure, and the word survives only as a concept an Agent owns — never again as a
|
||||
module name.
|
||||
|
||||
This shape depends on **named features with per-node opt-in** — without it, every
|
||||
independently-deployable unit inside a context becomes a module again and the count
|
||||
returns. That mechanism is the subject of a separate ADR.
|
||||
|
||||
## Decided since
|
||||
|
||||
| question | decision |
|
||||
|---|---|
|
||||
| Meetings — own context or part of agents? | **Neither.** A meeting is a thread in the record; `hal/stream` absorbs it. The category was wrong. |
|
||||
| Thoughts — autonomy or agent behaviour? | `hal/agents`. Any agent can have a thought. |
|
||||
| `hal/axon` vs `hal/cortex` | `hal/cortex` is vestigial — no `daemon/` source, unit inactive on every node, description copy-pasted from axon. Messaging belongs to `hal/stream`. |
|
||||
| `hal/hypothalamus` | Orphaned build output. A `dist/` directory with no `module.yml`, so not a module at all. |
|
||||
| Who owns agent identity? | `hal/agents`. Nodes and agents are decoupled; a node is only a place an agent can run. |
|
||||
| Where do third-party modules live? | Outside this repository. They run *on* the mesh, not *of* it. |
|
||||
|
||||
## Open questions
|
||||
|
||||
1. **`hal/scheduler`** — infrastructure (fire an event later, belongs to `hal/mesh`) or
|
||||
part of `hal/work`?
|
||||
2. **The executor.** Pulling and routing work items is `hal/work`; spawning a session is
|
||||
`hal/agents`; providing the place to run is `hal/mesh`. It currently straddles all
|
||||
three, which is the same kind of straddle that put credential-writing in the noxflow
|
||||
runtime.
|
||||
3. **Catalogue destination** — one repository, or per-application repositories, given the
|
||||
pipeline resolves dependencies across the registry rather than the filesystem?
|
||||
4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or
|
||||
does that belong to `hal/stream`? Everything imports it, which argues both ways.
|
||||
5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which
|
||||
user, on which node, a human agent acts as. Where does it live — an attribute of the
|
||||
agent, or of the agent-node binding?
|
||||
@@ -0,0 +1,46 @@
|
||||
# 001 — Module domain decomposition
|
||||
|
||||
- **Status:** ONGOING — ADR 0001 accepted; graduates when `02-DESIGN` carries the per-context specifications
|
||||
- **Initiated by:** jochen, 2026-08-22
|
||||
- **Areas touched:** every `hal/*` and `noxflow/*` module; the pipeline's dependency
|
||||
graph; the knowledge base; agent identity and credentials.
|
||||
|
||||
## Summary
|
||||
|
||||
HAL has 124 modules. That number is not a maintenance problem in itself — it is the
|
||||
**symptom of missing bounded contexts**. Modules are split not because they model
|
||||
different domains, but because splitting is the only lever the platform offers:
|
||||
|
||||
- no way to run one daemon on one node without making it a module
|
||||
(`hal/claude-licences` — one daemon, single-node)
|
||||
- no way to expose two of a kind from one module
|
||||
- no namespace separating the mesh from the software it runs
|
||||
|
||||
This effort establishes the **current state**, the **ideal state**, and the sequence
|
||||
between them.
|
||||
|
||||
## Trigger
|
||||
|
||||
A night of debugging that produced four fixes and one conclusion. Every fault was a
|
||||
boundary fault:
|
||||
|
||||
- Per-agent Claude credentials had to be written by the *noxflow runtime*, because
|
||||
`agents.claude_account` is in noxflow's database — even though agent identity is a
|
||||
mesh concept and node identity already lives in the mesh registry.
|
||||
- Whether `hal/brain` may depend on noxflow took three attempts to answer, twice
|
||||
wrongly, because the ownership boundary was never stated.
|
||||
- Authoritative documentation existed in `mesh_docs` and was not found, while a
|
||||
proposal in repo markdown was invisible to search entirely.
|
||||
|
||||
## Decisions taken (2026-08-22)
|
||||
|
||||
| Question | Decision |
|
||||
|---|---|
|
||||
| What should noxflow become? | Decompose into `hal/*` modules; noxflow returns to tasks/workflows |
|
||||
| Who owns agent identity? | `hal/agents` — a mesh concept, alongside nodes |
|
||||
| Where do third-party apps live? | Out of this repo. They run *on* the mesh; they are not *of* it |
|
||||
| Knowledge structure | Modelled on `papa-hq`; implementation choice left open |
|
||||
|
||||
## Open questions
|
||||
|
||||
Tracked in [`analysis.md`](analysis.md) under "Open questions".
|
||||
@@ -0,0 +1,258 @@
|
||||
# A mesh that runs locally — current state and obstacles
|
||||
|
||||
Evidence for Phase 0. Every claim here is either a file location or something measured on
|
||||
2026-08-22; where a claim was checked and found false, that is recorded too.
|
||||
|
||||
---
|
||||
|
||||
## 1. What exists today, and why neither is a mesh
|
||||
|
||||
### `test/pipeline/` — the only containerised HAL node, and it is dead
|
||||
|
||||
This harness builds `@hal/sdk` and the `hal/brain` daemon into a `node:23-alpine` image,
|
||||
compiles `hal/coordinator`'s tools against it, and runs `brainstem.js` with
|
||||
`HAL_MODE=daemon` against a containerised postgres and LavinMQ
|
||||
(`test/pipeline/Dockerfile`, `test/pipeline/docker-compose.yml`). It proves the valuable
|
||||
thing: **a node runtime needs no systemd, no `/services/`, and no nvm** — environment
|
||||
variables and a seeded `nodes` row are sufficient.
|
||||
|
||||
It also cannot build. Measured 2026-08-22:
|
||||
|
||||
```
|
||||
Step 7/16 : RUN npm run build -w modules/hal/sdk
|
||||
npm error No workspaces found:
|
||||
npm error --workspace=modules/hal/sdk
|
||||
```
|
||||
|
||||
The npm workspace it depends on was removed on 2026-06-04 (`21ef4a4e`, "kill npm
|
||||
workspace"), and the root `package.json` now carries a comment explaining why it will not
|
||||
come back. The harness has not been touched since before that commit.
|
||||
|
||||
So the repository's only end-to-end pipeline test has been unrunnable for two and a half
|
||||
months and nothing reported it — which is the same shape as the faults Phase 0 exists to
|
||||
catch. **A test nobody runs is indistinguishable from a test that passes.**
|
||||
|
||||
### `test/dev-mesh/` — a work plane, not a control plane
|
||||
|
||||
Four `runtime-*` services, a noxflow API, a meshboard and a shared postgres + LavinMQ on
|
||||
one network (`test/dev-mesh/docker-compose.yml`). It is genuinely useful and the topology
|
||||
is a good precedent: one network, per-service `HAL_NODE`, `EXECUTOR_STUB=1` so dispatch and
|
||||
run lifecycle are real while the model call is faked.
|
||||
|
||||
It is the wrong layer for Phase 0, on three counts:
|
||||
|
||||
| | dev-mesh does | Phase 0 needs |
|
||||
|---|---|---|
|
||||
| state | restores `pg_dump`s (`test/dev-mesh/init.sh`) | schema built by **migrations**, or fixture 3 is untestable by construction |
|
||||
| code | bind-mounts pre-built `dist/` from the host | artifacts **built and delivered** by the pipeline |
|
||||
| scope | noxflow runtime, API, meshboard | `hal/coordinator`, `hal/meshware`, MinIO, provisioning |
|
||||
|
||||
No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises
|
||||
what the mesh *runs*; Phase 0 needs what the mesh *is*.
|
||||
|
||||
### `dev_up` — confirmed host-coupled
|
||||
|
||||
The claim in the work breakdown holds. `startDevProvider()` starts providers through the
|
||||
host's user systemd instance and reads credentials from a host path
|
||||
(`modules/hal/developer/tools/dev-env.ts:146-178`):
|
||||
|
||||
```ts
|
||||
const unitName = `hal-module@${provider}.service`;
|
||||
execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" });
|
||||
const creds = resolveRunningProviderCreds(provider); // reads /services/<provider>/.env
|
||||
```
|
||||
|
||||
It borrows the host. There is no mesh to stand up.
|
||||
|
||||
---
|
||||
|
||||
## 2. What is already portable
|
||||
|
||||
- **Node identity is one environment variable.** `resolveNodeName()` is
|
||||
`process.env.HAL_NODE || hostname()`, lowercased (`modules/hal/sdk/src/mesh-config.ts:215`).
|
||||
No file, no registration handshake, no host coupling.
|
||||
- **Registration is a plain upsert.** `mesh_node_register` inserts into `nodes`
|
||||
(`modules/hal/mesh/tools/index.ts:490-540`); mandatory fields are `name` and `user_name`.
|
||||
- **The daemons are AMQP loops.** `hal/meshware` consumes `cmd.feature.{build,install,configure,start}`
|
||||
(`modules/hal/meshware/daemon/src/cerebellum.ts:780-857`); `hal/coordinator` consumes
|
||||
`event.gitea.push` and drives the cascade. Both are configured entirely by environment.
|
||||
- **`hal/brain` in daemon mode already runs in a plain node image**, per `test/pipeline/`.
|
||||
|
||||
---
|
||||
|
||||
## 3. The obstacles, located
|
||||
|
||||
Four couplings stand between the daemons and a container. All are in code, none are
|
||||
mysterious.
|
||||
|
||||
| # | coupling | location |
|
||||
|---|---|---|
|
||||
| 1 | meshware restarts itself via the host: `execFileSync("systemctl", ["--user","restart","hal-meshware.service"])` | `modules/hal/meshware/daemon/src/cerebellum.ts:826` |
|
||||
| 2 | meshware restarts the broker via a hardcoded host path: `execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"])` | `cerebellum.ts:843-844` |
|
||||
| 3 | **a module service *is* a systemd unit** — `hal-module@.service` runs `docker compose --project-directory /services/%i` | `modules/hal/meshware/systemd/hal-module@.service:10-11` |
|
||||
| 4 | env generation writes `homedir()`-relative host paths — `~/.config/hal/env`, `~/.config/hal/modules/*.env`, `/services/*/.env` | `modules/hal/sdk/src/env-generator.ts:188-223`, `:538-568`, `:584-606` |
|
||||
|
||||
1, 2 and 4 are small: a guard and a configurable root. **3 is the architectural one**, and
|
||||
it is the first open question below — in a container, "systemd unit that runs docker
|
||||
compose" has no natural translation.
|
||||
|
||||
The bootstrap scripts add their own host assumptions — nvm for node resolution, `sudo
|
||||
mkdir -p /services && chown`, and a `~/dotfiles` repository that is not in this repo
|
||||
(`install.d/adopt.sh:165-170`) — but Phase 0 does not have to run them. A container image
|
||||
can be built the way `test/pipeline/Dockerfile` builds one, bypassing the bootstrap path
|
||||
entirely. That is a deliberate divergence to record: **the local mesh would not test the
|
||||
bootstrap**, only the running mesh.
|
||||
|
||||
---
|
||||
|
||||
## 4. External services, and whether they can be local
|
||||
|
||||
| service | used for | local substitute |
|
||||
|---|---|---|
|
||||
| PostgreSQL | registry + pipeline state | yes — already the pattern in both harnesses |
|
||||
| LavinMQ | all coordinator↔meshware messaging | yes — `cloudamqp/lavinmq`, already used |
|
||||
| MinIO | per-feature artifact tarballs, `modules/{name}/{version}/{feature}.tar.gz` (`modules/hal/sdk/src/artifact-manager.ts:21,65`) | yes — `minio/minio`, needs only `REGISTRY_MINIO_*` |
|
||||
| Gitea | source, push webhook, **and** the `@hal/*` npm registry | container exists, but heavy; see open question 2 |
|
||||
| Docker registry | `docker push` for modules declaring `docker:` (`build-executor.ts:195-231`) | `registry:2`, or avoid modules that need it |
|
||||
| Traefik | writes routing config at install; does not gate success | omit |
|
||||
|
||||
Only Gitea is awkward, and only because it carries two roles at once.
|
||||
|
||||
---
|
||||
|
||||
## 5. The fixtures
|
||||
|
||||
Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they
|
||||
differ sharply in cost.
|
||||
|
||||
### Fixture B — a provider deploy rotates a shared credential without fanning out
|
||||
|
||||
**Cheapest, best understood, and the root cause is still open.** Documented at
|
||||
`troubleshooting/provision-adoption-rotates-live-credential`. The mechanism is two lines:
|
||||
|
||||
```ts
|
||||
async provision(project, username, options) {
|
||||
const password = generatePassword(); // ALWAYS a fresh password
|
||||
```
|
||||
|
||||
and, in the adoption branch of `provisionDatabase`, an `ALTER ROLE ... WITH PASSWORD` that
|
||||
rotates the live secret while updating only the provider's own `mesh_provisions` row
|
||||
(`modules/postgres/tools/index.ts`). Every consumer sharing that role keeps a stale
|
||||
password and fails permanently; the provider recovers alone, and that asymmetry is the
|
||||
tell.
|
||||
|
||||
Reproduction needs one provider node and two consumer nodes — the minimum interesting
|
||||
mesh. It re-fires on every provider deploy, so it does not need to be provoked, only
|
||||
observed. Confirmation is a timestamp comparison, not a hash: the stored values are
|
||||
`enc:v1:` with a random IV, so identical plaintexts hash differently.
|
||||
|
||||
Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors
|
||||
each; the provider failed for 22 minutes and recovered by itself.
|
||||
|
||||
### Fixture C — a migration ships nothing while the pipeline reports success
|
||||
|
||||
Three independent silent-skip points, any one of which produces it:
|
||||
|
||||
1. **Not-applicable and succeeded are the same status.** `MigrationsHandler.detect()` is
|
||||
`existsSync(join(moduleDir, "migrations"))` against the freshly cloned workspace
|
||||
(`modules/hal/sdk/src/feature-handlers/migrations.ts:32-41`). If the directory is not
|
||||
there — uncommitted, ignored, or a `module_path` that does not line up — the build
|
||||
reports `success` (`cerebellum.ts:634-638`).
|
||||
2. **Upload failure is a warning.** `uploadFeatureArtifact()` returns `{sha256: ""}` with a
|
||||
`warn` when none of the listed files exist (`artifact-manager.ts:41-44`), and callers
|
||||
catch and warn rather than throw (`cerebellum.ts:672-673`). The symmetric download
|
||||
failure on the target node is also non-fatal (`artifact-manager.ts:93-98`,
|
||||
`cerebellum.ts:426-431`, commented "Non-fatal: feature may work without artifact").
|
||||
3. **Missing packaged files are skipped in silence.** `addFileEntries()` does
|
||||
`if (!existsSync(abs)) continue;` (`module-builder-core.ts:134-135`) — an explicit
|
||||
`package:` entry that is not on disk simply is not in the tarball.
|
||||
|
||||
This is the same class as `troubleshooting/flavors-never-packaged` (`flavors/` was never
|
||||
staged, so a node kept its first copy forever) and is catalogued in
|
||||
`troubleshooting/deploy-reports-transport-not-effect`, which counted **0 of 119 modules
|
||||
implementing `verify`** — the stage that exists, is dispatched, has a working handler, and
|
||||
would turn every one of these into a red pipeline.
|
||||
|
||||
### Fixture A — a credential rotation does not reach a running session
|
||||
|
||||
The most valuable and the most work: it needs a *running session* to rotate underneath,
|
||||
which means the local mesh must be able to start one. The mechanism is understood — an
|
||||
`EnvironmentFile` is read once and `process.env` is a snapshot — and it is why credentials
|
||||
became files that get re-read. Defer it behind B and C.
|
||||
|
||||
**Recommendation:** B first. It needs three nodes and a provision, no build and no session,
|
||||
and it is the one whose root cause is still open — so reproducing it locally has value
|
||||
beyond proving the harness.
|
||||
|
||||
---
|
||||
|
||||
## 6. A claim that was checked and found false
|
||||
|
||||
The survey behind this document suspected that the live AMQP pipeline never writes
|
||||
`deployments` or `node_modules.installed_version`, because the code that does so sits in
|
||||
`installer-core.ts:1045-1126` on what is commented as the "CLI path".
|
||||
|
||||
Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours:
|
||||
|
||||
| node | deploys | newest |
|
||||
|---|---|---|
|
||||
| 1 | 24 | 17:26:17 |
|
||||
| 2 | 27 | 17:22:22 |
|
||||
| 3 | 47 | 17:25:54 |
|
||||
| 4 | 26 | 17:22:23 |
|
||||
|
||||
The record is live on all four nodes and minutes old. **The claim is refuted.** Recorded
|
||||
because the reasoning was plausible and someone will retrace it.
|
||||
|
||||
---
|
||||
|
||||
## 7. Questions
|
||||
|
||||
### Settled (jochen, 2026-08-22)
|
||||
|
||||
**Phase 0 is a development environment, not a fixture rig.** It is built as a supported
|
||||
surface to work in daily, and it is what `dev_up` should become. The definition of done —
|
||||
"demonstrated in the local mesh" — is therefore meant literally.
|
||||
|
||||
**The trigger is a real Gitea container.** The webhook relay is part of what is under test,
|
||||
including the changed-file enrichment that exists because the webhook truncates at 20
|
||||
commits (`modules/hal/gitea/tools/index.ts:136-195`). A synthetic `event.gitea.push` would
|
||||
skip it. Gitea also carries the `@hal/*` npm registry, so the local mesh needs it twice
|
||||
over either way.
|
||||
|
||||
### Open — blocking
|
||||
|
||||
1. **How does a node supervise a module service?** Today a module service is a systemd unit
|
||||
running `docker compose` against `/services/<module>`
|
||||
(`modules/hal/meshware/systemd/hal-module@.service:10-11`), which has no direct
|
||||
translation inside a container. The question raised in response is the better one, and
|
||||
is broader than Phase 0: **what would it cost to stop using systemd altogether and have
|
||||
the mesh supervise its own services?** That is a mesh-level architectural question, not
|
||||
a local-mesh implementation detail — if the answer is that the mesh should own
|
||||
supervision, Phase 0 should not build a container-only workaround first. Under
|
||||
investigation; findings will land in [`003-service-supervision`](../003-service-supervision/)
|
||||
and, if it goes ahead, ADR 0002.
|
||||
|
||||
2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable
|
||||
root keeps one code path; container-specific targets keep the host paths untouched. This
|
||||
decides whether env generation gets a seam or a conditional.
|
||||
|
||||
3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap
|
||||
scripts, and it will run one OS where the real mesh is heterogeneous by design
|
||||
(`00-GENESIS/context.md`). Stating the divergence up front is what stops "it works
|
||||
locally" from becoming its own class of silent failure. Now sharper, because a
|
||||
development environment people use daily is trusted far more than a rig — and drifting
|
||||
from production costs correspondingly more.
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this
|
||||
phase unblocks
|
||||
- [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) — Phase 0 tasks
|
||||
and checkpoint
|
||||
- `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open
|
||||
- `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped
|
||||
green, and the unused `verify` stage
|
||||
- `troubleshooting/flavors-never-packaged` — fixture C, previously seen in the wild
|
||||
@@ -0,0 +1,53 @@
|
||||
# 002 — A mesh that runs locally
|
||||
|
||||
- **Status:** GRADUATED — the design is [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md)
|
||||
- **Initiated by:** jochen, 2026-08-22
|
||||
- **Areas touched:** `install.d/`, `hal/meshware`, `hal/coordinator`, `hal/brain`,
|
||||
`hal/developer` (`dev_up`), `hal/sdk` (env generation, feature handlers, artifact
|
||||
manager), `test/pipeline/`, `test/dev-mesh/`, the provisioning path in
|
||||
`modules/postgres/`.
|
||||
|
||||
## Summary
|
||||
|
||||
Phase 0 of [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) requires
|
||||
a mesh that comes up in containers, runs its own pipeline, and reproduces known faults on
|
||||
demand. Nothing else in the decomposition starts until it exists, because every fault the
|
||||
decomposition addresses was found in production — there was nowhere else to find it.
|
||||
|
||||
This effort establishes what already runs in a container, what is welded to the host, and
|
||||
what it would take to close the gap. It does **not** choose an approach: the central
|
||||
question — how a containerised node executes a module service, when a module service is
|
||||
defined today as a systemd unit shelling to `docker compose` in `/services/` — is not
|
||||
answered by ADR 0001 and is recorded below rather than decided.
|
||||
|
||||
## What was established
|
||||
|
||||
- The two existing container harnesses are **neither of them a mesh**, and one of them has
|
||||
not been able to build since 2026-06-04.
|
||||
- Node identity is already portable — a single environment variable, no host handshake.
|
||||
- Four concrete host couplings block a containerised node, all with known locations.
|
||||
- All three Phase 0 fixtures are reproducible; one of them is documented in the knowledge
|
||||
base with an open root cause and is the cheapest place to start.
|
||||
|
||||
Detail and evidence in [`analysis.md`](analysis.md).
|
||||
|
||||
## Questions — all settled 2026-08-22
|
||||
|
||||
**Phase 0 is a development environment**, not a fixture rig — it is what the host-borrowing
|
||||
dev tooling becomes.
|
||||
|
||||
**The trigger is a real source-forge container**, because the webhook relay is part of what
|
||||
is under test.
|
||||
|
||||
**A node is a system container**, promotable to a virtual machine per node. This answered
|
||||
the question the effort was stuck on, and dissolved it rather than solving it: against a
|
||||
real node with a real init, the four host couplings catalogued in `analysis.md` §3 are not
|
||||
couplings — they are how a node works. They were obstacles only to a node modelled as an
|
||||
application container.
|
||||
|
||||
Consequently [`003-service-supervision`](../003-service-supervision/) **no longer blocks
|
||||
Phase 0**. It remains a live architecture question, on its own timeline.
|
||||
|
||||
The remaining questions in `analysis.md` — what replaces host paths in a container, and how
|
||||
faithful the lab must be — are answered in the design: nothing replaces them, because the
|
||||
paths are real; and the divergences are enumerated rather than discovered.
|
||||
@@ -0,0 +1,225 @@
|
||||
# Who supervises a service — the cost of leaving systemd
|
||||
|
||||
Measured 2026-08-22. Every claim is a file location or a count.
|
||||
|
||||
---
|
||||
|
||||
## 1. What systemd actually does for HAL
|
||||
|
||||
**14 modules** ship a `systemd/` directory. The units divide cleanly into three kinds:
|
||||
|
||||
| kind | count | what it is |
|
||||
|---|---|---|
|
||||
| long-running daemons | 9 | `hal/brain`, `hal/meshware`, `hal/coordinator`, `hal/cortex`, `hal/env-sync`, `hal/file-share`, `hal/thoughts`, `noxflow/runtime`, `desktop_notifications` — all `Restart=always`, `RestartSec=10` |
|
||||
| scheduled one-shots | 6 timers | docker-prune, docker-registry maintenance, claude-code sessions + usage, health, mailu cert-sync |
|
||||
| the template | 1 | `hal-module@.service` — the per-module Docker lifecycle |
|
||||
|
||||
Only `mailu` is system-scope. Everything else is a user unit under
|
||||
`~/.config/systemd/user`, auto-detected from directory presence — the module author names
|
||||
the file and that filename *is* the unit name
|
||||
(`modules/hal/sdk/src/feature-handlers/systemd.ts:31-42`).
|
||||
|
||||
The template is the piece `002` tripped over
|
||||
(`modules/hal/meshware/systemd/hal-module@.service`):
|
||||
|
||||
```
|
||||
ExecStart=docker compose --project-directory /services/%i --project-name %i up --remove-orphans --pull always
|
||||
ExecStop=docker compose --project-directory /services/%i --project-name %i down
|
||||
Restart=on-failure
|
||||
StartLimitIntervalSec=120
|
||||
StartLimitBurst=5
|
||||
```
|
||||
|
||||
Systemd's roles, then: restart-on-failure, start at login, env-file loading, ordering,
|
||||
per-module lifecycle, timers, and — via journald — **the only log store the 9 Node daemons
|
||||
have** (`modules/hal/sdk/src/tools/log-tail.ts:50-53`).
|
||||
|
||||
---
|
||||
|
||||
## 2. The question splits in two, and the halves disagree
|
||||
|
||||
### For `docker compose` stacks, systemd is mostly redundant
|
||||
|
||||
**44 of 44** module `docker-compose.yml` files declare a container-level `restart:` policy —
|
||||
`always` or `unless-stopped`, none missing. Docker's own daemon already restarts crashed
|
||||
containers, independently of systemd.
|
||||
|
||||
`hal-module@.service` does not supervise the containers. It supervises the **`docker compose
|
||||
up` foreground process**. It is a second layer on top of a restart policy that already
|
||||
works. What it genuinely adds is narrower than it looks:
|
||||
|
||||
- one uniform verb for every module type — `systemctl --user start hal-module@X` rather than
|
||||
remembering each module's compose invocation, relied on across a dozen call sites in
|
||||
`meshware.ts:124-154`, `dev-env.ts:142-163`, `installer-core.ts`, `infra.ts:166`
|
||||
- recovery when the `docker compose up --pull always` process itself dies — a failed image
|
||||
pull, not a crashed container
|
||||
- rate-limited restart (`StartLimitBurst=5`) so a broken stack does not spin
|
||||
|
||||
That is real, but it is a convenience layer, not a safety layer. **This half could go at
|
||||
moderate cost.**
|
||||
|
||||
### For the 9 Node daemons, systemd is load-bearing
|
||||
|
||||
There is no alternative supervisor anywhere in the repo. No PM2, no forever, no nodemon, no
|
||||
watchdog loop — all checked, zero hits. `Restart=always` is the only thing standing between
|
||||
a crashed daemon and a dead node.
|
||||
|
||||
**This half is the actual question.**
|
||||
|
||||
---
|
||||
|
||||
## 3. The hard part is fate-sharing
|
||||
|
||||
The difficulty is not systemd. It is that **a supervisor must not share fate with what it
|
||||
supervises**, and the codebase already has a scar from exactly this.
|
||||
|
||||
meshware cannot restart itself mid-request: killing the process before it ACKs the AMQP
|
||||
message loses the message. So it defers its own restart by two seconds after closing the
|
||||
connection (`modules/hal/meshware/daemon/src/cerebellum.ts:815-828`) — a commented
|
||||
workaround for a problem that only exists because the thing being restarted is the thing
|
||||
doing the restarting.
|
||||
|
||||
Any mesh-native supervisor inherits this recursively. Something has to be the outermost
|
||||
always-alive process, and if it is written by the mesh, the mesh must supervise it, and so
|
||||
on. The recursion only terminates at a process the mesh does not own. Today that is
|
||||
systemd. **A "more mesh" supervisor that is itself a mesh process is not a smaller problem;
|
||||
it is the same problem with a new name.**
|
||||
|
||||
This is the strongest argument for the status quo, and it is worth stating plainly before
|
||||
looking at alternatives.
|
||||
|
||||
---
|
||||
|
||||
## 4. The option neither of us named
|
||||
|
||||
There is a third answer that terminates the recursion in something that is not systemd and
|
||||
not written by us: **run HAL's own daemons as containers.**
|
||||
|
||||
Docker is already the outermost supervisor for 44 of 44 module stacks. It does not share
|
||||
fate with the mesh. It already has restart policies, backoff, and a log store that
|
||||
`log_tail` already speaks (`log-tail.ts:43-47` reads `docker logs` for Docker modules
|
||||
today). Extending it from "the things the mesh runs" to "the mesh itself" is not new
|
||||
machinery — it is applying machinery the mesh already trusts to one more case.
|
||||
|
||||
What that buys, beyond supervision:
|
||||
|
||||
- **Phase 0 stops being a translation.** A containerised node becomes the same shape as a
|
||||
production node, rather than a local approximation with a systemd-shaped hole in it. The
|
||||
divergence that `002` open question 3 worries about largely disappears.
|
||||
- **`/services/` and `~/.config/hal/` stop being special.** Mounts, not host paths.
|
||||
- **The GENESIS "dogfood everything" value gets easier**, not harder: the mesh's own
|
||||
components would ship and run exactly like everything else it carries.
|
||||
|
||||
What it costs, honestly:
|
||||
|
||||
- **journald → docker logs** for the 9 daemons. `log_tail` already handles both, but
|
||||
`systemd_journal` and the health checks that read unit state
|
||||
(`modules/hal/mesh/health.sh:34-40`, `modules/hal/health/hal-health.sh:335-395`) would
|
||||
need a container-aware path.
|
||||
- **Boot start** becomes Docker's `restart: always` plus the Docker daemon being enabled at
|
||||
boot — which is still one systemd unit, but the OS's own, not ours.
|
||||
- **A container needs the host to be reachable** for anything that touches the node itself.
|
||||
Some of these daemons exist precisely to write host files.
|
||||
|
||||
---
|
||||
|
||||
## 5. Why it cannot be all-or-nothing — and ADR 0001 already says so
|
||||
|
||||
Some of what runs under systemd today **cannot** be containerised, and the reason is
|
||||
already in the domain model. ADR 0001:
|
||||
|
||||
> a non-human agent acts through a spawned session — a human agent acts through a shell or
|
||||
> desktop
|
||||
|
||||
`hal/brain` has two modes (`modules/hal/brain/daemon/src/brainstem.ts:7-13`): a daemon mode
|
||||
that is an AMQP relay, and a **cortex mode that is an MCP server over stdio for an
|
||||
interactive Claude Code session**. The second is a human agent's modality. It runs in the
|
||||
human's shell, on the human's node, against the human's `~/.claude`. Containerising it is
|
||||
not a hard engineering problem, it is a category error.
|
||||
|
||||
The same holds for `desktop_notifications` and everything `hal/desktop-environment` touches.
|
||||
|
||||
So the line is not "systemd or not". It is:
|
||||
|
||||
| | belongs where |
|
||||
|---|---|
|
||||
| mesh daemons — meshware, coordinator, env-sync, thoughts, file-share, brain **in daemon mode** | supervisable by Docker; candidates to containerise |
|
||||
| human-modality surfaces — brain **in cortex mode**, desktop notifications, desktop environment | on the host, by definition |
|
||||
| module stacks | already Docker; systemd layer is the redundant part |
|
||||
|
||||
This split is not a compromise between the options. It is what the domain model implies,
|
||||
and it is a decent sign that the model is doing work.
|
||||
|
||||
---
|
||||
|
||||
## 6. Options, with costs
|
||||
|
||||
| | option | cost | what it buys |
|
||||
|---|---|---|---|
|
||||
| **A** | **Keep systemd; systemd-in-container for Phase 0** | privileged containers, heavy images, slow iteration; local mesh keeps a shape production does not have | nothing changes in production; smallest change to the mesh |
|
||||
| **B** | **Drop the `hal-module@` layer only** — let Docker's restart policies supervise stacks directly | reimplement uniform start/stop across ~12 call sites; lose rate-limited restart and pull-failure recovery | removes the redundant layer; does **not** solve Phase 0 on its own, since the daemons still need supervising |
|
||||
| **C** | **Containerise the mesh daemons; Docker supervises** | container-aware `systemd_journal`/health; host access for daemons that write host files; the human-modality surfaces stay on the host regardless | terminates the fate-sharing recursion without writing a supervisor; makes local and production the same shape; Phase 0 becomes much less of a special case |
|
||||
| **D** | **Write a mesh-native supervisor** | the fate-sharing recursion (§3), plus matching systemd's maturity — backoff, resource limits, clean SIGTERM (`noxflow/runtime` already depends on `TimeoutStopSec=60`) — on machines that are somebody's daily driver | most "mesh"; least justified by the evidence |
|
||||
|
||||
**D is the option the phrasing "more hal mesh approach" points at, and the evidence argues
|
||||
against it.** Supervision is not a domain concern the mesh is better placed to solve than
|
||||
the OS; the mesh's distinguishing feature is brokering capabilities, not restarting
|
||||
processes. C gets the benefit D is reaching for — the mesh not depending on host-specific
|
||||
init — without the recursion.
|
||||
|
||||
**C and B compose.** C is the one that pays for Phase 0.
|
||||
|
||||
Worth noting what is *not* in this table: the dozens of `systemctl` call sites that
|
||||
configure the **host OS's own** units — NetworkManager, resolved, oomd, zram, docker.service,
|
||||
fail2ban, sshd, zfs, across `modules/asusd`, `g14-power`, `wireguard`, `sshd`, `zfs` and
|
||||
others. Those are not HAL supervising itself; they are HAL configuring an Arch box. They
|
||||
are out of scope for every option above and do not go away under any of them.
|
||||
|
||||
---
|
||||
|
||||
## 7. An incidental finding
|
||||
|
||||
Documentation describes an automatic node rescue: `hal-rescue.sh:20` states it is "triggered
|
||||
automatically by `hal-health.timer` when hal-meshware is failed".
|
||||
|
||||
**It is not.** `hal-health.sh` contains no call to `hal-rescue.sh` (checked, zero matches),
|
||||
and **no unit in the repository declares `OnFailure=`** (checked, zero matches). The only
|
||||
real triggers are the manual `rescue_node` tool
|
||||
(`modules/hal/sdk/src/tools/deploy-rescue.ts:77-92`) and running `install.d/rescue.sh` by
|
||||
hand.
|
||||
|
||||
This is worth recording for two reasons. It weakens any argument that systemd-adjacent
|
||||
self-healing is already wired — it is not. And it is another instance of the pattern this
|
||||
whole refactor is about: **a documented mechanism that does not exist, believed because it
|
||||
was written down.** `00-GENESIS/how-we-build.md` calls this out as a rule; here it is again,
|
||||
found by grep.
|
||||
|
||||
---
|
||||
|
||||
## 8. Open question
|
||||
|
||||
**Which supervision model does the mesh adopt?** A, B, C, D or a combination.
|
||||
|
||||
**This no longer gates Phase 0.** When this was written, the local mesh was assumed to be
|
||||
built from application containers, which forced the question — there is no natural way to
|
||||
run an init system inside one. The decision of 2026-08-22 to build development nodes as
|
||||
**system containers** (see [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md))
|
||||
removes that pressure entirely: a system container runs a real init, so the existing model
|
||||
works unmodified and the lab needs no answer here to exist.
|
||||
|
||||
What remains is the question on its own merits, which is worth keeping open because the
|
||||
evidence above still holds: option B removes a layer that 44 of 44 module stacks have already
|
||||
made redundant, and option C would let the mesh stop depending on host-specific init. Neither
|
||||
is urgent. Both are now cheap to *try*, because there is somewhere to try them.
|
||||
|
||||
Option D — a mesh-written supervisor — remains the one the evidence argues against, for the
|
||||
fate-sharing reason in §3.
|
||||
|
||||
## References
|
||||
|
||||
- [`002-local-mesh`](../002-local-mesh/analysis.md) — the effort this came out of
|
||||
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — agent modality, which
|
||||
decides what cannot leave the host
|
||||
- `modules/hal/meshware/daemon/src/cerebellum.ts:815-828` — the self-restart workaround
|
||||
- `modules/hal/meshware/systemd/hal-module@.service` — the per-module Docker lifecycle
|
||||
- `modules/hal/sdk/src/feature-handlers/systemd.ts` — detection, install, start
|
||||
@@ -0,0 +1,47 @@
|
||||
# 003 — Who supervises a service
|
||||
|
||||
- **Status:** ONGOING — evidence gathered, options costed, decision open.
|
||||
**No longer blocks Phase 0** (see below).
|
||||
- **Initiated by:** jochen, 2026-08-22, in response to
|
||||
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
|
||||
- **Areas touched:** every module shipping a `systemd/` directory (14), the
|
||||
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
|
||||
`systemd_journal`, the bootstrap scripts.
|
||||
|
||||
## The question
|
||||
|
||||
`002` asked how a containerised node runs a module service, given that a module service is
|
||||
defined today as a systemd unit running `docker compose` against `/services/`. The response
|
||||
was the better question:
|
||||
|
||||
> If it's possible to run systemd inside a container, that's the way to go I think. However,
|
||||
> what would the cost be to step away from systemd to run our services and set it up in a
|
||||
> different way? More hal mesh approach.
|
||||
|
||||
This effort answers the cost half. It does not choose.
|
||||
|
||||
## Summary of findings
|
||||
|
||||
- **The question splits in two**, and the halves have opposite answers. Supervising
|
||||
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
|
||||
files already declare a restart policy, so Docker is already the supervisor. Supervising
|
||||
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
|
||||
only thing between a crash and a dead node.
|
||||
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
|
||||
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
|
||||
recursively unless it sits outside the mesh's own process tree — at which point it is an
|
||||
OS-level supervisor again, just reinvented.
|
||||
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
|
||||
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
|
||||
production then have the same shape rather than a translation layer between them.
|
||||
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
|
||||
shell and a desktop. Those parts are on the host by definition.
|
||||
- One incidental finding: the automatic node rescue that documentation describes **does not
|
||||
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
|
||||
|
||||
Detail and costs in [`analysis.md`](analysis.md).
|
||||
|
||||
## Decision needed
|
||||
|
||||
Which supervision model the mesh adopts, recorded in ADR 0002 before Phase 0 builds
|
||||
anything. The options and their costs are in `analysis.md` under "Options".
|
||||
@@ -0,0 +1,198 @@
|
||||
# Reproducing the mesh network in a lab
|
||||
|
||||
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
|
||||
file location or a queried row.
|
||||
|
||||
---
|
||||
|
||||
## 1. The network is data, not configuration
|
||||
|
||||
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
|
||||
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
|
||||
generated by module hooks from mesh-DB rows:
|
||||
|
||||
| layer | generated by | from |
|
||||
|---|---|---|
|
||||
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
|
||||
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
|
||||
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
|
||||
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
|
||||
|
||||
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
|
||||
results is produced by the same code production runs, which is the difference between testing
|
||||
the network and testing a model of it.
|
||||
|
||||
---
|
||||
|
||||
## 2. The constraint that decides whether the lab works
|
||||
|
||||
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
|
||||
|
||||
```js
|
||||
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
|
||||
|
||||
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
|
||||
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
|
||||
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
|
||||
```
|
||||
|
||||
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
|
||||
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
|
||||
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
|
||||
WireGuard fault rather than an addressing choice.
|
||||
|
||||
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
|
||||
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
|
||||
that regex. The code then treats it exactly as it treats a real hosting provider address.
|
||||
|
||||
This is the single most important fact in this document.
|
||||
|
||||
---
|
||||
|
||||
## 3. The topology being reproduced
|
||||
|
||||
The shape below is what a mesh of this kind looks like: one node with a routable address, one
|
||||
publicly named but behind a household NAT, one stationary workstation, one that roams.
|
||||
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
|
||||
|
||||
| node | profile | site | underlay | WG | accessors |
|
||||
|---|---|---|---|---|---|
|
||||
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
|
||||
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
|
||||
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
|
||||
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
|
||||
|
||||
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
|
||||
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
|
||||
`10.10.0.1` to the node it intends as hub or there will be no hub.
|
||||
|
||||
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
|
||||
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
|
||||
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
|
||||
specific `/32` route to a dead endpoint blackholes rather than falling back.
|
||||
|
||||
The four interesting pairs, all of which the lab must reproduce:
|
||||
|
||||
| pair | behaviour | branch taken |
|
||||
|---|---|---|
|
||||
| anything → `anchor` | `Endpoint` written | underlay non-private |
|
||||
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
|
||||
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
|
||||
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
|
||||
|
||||
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
|
||||
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
|
||||
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
|
||||
imply reachability; the address does."* An earlier version aimed the hub at that node's public
|
||||
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
|
||||
|
||||
---
|
||||
|
||||
## 4. The lab
|
||||
|
||||
```
|
||||
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
|
||||
│
|
||||
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
|
||||
│
|
||||
└── router VM 203.0.113.1 / 192.168.1.1
|
||||
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
|
||||
│
|
||||
br-lan 192.168.1.0/24 identical to production, same host addresses
|
||||
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
|
||||
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
|
||||
|
||||
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
|
||||
underlay NULL profile=workstation site=NULL WG 10.10.0.4
|
||||
```
|
||||
|
||||
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
|
||||
WireGuard plan. Only the public segment is substituted, and only because it must be.
|
||||
|
||||
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
|
||||
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
|
||||
It also gives somewhere to break things — drop the forward and observe whether the mesh
|
||||
notices or whether the public name simply stops working.
|
||||
|
||||
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
|
||||
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
|
||||
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
|
||||
DNS exists.
|
||||
|
||||
---
|
||||
|
||||
## 5. What a node needs before any of this works
|
||||
|
||||
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
|
||||
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
|
||||
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
|
||||
`mesh-ca`, and `traefik` where it serves.
|
||||
|
||||
On disk beforehand, because the node must reach the mesh DB before it can read any of the
|
||||
above: registry database and object-store host and credentials, plus an npm token. This
|
||||
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
|
||||
and is why the `/etc/hosts` floor exists.
|
||||
|
||||
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
|
||||
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
|
||||
leaf.
|
||||
|
||||
---
|
||||
|
||||
## 6. Certificates — the lab issues its own
|
||||
|
||||
**Settled 2026-08-22: the lab runs its own ACME issuer.**
|
||||
|
||||
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
|
||||
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
|
||||
testing, the lab stands up an ACME server on its wan segment.
|
||||
|
||||
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
|
||||
public names from a public authority and internal names from the mesh CA; a lab with one CA
|
||||
would hide any bug living in that split. So:
|
||||
|
||||
| | production | lab |
|
||||
|---|---|---|
|
||||
| public names | a public ACME authority | a test ACME server on the wan segment |
|
||||
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
|
||||
|
||||
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
|
||||
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
|
||||
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
|
||||
tests less than the real thing, not more.
|
||||
|
||||
### It also exercises the port forward
|
||||
|
||||
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
|
||||
|
||||
- the hub is directly reachable on the wan segment — straightforward
|
||||
- **the published-but-NATed node is reachable only through the router's forwarded port**
|
||||
|
||||
So certificate issuance for that node passes only if the forward is correct. That is exactly
|
||||
why its certificate works in production, and it makes "the forward is missing" a reproducible
|
||||
failure rather than a mystery.
|
||||
|
||||
### Required change: `caServer` must be configurable
|
||||
|
||||
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
|
||||
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
|
||||
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
|
||||
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
|
||||
overrides it per node.
|
||||
|
||||
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
|
||||
one means every certificate experiment on a real node consumes production issuance quota, and
|
||||
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
|
||||
|
||||
---
|
||||
|
||||
## 7. Incidental finding: `scope:` is read by nothing
|
||||
|
||||
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
|
||||
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
|
||||
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
|
||||
`modules/unifi/module.yml:52-93` does deliberately.
|
||||
|
||||
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
|
||||
This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is
|
||||
indistinguishable from a wrong one, and costs more, because people believe it.*
|
||||
@@ -0,0 +1,44 @@
|
||||
# 004 — Reproducing the mesh network in a lab
|
||||
|
||||
- **Status:** ONGOING — topology established and mapped; not yet stood up
|
||||
- **Initiated by:** jochen, 2026-08-22 — *"the most difficult part of our VM setup will be
|
||||
the networking part"*
|
||||
- **Areas touched:** `modules/wireguard`, `modules/dnsmasq-app`, `modules/traefik`,
|
||||
`modules/mesh-ca`, `node_accessors`, `nodes.site` / `nodes.underlay_addr`.
|
||||
|
||||
## Summary
|
||||
|
||||
The network is **entirely generated from mesh-DB rows by module hooks**. `install.d` performs
|
||||
no network configuration whatsoever — no WireGuard, no DNS, no firewall. That makes a faithful
|
||||
lab primarily a *data* problem rather than a networking problem, and means the lab exercises
|
||||
the real code path instead of a reimplementation of it.
|
||||
|
||||
One constraint decides whether the lab works at all: the WireGuard endpoint rule tests the
|
||||
underlay address against an RFC1918 regex to decide reachability. **A simulated public segment
|
||||
addressed from RFC1918 space silently prevents the mesh from forming** — no endpoint is written
|
||||
for the hub, so nothing can ever initiate. The simulated public segment must therefore use
|
||||
TEST-NET-3 (`203.0.113.0/24`).
|
||||
|
||||
With that one substitution the lab reproduces the production topology exactly, including the
|
||||
case that is hardest to get right: a node that is publicly *named* but sits behind NAT, whose
|
||||
endpoint the hub can only learn from a handshake.
|
||||
|
||||
Detail in [`analysis.md`](analysis.md).
|
||||
|
||||
## Settled
|
||||
|
||||
**The lab issues its own certificates.** Public names are certified by an ACME server on the
|
||||
lab's wan segment; `.internal` names keep the mesh CA. The lab preserves production's two-CA
|
||||
split rather than collapsing it, because a single-CA lab would hide any bug living in that
|
||||
split. It also makes the router's port forward load-bearing — HTTP-01 must reach the
|
||||
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
|
||||
failure instead of a mystery.
|
||||
|
||||
Requires one change: `caServer` is not set on the reverse proxy today, so it defaults to the
|
||||
public authority's **production** endpoint. It must become configurable, defaulting to
|
||||
production so real nodes are unaffected.
|
||||
|
||||
## Open
|
||||
|
||||
- Not yet stood up. `incus` is declared in `modules/hal/developer/module.yml` and merged
|
||||
(PR #944); the lab itself is unbuilt.
|
||||
@@ -0,0 +1,16 @@
|
||||
# 01-RESEARCH
|
||||
|
||||
Investigations that have not yet hardened into design.
|
||||
|
||||
## Structure
|
||||
|
||||
Each effort lives in `NNN-descriptive-name/` and **must** contain `status.md` with:
|
||||
|
||||
- a short summary of the effort
|
||||
- who initiated it
|
||||
- the areas it touches
|
||||
- current status: `ONGOING`, `GRADUATED`, or `ABANDONED`
|
||||
|
||||
An effort graduates by producing an ADR and a `02-DESIGN` entry. It is abandoned in
|
||||
place — never deleted. What was rejected, and why, is the more expensive half to
|
||||
rediscover.
|
||||
Reference in New Issue
Block a user