HQ — the mesh's own documentation

What the mesh is, what it is becoming, and why. Implementation lives in the
code repositories; the reasoning lives here.

  00-GENESIS   mission, engineering context, effect, and the rules that hold
  01-RESEARCH  investigations, before they harden into design
  02-DESIGN    the authoritative specification
  adr          numbered decisions — what was chosen, and what was rejected
  DECISIONS.md the ledger: every decision, in the order it was taken

Written for a reader who is not its author and has no access to the mesh it
describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918;
nodes are named by role.

Single initial commit by intent. The prior history came from a private
repository and carried operational detail — a routable address identified as a
VPN hub, real domain names, a hosting provider — which sanitising a tip commit
would not have removed from the log.
This commit is contained in:
2026-08-22 22:01:32 +02:00
commit cf9357e8e9
22 changed files with 2376 additions and 0 deletions
@@ -0,0 +1,237 @@
# Current state → ideal state
## 1. The count is a symptom
124 modules. The number is not the problem; the reason for it is. Modules are split
because splitting is the **only granularity lever the platform offers**:
| symptom | module | why it exists |
|---|---|---|
| one daemon must run on one node | `hal/claude-licences` | no per-node feature opt-in |
| a library must not drag a 450 MB dep | `hal/claude` | no way to expose two npm packages |
| a domain needs a daemon *and* a library | `hal/claude-code` + `hal/claude` | one feature of each type per module |
| a handful of SDK verbs need a home | `infra` (10 lines) | tools must belong to *some* module |
A feature is a **handler type**, discovered by directory presence, singleton per module
(`FEATURE_HANDLERS` in `modules/hal/sdk/src/feature-handlers/index.ts`). So "one more
deployable unit" always means "one more module".
Splitting also fragments domains. Claude licence logic sits in three modules —
resolution in `hal/claude`, health in `hal/claude-code`, refresh in
`hal/claude-licences` — so following one token means reading three.
## 2. Two things are conflated in one repo
`noxflow` is 45 tables. Its original purpose — a HAL-native Jira — is 7 of them:
| domain | tables |
|---|---:|
| tasks (the original concept) | 7 |
| agents / HR | 13 |
| knowledge (a Confluence) | 8 |
| meetings | 7 |
| scheduling / ops | 2 |
| other | 8 |
Agents alone are nearly double the concept the module was built for. This is the direct
cause of concrete faults: per-agent Claude credentials had to be written by the noxflow
runtime purely because `agents.claude_account` lives in noxflow's database, while node
identity — the same kind of fact — lives in the mesh registry.
Separately, 91 of the 124 modules are third-party software (`plex`, `sonarr`,
`firefox`, `postgres`). These **run on the mesh; they are not part of it**. They belong
outside this repository. The pipeline and provisioning system are deliberately generic,
so HAL's own modules dogfood exactly what external modules use — that property is what
makes the separation safe.
## 3. Knowledge is split three ways
| store | owner | size | search |
|---|---|---:|---|
| `mesh_docs` (+ revisions) | `hal/hippocampus`, in the mesh DB | 144 docs, 73 touched in 30d | yes |
| `knowledge_*` (8 tables) | `noxflow/runtime`, in the noxflow DB | 20 pages | yes |
| repo markdown | git | `VISION.md`, `FLAVOR-PROPOSAL.md`, `docs/` | **no** |
The librarian is split across two of them: `librarian_file_document` writes via
hippocampus, `librarian_ask` and `librarian_retrieve` read via noxflow. Filing and
retrieval, different modules, different databases.
Both DB stores have revisions; only noxflow has a review workflow (`_proposals`,
`_promotion_requests`). Neither has file access. Git has file access and history but no
search integration. **Every piece needed already exists — none of them together.**
## 4. `hal/developer` and the missing local mesh
`hal/developer` (≈2,600 lines) exposes `dev_stage`, `dev_diff`, `dev_validate`,
`dev_typecheck`, `dev_verify_flavor`, `dev_bootstrap`, `dev_build`, `dev_up`, `dev_down`,
`dev_env`. It is the tooling that both reflects and enforces the module model, so it
moves with any change to that model.
`dev_up` builds a dev environment **for a single module**: resolves `requires:`, starts
providers "via systemctl (same as production)", provisions workspace-namespaced
databases. It borrows the host node's services.
**There is no way to run a mesh locally.** Consequently anything that only manifests
*between* nodes — credential rotation reaching a running session, provision credentials
fanning out, cascade ordering, artifact packaging — is discoverable only in production.
Every fault fixed on 2026-08-22 was of that kind.
A containerised mesh (N nodes, a broker, a registry, a coordinator) is therefore not a
convenience. It is the precondition for a refactor of this size being verifiable at all.
## 5. `hal/sdk` is where domains go to hide
155 files, 34,636 lines. It contains code from essentially every bounded context:
| file | domain it belongs to |
|---|---|
| `installer-core.ts` (1,495) | runtime |
| `tools/provisions.ts` (1,311) | provisioning |
| `tools/deploy-service.ts` (1,081) | deploy |
| `workflow-engine.ts` (970), `task-commands.ts` (588) | **work** — noxflow's domain, inside HAL's SDK |
| `artifact-manager.ts` (948), `build-executor.ts` (760), `feature-handlers/*` | delivery |
| `env-generator.ts` (944) | config |
| `amqp-client.ts` (899) | comms |
| `meshware.ts` (799) | runtime |
| `module-registry.ts` (743) | mesh |
| `claude-credentials.ts` | **ai** |
Nothing here was misplaced carelessly. Each landed in the SDK because the SDK is the one
package every module may import — so putting a shared type there is the way to avoid a
circular dependency between two modules that both need it.
The result is that the dependency graph is trivially acyclic and completely
uninformative: everything depends on one 34k-line package, so a change anywhere in it
rebuilds everything. On 2026-08-22 a three-line fix in `feature-handlers/migrations.ts`
produced a 1,445-job, 26-minute mesh-wide cascade.
The last row is from that same day, and is the clearest illustration: per-agent Claude
config-directory helpers were added to `@hal/sdk/claude-credentials.ts` because it was
the only place both `hal/claude` and the noxflow runtime could import from. The
alternative — `hal/ai` owning it and noxflow depending on `hal/ai` — was unavailable
without answering the ownership question this effort exists to answer.
**Extraction is therefore blocked on the decomposition, not the reverse.** A domain
module can only own its shared code once the domain has a module. The SDK should retain
only what is genuinely cross-cutting: transport, manifest types, logging, process
helpers.
## 6. The crux: two notions of "agent"
The original vision was **a node is an agent with thinking abilities**. Later, noxflow
implemented actual agents — employees with skills, workspaces and tasks. Both survive,
and they collide.
The collision is visible in the data. There are **two agent rows per node**:
| agent | display name | skills | bound to |
|---|---|---|---|
| one named after each node | the node's name | `{deploy,verify,operate}` | its node |
| `hal-<node>`, one per node | "HAL (\<node\>)" | `{}` | its node |
One carries the **work**; the other carries only **identity** — no skills at all, existing
to hold a licence for HAL's own sessions (`resolveHalModuleEnv`). The node-as-agent idea
was implemented twice, half each.
The same split appears in the schema: `nodes.hal_claude_account` and
`agents.claude_account` are the same fact in two tables in two databases. What was
documented as "three licence touchpoints" (NODE, HAL, AGENT) is **one concept modelled
three times**.
### Resolution
**Nodes and agents are decoupled.** The agents module owns agents; agents run on nodes.
A node is a place where an agent can run — that is the whole of the relationship. There
is no resident agent, no node-owned identity, no ownership in either direction.
Agents named after a node still exist, but they are **ordinary agents defined through the
agents module** that happen to hold infra skills and be named after the node they usually run
on. Nothing about them is structural.
Two consequences follow:
1. **`nodes.node_license` and `nodes.hal_claude_account` should not exist.** A node does
not authenticate to a model provider — agents do. Both collapse into
`agents.claude_account`, and what was documented as three licence touchpoints becomes
one.
2. **The pair of agent rows per node merges.** `hal-<node>` exists only to hold a licence
for HAL's own sessions; once licences belong to agents and any module may employ an
agent, it has no reason to be separate from `<node>`.
What remains unaccounted for is the operator's own interactive session:
`~/.claude/.credentials.json` on a node serves a **human at a prompt**, not an agent.
Under this decoupling that is not a node property either — it is the operator's
credential on whichever machine they are sitting at. Whether the operator is modelled as
a persona (soulstream treats humans and agents as peers with identical credentials) is an
open question, and the last remaining reason `nodes` carries a licence column today.
## 7. The record, not "meetings"
soulstream models work as **one signed log** in which humans and agents apply identical
changes, with work arriving by **mention**. Memory and decisions live in that single
auditable stream — no hidden state between components. Its stated economics,
*"cheap awareness and expensive thinking"*, is the same principle as the budget guard:
constant monitoring, selective deep processing.
Under that model "meetings" is not a domain. A meeting is a **thread in the record**.
Which means six of today's modules and table groups are one thing seen from different
angles:
`hal/axon` (DMs, rooms) · `hal/cortex` · `hal/synapse` · `hal/notifications` ·
`conversations` · `meetings`
A notification, in particular, is just a mention not yet read — which is why it needed
its own table, its own id column, and its own failure mode.
## 8. Ideal state
The mesh repository contains only the mesh, decomposed by domain rather than by
deployment accident:
| context | what it is | absorbs today's |
|---|---|---|
| `hal/infra` | **the room** — nodes, registry, provisioning, pipeline, node runtime | `hal/mesh`, `hal/brain`, `hal/coordinator`, `hal/meshware`, `hal/developer`, `hal/bootstrap` |
| `hal/agents` | **the name and the thinking** — Agent owns identity, licence, runs, memory, thoughts | noxflow agents (13 tables), `hal/thoughts` |
| `hal/stream` | **the record** — mentions, threads, meetings, notifications | `hal/axon`, `hal/cortex`, `hal/synapse`, `hal/notifications`, meetings (7), conversations |
| `hal/knowledge` | documents, spaces, revisions, review, search | `hal/hippocampus` + noxflow `knowledge_*` |
| `hal/work` | tasks, workflows, boards — the original noxflow | noxflow tasks (7 tables) |
| `hal/ai` | provider integration, flavored per vendor | `hal/claude*`, `claude-code` |
| `hal/config` | env and config distribution, secrets, PKI | `hal/env-sync`, `hal/config-sync`, `hal/secrets`, `mesh-ca` |
| `hal/observability` | health, dashboards, logs | `hal/health`, `hal/meshboard` |
Eight contexts instead of 33 platform modules, with the catalogue elsewhere.
The naming corrects itself in the process: what is called "brain" today is
infrastructure, and the word survives only as a concept an Agent owns — never again as a
module name.
This shape depends on **named features with per-node opt-in** — without it, every
independently-deployable unit inside a context becomes a module again and the count
returns. That mechanism is the subject of a separate ADR.
## Decided since
| question | decision |
|---|---|
| Meetings — own context or part of agents? | **Neither.** A meeting is a thread in the record; `hal/stream` absorbs it. The category was wrong. |
| Thoughts — autonomy or agent behaviour? | `hal/agents`. Any agent can have a thought. |
| `hal/axon` vs `hal/cortex` | `hal/cortex` is vestigial — no `daemon/` source, unit inactive on every node, description copy-pasted from axon. Messaging belongs to `hal/stream`. |
| `hal/hypothalamus` | Orphaned build output. A `dist/` directory with no `module.yml`, so not a module at all. |
| Who owns agent identity? | `hal/agents`. Nodes and agents are decoupled; a node is only a place an agent can run. |
| Where do third-party modules live? | Outside this repository. They run *on* the mesh, not *of* it. |
## Open questions
1. **`hal/scheduler`** — infrastructure (fire an event later, belongs to `hal/mesh`) or
part of `hal/work`?
2. **The executor.** Pulling and routing work items is `hal/work`; spawning a session is
`hal/agents`; providing the place to run is `hal/mesh`. It currently straddles all
three, which is the same kind of straddle that put credential-writing in the noxflow
runtime.
3. **Catalogue destination** — one repository, or per-application repositories, given the
pipeline resolves dependencies across the registry rather than the filesystem?
4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or
does that belong to `hal/stream`? Everything imports it, which argues both ways.
5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which
user, on which node, a human agent acts as. Where does it live — an attribute of the
agent, or of the agent-node binding?
@@ -0,0 +1,46 @@
# 001 — Module domain decomposition
- **Status:** ONGOING — ADR 0001 accepted; graduates when `02-DESIGN` carries the per-context specifications
- **Initiated by:** jochen, 2026-08-22
- **Areas touched:** every `hal/*` and `noxflow/*` module; the pipeline's dependency
graph; the knowledge base; agent identity and credentials.
## Summary
HAL has 124 modules. That number is not a maintenance problem in itself — it is the
**symptom of missing bounded contexts**. Modules are split not because they model
different domains, but because splitting is the only lever the platform offers:
- no way to run one daemon on one node without making it a module
(`hal/claude-licences` — one daemon, single-node)
- no way to expose two of a kind from one module
- no namespace separating the mesh from the software it runs
This effort establishes the **current state**, the **ideal state**, and the sequence
between them.
## Trigger
A night of debugging that produced four fixes and one conclusion. Every fault was a
boundary fault:
- Per-agent Claude credentials had to be written by the *noxflow runtime*, because
`agents.claude_account` is in noxflow's database — even though agent identity is a
mesh concept and node identity already lives in the mesh registry.
- Whether `hal/brain` may depend on noxflow took three attempts to answer, twice
wrongly, because the ownership boundary was never stated.
- Authoritative documentation existed in `mesh_docs` and was not found, while a
proposal in repo markdown was invisible to search entirely.
## Decisions taken (2026-08-22)
| Question | Decision |
|---|---|
| What should noxflow become? | Decompose into `hal/*` modules; noxflow returns to tasks/workflows |
| Who owns agent identity? | `hal/agents` — a mesh concept, alongside nodes |
| Where do third-party apps live? | Out of this repo. They run *on* the mesh; they are not *of* it |
| Knowledge structure | Modelled on `papa-hq`; implementation choice left open |
## Open questions
Tracked in [`analysis.md`](analysis.md) under "Open questions".
+258
View File
@@ -0,0 +1,258 @@
# A mesh that runs locally — current state and obstacles
Evidence for Phase 0. Every claim here is either a file location or something measured on
2026-08-22; where a claim was checked and found false, that is recorded too.
---
## 1. What exists today, and why neither is a mesh
### `test/pipeline/` — the only containerised HAL node, and it is dead
This harness builds `@hal/sdk` and the `hal/brain` daemon into a `node:23-alpine` image,
compiles `hal/coordinator`'s tools against it, and runs `brainstem.js` with
`HAL_MODE=daemon` against a containerised postgres and LavinMQ
(`test/pipeline/Dockerfile`, `test/pipeline/docker-compose.yml`). It proves the valuable
thing: **a node runtime needs no systemd, no `/services/`, and no nvm** — environment
variables and a seeded `nodes` row are sufficient.
It also cannot build. Measured 2026-08-22:
```
Step 7/16 : RUN npm run build -w modules/hal/sdk
npm error No workspaces found:
npm error --workspace=modules/hal/sdk
```
The npm workspace it depends on was removed on 2026-06-04 (`21ef4a4e`, "kill npm
workspace"), and the root `package.json` now carries a comment explaining why it will not
come back. The harness has not been touched since before that commit.
So the repository's only end-to-end pipeline test has been unrunnable for two and a half
months and nothing reported it — which is the same shape as the faults Phase 0 exists to
catch. **A test nobody runs is indistinguishable from a test that passes.**
### `test/dev-mesh/` — a work plane, not a control plane
Four `runtime-*` services, a noxflow API, a meshboard and a shared postgres + LavinMQ on
one network (`test/dev-mesh/docker-compose.yml`). It is genuinely useful and the topology
is a good precedent: one network, per-service `HAL_NODE`, `EXECUTOR_STUB=1` so dispatch and
run lifecycle are real while the model call is faked.
It is the wrong layer for Phase 0, on three counts:
| | dev-mesh does | Phase 0 needs |
|---|---|---|
| state | restores `pg_dump`s (`test/dev-mesh/init.sh`) | schema built by **migrations**, or fixture 3 is untestable by construction |
| code | bind-mounts pre-built `dist/` from the host | artifacts **built and delivered** by the pipeline |
| scope | noxflow runtime, API, meshboard | `hal/coordinator`, `hal/meshware`, MinIO, provisioning |
No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises
what the mesh *runs*; Phase 0 needs what the mesh *is*.
### `dev_up` — confirmed host-coupled
The claim in the work breakdown holds. `startDevProvider()` starts providers through the
host's user systemd instance and reads credentials from a host path
(`modules/hal/developer/tools/dev-env.ts:146-178`):
```ts
const unitName = `hal-module@${provider}.service`;
execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" });
const creds = resolveRunningProviderCreds(provider); // reads /services/<provider>/.env
```
It borrows the host. There is no mesh to stand up.
---
## 2. What is already portable
- **Node identity is one environment variable.** `resolveNodeName()` is
`process.env.HAL_NODE || hostname()`, lowercased (`modules/hal/sdk/src/mesh-config.ts:215`).
No file, no registration handshake, no host coupling.
- **Registration is a plain upsert.** `mesh_node_register` inserts into `nodes`
(`modules/hal/mesh/tools/index.ts:490-540`); mandatory fields are `name` and `user_name`.
- **The daemons are AMQP loops.** `hal/meshware` consumes `cmd.feature.{build,install,configure,start}`
(`modules/hal/meshware/daemon/src/cerebellum.ts:780-857`); `hal/coordinator` consumes
`event.gitea.push` and drives the cascade. Both are configured entirely by environment.
- **`hal/brain` in daemon mode already runs in a plain node image**, per `test/pipeline/`.
---
## 3. The obstacles, located
Four couplings stand between the daemons and a container. All are in code, none are
mysterious.
| # | coupling | location |
|---|---|---|
| 1 | meshware restarts itself via the host: `execFileSync("systemctl", ["--user","restart","hal-meshware.service"])` | `modules/hal/meshware/daemon/src/cerebellum.ts:826` |
| 2 | meshware restarts the broker via a hardcoded host path: `execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"])` | `cerebellum.ts:843-844` |
| 3 | **a module service *is* a systemd unit** — `hal-module@.service` runs `docker compose --project-directory /services/%i` | `modules/hal/meshware/systemd/hal-module@.service:10-11` |
| 4 | env generation writes `homedir()`-relative host paths — `~/.config/hal/env`, `~/.config/hal/modules/*.env`, `/services/*/.env` | `modules/hal/sdk/src/env-generator.ts:188-223`, `:538-568`, `:584-606` |
1, 2 and 4 are small: a guard and a configurable root. **3 is the architectural one**, and
it is the first open question below — in a container, "systemd unit that runs docker
compose" has no natural translation.
The bootstrap scripts add their own host assumptions — nvm for node resolution, `sudo
mkdir -p /services && chown`, and a `~/dotfiles` repository that is not in this repo
(`install.d/adopt.sh:165-170`) — but Phase 0 does not have to run them. A container image
can be built the way `test/pipeline/Dockerfile` builds one, bypassing the bootstrap path
entirely. That is a deliberate divergence to record: **the local mesh would not test the
bootstrap**, only the running mesh.
---
## 4. External services, and whether they can be local
| service | used for | local substitute |
|---|---|---|
| PostgreSQL | registry + pipeline state | yes — already the pattern in both harnesses |
| LavinMQ | all coordinator↔meshware messaging | yes — `cloudamqp/lavinmq`, already used |
| MinIO | per-feature artifact tarballs, `modules/{name}/{version}/{feature}.tar.gz` (`modules/hal/sdk/src/artifact-manager.ts:21,65`) | yes — `minio/minio`, needs only `REGISTRY_MINIO_*` |
| Gitea | source, push webhook, **and** the `@hal/*` npm registry | container exists, but heavy; see open question 2 |
| Docker registry | `docker push` for modules declaring `docker:` (`build-executor.ts:195-231`) | `registry:2`, or avoid modules that need it |
| Traefik | writes routing config at install; does not gate success | omit |
Only Gitea is awkward, and only because it carries two roles at once.
---
## 5. The fixtures
Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they
differ sharply in cost.
### Fixture B — a provider deploy rotates a shared credential without fanning out
**Cheapest, best understood, and the root cause is still open.** Documented at
`troubleshooting/provision-adoption-rotates-live-credential`. The mechanism is two lines:
```ts
async provision(project, username, options) {
const password = generatePassword(); // ALWAYS a fresh password
```
and, in the adoption branch of `provisionDatabase`, an `ALTER ROLE ... WITH PASSWORD` that
rotates the live secret while updating only the provider's own `mesh_provisions` row
(`modules/postgres/tools/index.ts`). Every consumer sharing that role keeps a stale
password and fails permanently; the provider recovers alone, and that asymmetry is the
tell.
Reproduction needs one provider node and two consumer nodes — the minimum interesting
mesh. It re-fires on every provider deploy, so it does not need to be provoked, only
observed. Confirmation is a timestamp comparison, not a hash: the stored values are
`enc:v1:` with a random IV, so identical plaintexts hash differently.
Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors
each; the provider failed for 22 minutes and recovered by itself.
### Fixture C — a migration ships nothing while the pipeline reports success
Three independent silent-skip points, any one of which produces it:
1. **Not-applicable and succeeded are the same status.** `MigrationsHandler.detect()` is
`existsSync(join(moduleDir, "migrations"))` against the freshly cloned workspace
(`modules/hal/sdk/src/feature-handlers/migrations.ts:32-41`). If the directory is not
there — uncommitted, ignored, or a `module_path` that does not line up — the build
reports `success` (`cerebellum.ts:634-638`).
2. **Upload failure is a warning.** `uploadFeatureArtifact()` returns `{sha256: ""}` with a
`warn` when none of the listed files exist (`artifact-manager.ts:41-44`), and callers
catch and warn rather than throw (`cerebellum.ts:672-673`). The symmetric download
failure on the target node is also non-fatal (`artifact-manager.ts:93-98`,
`cerebellum.ts:426-431`, commented "Non-fatal: feature may work without artifact").
3. **Missing packaged files are skipped in silence.** `addFileEntries()` does
`if (!existsSync(abs)) continue;` (`module-builder-core.ts:134-135`) — an explicit
`package:` entry that is not on disk simply is not in the tarball.
This is the same class as `troubleshooting/flavors-never-packaged` (`flavors/` was never
staged, so a node kept its first copy forever) and is catalogued in
`troubleshooting/deploy-reports-transport-not-effect`, which counted **0 of 119 modules
implementing `verify`** — the stage that exists, is dispatched, has a working handler, and
would turn every one of these into a red pipeline.
### Fixture A — a credential rotation does not reach a running session
The most valuable and the most work: it needs a *running session* to rotate underneath,
which means the local mesh must be able to start one. The mechanism is understood — an
`EnvironmentFile` is read once and `process.env` is a snapshot — and it is why credentials
became files that get re-read. Defer it behind B and C.
**Recommendation:** B first. It needs three nodes and a provision, no build and no session,
and it is the one whose root cause is still open — so reproducing it locally has value
beyond proving the harness.
---
## 6. A claim that was checked and found false
The survey behind this document suspected that the live AMQP pipeline never writes
`deployments` or `node_modules.installed_version`, because the code that does so sits in
`installer-core.ts:1045-1126` on what is commented as the "CLI path".
Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours:
| node | deploys | newest |
|---|---|---|
| 1 | 24 | 17:26:17 |
| 2 | 27 | 17:22:22 |
| 3 | 47 | 17:25:54 |
| 4 | 26 | 17:22:23 |
The record is live on all four nodes and minutes old. **The claim is refuted.** Recorded
because the reasoning was plausible and someone will retrace it.
---
## 7. Questions
### Settled (jochen, 2026-08-22)
**Phase 0 is a development environment, not a fixture rig.** It is built as a supported
surface to work in daily, and it is what `dev_up` should become. The definition of done —
"demonstrated in the local mesh" — is therefore meant literally.
**The trigger is a real Gitea container.** The webhook relay is part of what is under test,
including the changed-file enrichment that exists because the webhook truncates at 20
commits (`modules/hal/gitea/tools/index.ts:136-195`). A synthetic `event.gitea.push` would
skip it. Gitea also carries the `@hal/*` npm registry, so the local mesh needs it twice
over either way.
### Open — blocking
1. **How does a node supervise a module service?** Today a module service is a systemd unit
running `docker compose` against `/services/<module>`
(`modules/hal/meshware/systemd/hal-module@.service:10-11`), which has no direct
translation inside a container. The question raised in response is the better one, and
is broader than Phase 0: **what would it cost to stop using systemd altogether and have
the mesh supervise its own services?** That is a mesh-level architectural question, not
a local-mesh implementation detail — if the answer is that the mesh should own
supervision, Phase 0 should not build a container-only workaround first. Under
investigation; findings will land in [`003-service-supervision`](../003-service-supervision/)
and, if it goes ahead, ADR 0002.
2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable
root keeps one code path; container-specific targets keep the host paths untouched. This
decides whether env generation gets a seam or a conditional.
3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap
scripts, and it will run one OS where the real mesh is heterogeneous by design
(`00-GENESIS/context.md`). Stating the divergence up front is what stops "it works
locally" from becoming its own class of silent failure. Now sharper, because a
development environment people use daily is trusted far more than a rig — and drifting
from production costs correspondingly more.
---
## References
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this
phase unblocks
- [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) — Phase 0 tasks
and checkpoint
- `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open
- `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped
green, and the unused `verify` stage
- `troubleshooting/flavors-never-packaged` — fixture C, previously seen in the wild
+53
View File
@@ -0,0 +1,53 @@
# 002 — A mesh that runs locally
- **Status:** GRADUATED — the design is [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md)
- **Initiated by:** jochen, 2026-08-22
- **Areas touched:** `install.d/`, `hal/meshware`, `hal/coordinator`, `hal/brain`,
`hal/developer` (`dev_up`), `hal/sdk` (env generation, feature handlers, artifact
manager), `test/pipeline/`, `test/dev-mesh/`, the provisioning path in
`modules/postgres/`.
## Summary
Phase 0 of [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) requires
a mesh that comes up in containers, runs its own pipeline, and reproduces known faults on
demand. Nothing else in the decomposition starts until it exists, because every fault the
decomposition addresses was found in production — there was nowhere else to find it.
This effort establishes what already runs in a container, what is welded to the host, and
what it would take to close the gap. It does **not** choose an approach: the central
question — how a containerised node executes a module service, when a module service is
defined today as a systemd unit shelling to `docker compose` in `/services/` — is not
answered by ADR 0001 and is recorded below rather than decided.
## What was established
- The two existing container harnesses are **neither of them a mesh**, and one of them has
not been able to build since 2026-06-04.
- Node identity is already portable — a single environment variable, no host handshake.
- Four concrete host couplings block a containerised node, all with known locations.
- All three Phase 0 fixtures are reproducible; one of them is documented in the knowledge
base with an open root cause and is the cheapest place to start.
Detail and evidence in [`analysis.md`](analysis.md).
## Questions — all settled 2026-08-22
**Phase 0 is a development environment**, not a fixture rig — it is what the host-borrowing
dev tooling becomes.
**The trigger is a real source-forge container**, because the webhook relay is part of what
is under test.
**A node is a system container**, promotable to a virtual machine per node. This answered
the question the effort was stuck on, and dissolved it rather than solving it: against a
real node with a real init, the four host couplings catalogued in `analysis.md` §3 are not
couplings — they are how a node works. They were obstacles only to a node modelled as an
application container.
Consequently [`003-service-supervision`](../003-service-supervision/) **no longer blocks
Phase 0**. It remains a live architecture question, on its own timeline.
The remaining questions in `analysis.md` — what replaces host paths in a container, and how
faithful the lab must be — are answered in the design: nothing replaces them, because the
paths are real; and the divergences are enumerated rather than discovered.
@@ -0,0 +1,225 @@
# Who supervises a service — the cost of leaving systemd
Measured 2026-08-22. Every claim is a file location or a count.
---
## 1. What systemd actually does for HAL
**14 modules** ship a `systemd/` directory. The units divide cleanly into three kinds:
| kind | count | what it is |
|---|---|---|
| long-running daemons | 9 | `hal/brain`, `hal/meshware`, `hal/coordinator`, `hal/cortex`, `hal/env-sync`, `hal/file-share`, `hal/thoughts`, `noxflow/runtime`, `desktop_notifications` — all `Restart=always`, `RestartSec=10` |
| scheduled one-shots | 6 timers | docker-prune, docker-registry maintenance, claude-code sessions + usage, health, mailu cert-sync |
| the template | 1 | `hal-module@.service` — the per-module Docker lifecycle |
Only `mailu` is system-scope. Everything else is a user unit under
`~/.config/systemd/user`, auto-detected from directory presence — the module author names
the file and that filename *is* the unit name
(`modules/hal/sdk/src/feature-handlers/systemd.ts:31-42`).
The template is the piece `002` tripped over
(`modules/hal/meshware/systemd/hal-module@.service`):
```
ExecStart=docker compose --project-directory /services/%i --project-name %i up --remove-orphans --pull always
ExecStop=docker compose --project-directory /services/%i --project-name %i down
Restart=on-failure
StartLimitIntervalSec=120
StartLimitBurst=5
```
Systemd's roles, then: restart-on-failure, start at login, env-file loading, ordering,
per-module lifecycle, timers, and — via journald — **the only log store the 9 Node daemons
have** (`modules/hal/sdk/src/tools/log-tail.ts:50-53`).
---
## 2. The question splits in two, and the halves disagree
### For `docker compose` stacks, systemd is mostly redundant
**44 of 44** module `docker-compose.yml` files declare a container-level `restart:` policy —
`always` or `unless-stopped`, none missing. Docker's own daemon already restarts crashed
containers, independently of systemd.
`hal-module@.service` does not supervise the containers. It supervises the **`docker compose
up` foreground process**. It is a second layer on top of a restart policy that already
works. What it genuinely adds is narrower than it looks:
- one uniform verb for every module type — `systemctl --user start hal-module@X` rather than
remembering each module's compose invocation, relied on across a dozen call sites in
`meshware.ts:124-154`, `dev-env.ts:142-163`, `installer-core.ts`, `infra.ts:166`
- recovery when the `docker compose up --pull always` process itself dies — a failed image
pull, not a crashed container
- rate-limited restart (`StartLimitBurst=5`) so a broken stack does not spin
That is real, but it is a convenience layer, not a safety layer. **This half could go at
moderate cost.**
### For the 9 Node daemons, systemd is load-bearing
There is no alternative supervisor anywhere in the repo. No PM2, no forever, no nodemon, no
watchdog loop — all checked, zero hits. `Restart=always` is the only thing standing between
a crashed daemon and a dead node.
**This half is the actual question.**
---
## 3. The hard part is fate-sharing
The difficulty is not systemd. It is that **a supervisor must not share fate with what it
supervises**, and the codebase already has a scar from exactly this.
meshware cannot restart itself mid-request: killing the process before it ACKs the AMQP
message loses the message. So it defers its own restart by two seconds after closing the
connection (`modules/hal/meshware/daemon/src/cerebellum.ts:815-828`) — a commented
workaround for a problem that only exists because the thing being restarted is the thing
doing the restarting.
Any mesh-native supervisor inherits this recursively. Something has to be the outermost
always-alive process, and if it is written by the mesh, the mesh must supervise it, and so
on. The recursion only terminates at a process the mesh does not own. Today that is
systemd. **A "more mesh" supervisor that is itself a mesh process is not a smaller problem;
it is the same problem with a new name.**
This is the strongest argument for the status quo, and it is worth stating plainly before
looking at alternatives.
---
## 4. The option neither of us named
There is a third answer that terminates the recursion in something that is not systemd and
not written by us: **run HAL's own daemons as containers.**
Docker is already the outermost supervisor for 44 of 44 module stacks. It does not share
fate with the mesh. It already has restart policies, backoff, and a log store that
`log_tail` already speaks (`log-tail.ts:43-47` reads `docker logs` for Docker modules
today). Extending it from "the things the mesh runs" to "the mesh itself" is not new
machinery — it is applying machinery the mesh already trusts to one more case.
What that buys, beyond supervision:
- **Phase 0 stops being a translation.** A containerised node becomes the same shape as a
production node, rather than a local approximation with a systemd-shaped hole in it. The
divergence that `002` open question 3 worries about largely disappears.
- **`/services/` and `~/.config/hal/` stop being special.** Mounts, not host paths.
- **The GENESIS "dogfood everything" value gets easier**, not harder: the mesh's own
components would ship and run exactly like everything else it carries.
What it costs, honestly:
- **journald → docker logs** for the 9 daemons. `log_tail` already handles both, but
`systemd_journal` and the health checks that read unit state
(`modules/hal/mesh/health.sh:34-40`, `modules/hal/health/hal-health.sh:335-395`) would
need a container-aware path.
- **Boot start** becomes Docker's `restart: always` plus the Docker daemon being enabled at
boot — which is still one systemd unit, but the OS's own, not ours.
- **A container needs the host to be reachable** for anything that touches the node itself.
Some of these daemons exist precisely to write host files.
---
## 5. Why it cannot be all-or-nothing — and ADR 0001 already says so
Some of what runs under systemd today **cannot** be containerised, and the reason is
already in the domain model. ADR 0001:
> a non-human agent acts through a spawned session — a human agent acts through a shell or
> desktop
`hal/brain` has two modes (`modules/hal/brain/daemon/src/brainstem.ts:7-13`): a daemon mode
that is an AMQP relay, and a **cortex mode that is an MCP server over stdio for an
interactive Claude Code session**. The second is a human agent's modality. It runs in the
human's shell, on the human's node, against the human's `~/.claude`. Containerising it is
not a hard engineering problem, it is a category error.
The same holds for `desktop_notifications` and everything `hal/desktop-environment` touches.
So the line is not "systemd or not". It is:
| | belongs where |
|---|---|
| mesh daemons — meshware, coordinator, env-sync, thoughts, file-share, brain **in daemon mode** | supervisable by Docker; candidates to containerise |
| human-modality surfaces — brain **in cortex mode**, desktop notifications, desktop environment | on the host, by definition |
| module stacks | already Docker; systemd layer is the redundant part |
This split is not a compromise between the options. It is what the domain model implies,
and it is a decent sign that the model is doing work.
---
## 6. Options, with costs
| | option | cost | what it buys |
|---|---|---|---|
| **A** | **Keep systemd; systemd-in-container for Phase 0** | privileged containers, heavy images, slow iteration; local mesh keeps a shape production does not have | nothing changes in production; smallest change to the mesh |
| **B** | **Drop the `hal-module@` layer only** — let Docker's restart policies supervise stacks directly | reimplement uniform start/stop across ~12 call sites; lose rate-limited restart and pull-failure recovery | removes the redundant layer; does **not** solve Phase 0 on its own, since the daemons still need supervising |
| **C** | **Containerise the mesh daemons; Docker supervises** | container-aware `systemd_journal`/health; host access for daemons that write host files; the human-modality surfaces stay on the host regardless | terminates the fate-sharing recursion without writing a supervisor; makes local and production the same shape; Phase 0 becomes much less of a special case |
| **D** | **Write a mesh-native supervisor** | the fate-sharing recursion (§3), plus matching systemd's maturity — backoff, resource limits, clean SIGTERM (`noxflow/runtime` already depends on `TimeoutStopSec=60`) — on machines that are somebody's daily driver | most "mesh"; least justified by the evidence |
**D is the option the phrasing "more hal mesh approach" points at, and the evidence argues
against it.** Supervision is not a domain concern the mesh is better placed to solve than
the OS; the mesh's distinguishing feature is brokering capabilities, not restarting
processes. C gets the benefit D is reaching for — the mesh not depending on host-specific
init — without the recursion.
**C and B compose.** C is the one that pays for Phase 0.
Worth noting what is *not* in this table: the dozens of `systemctl` call sites that
configure the **host OS's own** units — NetworkManager, resolved, oomd, zram, docker.service,
fail2ban, sshd, zfs, across `modules/asusd`, `g14-power`, `wireguard`, `sshd`, `zfs` and
others. Those are not HAL supervising itself; they are HAL configuring an Arch box. They
are out of scope for every option above and do not go away under any of them.
---
## 7. An incidental finding
Documentation describes an automatic node rescue: `hal-rescue.sh:20` states it is "triggered
automatically by `hal-health.timer` when hal-meshware is failed".
**It is not.** `hal-health.sh` contains no call to `hal-rescue.sh` (checked, zero matches),
and **no unit in the repository declares `OnFailure=`** (checked, zero matches). The only
real triggers are the manual `rescue_node` tool
(`modules/hal/sdk/src/tools/deploy-rescue.ts:77-92`) and running `install.d/rescue.sh` by
hand.
This is worth recording for two reasons. It weakens any argument that systemd-adjacent
self-healing is already wired — it is not. And it is another instance of the pattern this
whole refactor is about: **a documented mechanism that does not exist, believed because it
was written down.** `00-GENESIS/how-we-build.md` calls this out as a rule; here it is again,
found by grep.
---
## 8. Open question
**Which supervision model does the mesh adopt?** A, B, C, D or a combination.
**This no longer gates Phase 0.** When this was written, the local mesh was assumed to be
built from application containers, which forced the question — there is no natural way to
run an init system inside one. The decision of 2026-08-22 to build development nodes as
**system containers** (see [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md))
removes that pressure entirely: a system container runs a real init, so the existing model
works unmodified and the lab needs no answer here to exist.
What remains is the question on its own merits, which is worth keeping open because the
evidence above still holds: option B removes a layer that 44 of 44 module stacks have already
made redundant, and option C would let the mesh stop depending on host-specific init. Neither
is urgent. Both are now cheap to *try*, because there is somewhere to try them.
Option D — a mesh-written supervisor — remains the one the evidence argues against, for the
fate-sharing reason in §3.
## References
- [`002-local-mesh`](../002-local-mesh/analysis.md) — the effort this came out of
- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — agent modality, which
decides what cannot leave the host
- `modules/hal/meshware/daemon/src/cerebellum.ts:815-828` — the self-restart workaround
- `modules/hal/meshware/systemd/hal-module@.service` — the per-module Docker lifecycle
- `modules/hal/sdk/src/feature-handlers/systemd.ts` — detection, install, start
@@ -0,0 +1,47 @@
# 003 — Who supervises a service
- **Status:** ONGOING — evidence gathered, options costed, decision open.
**No longer blocks Phase 0** (see below).
- **Initiated by:** jochen, 2026-08-22, in response to
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
- **Areas touched:** every module shipping a `systemd/` directory (14), the
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
`systemd_journal`, the bootstrap scripts.
## The question
`002` asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running `docker compose` against `/services/`. The response
was the better question:
> If it's possible to run systemd inside a container, that's the way to go I think. However,
> what would the cost be to step away from systemd to run our services and set it up in a
> different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
## Summary of findings
- **The question splits in two**, and the halves have opposite answers. Supervising
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
files already declare a restart policy, so Docker is already the supervisor. Supervising
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
only thing between a crash and a dead node.
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
recursively unless it sits outside the mesh's own process tree — at which point it is an
OS-level supervisor again, just reinvented.
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
production then have the same shape rather than a translation layer between them.
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes **does not
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
Detail and costs in [`analysis.md`](analysis.md).
## Decision needed
Which supervision model the mesh adopts, recorded in ADR 0002 before Phase 0 builds
anything. The options and their costs are in `analysis.md` under "Options".
+198
View File
@@ -0,0 +1,198 @@
# Reproducing the mesh network in a lab
Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a
file location or a queried row.
---
## 1. The network is data, not configuration
`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at
all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is
generated by module hooks from mesh-DB rows:
| layer | generated by | from |
|---|---|---|
| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` |
| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` |
| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` |
| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB |
**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that
results is produced by the same code production runs, which is the difference between testing
the network and testing a model of it.
---
## 2. The constraint that decides whether the lab works
`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`:
```js
const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a);
if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN
else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public
// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake
```
A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the
hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can
initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a
WireGuard fault rather than an addressing choice.
**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by
RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by
that regex. The code then treats it exactly as it treats a real hosting provider address.
This is the single most important fact in this document.
---
## 3. The topology being reproduced
The shape below is what a mesh of this kind looks like: one node with a routable address, one
publicly named but behind a household NAT, one stationary workstation, one that roams.
Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout.
| node | profile | site | underlay | WG | accessors |
|---|---|---|---|---|---|
| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) |
| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) |
| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) |
| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) |
**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'`
*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign
`10.10.0.1` to the node it intends as hub or there will be no hub.
**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null
`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL`
site means roaming and hub-only — deliberately, because WireGuard has no failover and a more
specific `/32` route to a dead endpoint blackholes rather than falling back.
The four interesting pairs, all of which the lab must reproduce:
| pair | behaviour | branch taken |
|---|---|---|
| anything → `anchor` | `Endpoint` written | underlay non-private |
| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located |
| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private |
| `laptop` → anything | hub only, always initiates | `site` is `NULL` |
The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224`
record what it cost to get right: testing `profile === "server"` was tried and was wrong,
because a home-hosted node **is** a server yet is not publicly reachable — *"role does not
imply reachability; the address does."* An earlier version aimed the hub at that node's public
name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake.
---
## 4. The lab
```
br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public
│
├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1
│
└── router VM 203.0.113.1 / 192.168.1.1
│ NAT, plus one forwarded port to reproduce a published-but-NATed node
│
br-lan 192.168.1.0/24 identical to production, same host addresses
├── a 192.168.1.135 profile=server site=home WG 10.10.0.2
└── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3
c — attach to br-lan, or br-wan ("away"), or detach ("asleep")
underlay NULL profile=workstation site=NULL WG 10.10.0.4
```
Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire
WireGuard plan. Only the public segment is substituted, and only because it must be.
**The router earns its own VM.** It is what makes the published-but-NATed case real: that node
is reachable from outside only through a forwarded port, and the hub must learn its endpoint.
It also gives somewhere to break things — drop the forward and observe whether the mesh
notices or whether the public name simply stops working.
**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names
need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an
`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own
DNS exists.
---
## 5. What a node needs before any of this works
Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`),
`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it,
`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`,
`mesh-ca`, and `traefik` where it serves.
On disk beforehand, because the node must reach the mesh DB before it can read any of the
above: registry database and object-store host and credentials, plus an npm token. This
ordering — contact the mesh before the mesh has configured you — is itself worth reproducing,
and is why the `/etc/hosts` floor exists.
Everything else is generated: the WireGuard keypair locally (the private key never leaves the
node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS
leaf.
---
## 6. Certificates — the lab issues its own
**Settled 2026-08-22: the lab runs its own ACME issuer.**
Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public
reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate
testing, the lab stands up an ACME server on its wan segment.
**The lab keeps production's two-CA split rather than collapsing it.** Production issues
public names from a public authority and internal names from the mesh CA; a lab with one CA
would hide any bug living in that split. So:
| | production | lab |
|---|---|---|
| public names | a public ACME authority | a test ACME server on the wan segment |
| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged |
**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers
deliberately vary their behaviour — validation timing, nonce handling, chain composition — to
expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite
tests less than the real thing, not more.
### It also exercises the port forward
HTTP-01 means the issuer must reach the node being certified on port 80. In the lab:
- the hub is directly reachable on the wan segment — straightforward
- **the published-but-NATed node is reachable only through the router's forwarded port**
So certificate issuance for that node passes only if the forward is correct. That is exactly
why its certificate works in production, and it makes "the forward is missing" a reproducible
failure rather than a mystery.
### Required change: `caServer` must be configurable
`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address
and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's
*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag
fed by an environment value, defaulting to production so real nodes are unaffected and the lab
overrides it per node.
Worth noting independently of the lab: aiming at the production endpoint rather than a staging
one means every certificate experiment on a real node consumes production issuance quota, and
a retry loop can exhaust it for a week. The lab issuer removes that exposure.
---
## 7. Incidental finding: `scope:` is read by nothing
Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`,
`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is
**referenced by no code** in the firewall path. Real scoping is done with `from:`, as
`modules/unifi/module.yml:52-93` does deliberately.
So a manifest can appear to restrict a port to the public scope and in fact restrict nothing.
This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is
indistinguishable from a wrong one, and costs more, because people believe it.*
+44
View File
@@ -0,0 +1,44 @@
# 004 — Reproducing the mesh network in a lab
- **Status:** ONGOING — topology established and mapped; not yet stood up
- **Initiated by:** jochen, 2026-08-22 — *"the most difficult part of our VM setup will be
the networking part"*
- **Areas touched:** `modules/wireguard`, `modules/dnsmasq-app`, `modules/traefik`,
`modules/mesh-ca`, `node_accessors`, `nodes.site` / `nodes.underlay_addr`.
## Summary
The network is **entirely generated from mesh-DB rows by module hooks**. `install.d` performs
no network configuration whatsoever — no WireGuard, no DNS, no firewall. That makes a faithful
lab primarily a *data* problem rather than a networking problem, and means the lab exercises
the real code path instead of a reimplementation of it.
One constraint decides whether the lab works at all: the WireGuard endpoint rule tests the
underlay address against an RFC1918 regex to decide reachability. **A simulated public segment
addressed from RFC1918 space silently prevents the mesh from forming** — no endpoint is written
for the hub, so nothing can ever initiate. The simulated public segment must therefore use
TEST-NET-3 (`203.0.113.0/24`).
With that one substitution the lab reproduces the production topology exactly, including the
case that is hardest to get right: a node that is publicly *named* but sits behind NAT, whose
endpoint the hub can only learn from a handshake.
Detail in [`analysis.md`](analysis.md).
## Settled
**The lab issues its own certificates.** Public names are certified by an ACME server on the
lab's wan segment; `.internal` names keep the mesh CA. The lab preserves production's two-CA
split rather than collapsing it, because a single-CA lab would hide any bug living in that
split. It also makes the router's port forward load-bearing — HTTP-01 must reach the
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure instead of a mystery.
Requires one change: `caServer` is not set on the reverse proxy today, so it defaults to the
public authority's **production** endpoint. It must become configurable, defaulting to
production so real nodes are unaffected.
## Open
- Not yet stood up. `incus` is declared in `modules/hal/developer/module.yml` and merged
(PR #944); the lab itself is unbuilt.
+16
View File
@@ -0,0 +1,16 @@
# 01-RESEARCH
Investigations that have not yet hardened into design.
## Structure
Each effort lives in `NNN-descriptive-name/` and **must** contain `status.md` with:
- a short summary of the effort
- who initiated it
- the areas it touches
- current status: `ONGOING`, `GRADUATED`, or `ABANDONED`
An effort graduates by producing an ADR and a `02-DESIGN` entry. It is abandoned in
place — never deleted. What was rejected, and why, is the more expensive half to
rediscover.