commit cf9357e8e950e8f70666f88e74860070676b3066 Author: jochen Date: Sat Aug 22 22:01:32 2026 +0200 HQ — the mesh's own documentation What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log. diff --git a/00-GENESIS/README.md b/00-GENESIS/README.md new file mode 100644 index 0000000..3dcb6ff --- /dev/null +++ b/00-GENESIS/README.md @@ -0,0 +1,28 @@ +# 00-GENESIS + +The **northern star**. What HAL is, the environment it runs in, and what changes when it +works. Every research effort and design decision is checked against this folder. + +| File | Purpose | +|------|---------| +| [`mission.md`](mission.md) | Vision, mission, and the values that decide arguments | +| [`context.md`](context.md) | The environment — conditions, not aspirations | +| [`effect.md`](effect.md) | What is different when the work is done | +| [`how-we-build.md`](how-we-build.md) | Rules that hold across the mesh, each one earned | + +## Rules + +- Markdown only. +- **Stable by nature.** Changes here reflect a genuine shift in intent, not iteration. +- Research and design must be traceable back to what is written here. + +## Note on `VISION.md` + +The repository root carries `VISION.md`, an architecture overview predating this folder. +It is a useful description of *how* the mesh works and should be folded into +[`02-DESIGN`](../02-DESIGN/), not here — GENESIS answers *why*. + +It has also drifted: it lists "Symlinks, not copies" as a key design principle, while the +operating rules forbid creating symlinks at all after one caused production data loss. +A founding document contradicting a hard rule is precisely the failure this folder exists +to prevent. diff --git a/00-GENESIS/context.md b/00-GENESIS/context.md new file mode 100644 index 0000000..fb9499f --- /dev/null +++ b/00-GENESIS/context.md @@ -0,0 +1,39 @@ +# Engineering Context + +The conditions the mesh is built for. Properties, not an inventory — no node here is +named, and nothing should be designed around a particular one existing. + +## Mandatory + +- **Nodes are heterogeneous.** Desktops, laptops and servers, with different hardware, + different operating systems and wildly different uptime. A design that assumes uniform + nodes does not survive contact. +- **Some nodes are mobile and frequently absent.** They sleep, change networks and lose + addressability. A node being unreachable is ordinary operation, never an incident. +- **At least one node must be stably addressable.** Central components — transport, + registry, artifact storage — can only live where they can always be reached. That is a + property some node must have, not an identity a particular node holds. +- **Human agents are few — often one — and usually asleep.** There is no team, no rota, no + second reviewer. Anything requiring a human to notice it will be noticed late. +- **Nodes are personal.** A human agent works on the same node the mesh runs on. The mesh + is a guest there and must not make a node worse to use. + +## Default + +- **Self-hosted throughout.** Transport, state, artifacts and memory run on nodes the mesh + owns, not a managed service. +- **A hosted model provider** supplies the thinking for non-human agents, drawn from a + shared pool of subscriptions — which is why budget pacing is a first-class concern. +- **Long-lived user services** rather than an orchestrator. No cluster scheduler, no cloud + control plane. + +Defaults, not mandates. A second model provider is anticipated by design; nothing in the +domain may assume one vendor's credential lifecycle. + +## Deviations + +- **No enterprise identity.** No directory, no SSO. Identity is mesh-internal. +- **Public exposure is minimal.** Only nodes that must terminate public traffic do so. +- **Agents share a pool of provider subscriptions** rather than holding billing + relationships of their own. A consequence of personal-scale infrastructure, and the + reason spend must be paced rather than merely billed. diff --git a/00-GENESIS/effect.md b/00-GENESIS/effect.md new file mode 100644 index 0000000..42c5e14 --- /dev/null +++ b/00-GENESIS/effect.md @@ -0,0 +1,43 @@ +# Effect + +Imagine the mesh works as intended. What is different? + +## You ask, and it happens + +An agent says what it wants — from a terminal, a phone, a message — and the mesh takes it +from there. It works out which nodes are involved, does the work, and returns a +result. Nobody opens a console, recalls which node holds what, or follows a runbook +written months ago. + +The interface is intent. The mesh handles the rest. + +## The nodes look after themselves + +Updates land, services recover, disks are kept clear, certificates renew, and the mesh +notices when something is wrong before you do. Maintenance stops being a thing you +schedule and becomes a thing that has already happened. + +When something genuinely needs a decision, you are asked — with the context, not a log +line. + +## Work continues while nobody is watching + +Agents keep working overnight and across the week. What they did is legible afterwards +because it is all in one record: what was asked, what was decided, what changed. The +overnight work can be trusted, which is what makes it worth doing at all. + +## The mesh remembers + +Nothing has to be explained twice. What was learned — how a thing works, why a decision +went the way it did, what broke last time — is available to whoever needs it next, +whether that is an agent working at 4am or a human agent six months later. + +## A human agent's environment is part of it + +The node a human agent sits at is not outside the mesh looking in. The desktop, the +notifications, the shell are how that agent acts — maintained by the mesh, exactly as a +spawned session is for an agent that is not human. + +## The difference for a human agent + +Less time spent operating the mesh. More time spent deciding what it should do. diff --git a/00-GENESIS/how-we-build.md b/00-GENESIS/how-we-build.md new file mode 100644 index 0000000..a9ebfae --- /dev/null +++ b/00-GENESIS/how-we-build.md @@ -0,0 +1,42 @@ +# How we build + +Working notes on the rules that hold across the mesh. Short, and each one earned. + +## Name a context after its aggregate, not after a metaphor + +`hal/agents` owns **Agent**. A *brain* — memory, thoughts, cognition — is something an +agent **has**, a concept inside the aggregate. It is not a module. + +The cost of getting this wrong is visible today: `hal/brain` names the node runtime, so +the most evocative word in the system points at infrastructure, and `hal/cortex` +describes itself as "mesh messaging" in its manifest while the anatomy documentation +calls it the interactive runtime — and it runs on no node at all. + +Anatomy makes attractive names and poor boundaries. Name the thing the domain calls it. + +## Ubiquitous language is checked, not assumed + +If a document states a rule about the mesh, say how the rule is verified. This repository +has a documented requirement that every module exposing tools declares `brain` as a +dependency. Zero modules do. An unenforced rule is indistinguishable from a wrong one, +and costs more, because people believe it. + +## Contexts integrate through the record, never through a shared schema + +Publish to the stream; do not join across a boundary. Today five domains share one +45-table schema, which is why work that belongs to one context keeps having to be +implemented in another. + +## A failed step must stop the steps after it + +Scripted work runs as a sequence, and a sequence that continues past a failure does the +next thing in the wrong place. Gate each step on the last: `cd X || exit`, not `cd X` +followed by a newline. + +Earned the obvious way. A `git worktree add` failed because the branch name collided with +an existing namespace; the `cd` into that worktree failed too; and the `cp`, `git add` and +`git commit` that followed ran in the shared checkout and committed to local `main`. The +error was printed and scrolled past. + +This is the same shape as the faults this refactor exists to remove — a step reported +failure, nothing stopped, and the damage happened somewhere nobody was looking. diff --git a/00-GENESIS/mission.md b/00-GENESIS/mission.md new file mode 100644 index 0000000..ab0113c --- /dev/null +++ b/00-GENESIS/mission.md @@ -0,0 +1,67 @@ +# Mission + +## Vision + +**A mesh that controls itself.** + +An agent states an intent — in words, from wherever they already are — and the mesh +carries it out. It takes the request in, works out what it means, does the work across +whichever nodes it needs, and returns a result. No console to open, no runbook to follow, +no remembering which node holds which thing. + +Not automation, which does what it was told to do in advance. Self-control: the mesh +holds the context, decides how, and acts. + +## Mission + +Build the layer that turns a set of nodes into one self-controlling mesh. + +- **Intake, process, deliver.** A request arrives, is understood, becomes work, and + returns an answer. That loop is the product; everything else exists to make it possible. +- **Agents inhabit the mesh.** They are not scripts that run and exit. They hold identity, + memory and skills, run on whichever node has room, and act continuously. +- **The mesh brokers everything the work needs.** Storage, credentials, compute, + knowledge, delivery — requested by capability, resolved by the mesh, never by the + requester knowing where things are. + +## Agents, some of whom are human + +There is one kind of participant: the **agent**. Some agents are human and some are not, +and the mesh does not treat that as a category difference. Both hold identity, both hold +credentials, both act, remember and coordinate. What differs is **modality** — how an +agent acts: + +- a non-human agent acts through a spawned session and the record +- a human agent acts through a shell, a desktop, a message from a phone + +That is why a desktop environment is as much a core concern as a knowledge store. One is +how some agents remember; the other is how some agents act. Neither is a courtesy +extended to a user outside the system. + +### What belongs in the mesh's own domain + +A **core** module supports an agent's *participation* — acting, remembering, +coordinating, or interfacing with the mesh. + +A media server supports a human, but not their participation. It is therefore not a core +module. It is still a perfectly valid HAL module — the +mesh installs it, provisions for it, brokers its capabilities and ships it through the +same pipeline. Entirely legitimate as a module, and no part of the mesh's own domain. + +The distinction is **core module** versus **module the mesh runs**, not module versus +not-a-module. Both use the same manifest, the same pipeline, the same provisioning — +which is exactly what makes the mesh's own components no more privileged than anything +else it carries. + +## Core values + +- **Evidence over assertion.** A claim that cannot be checked will quietly stop being + true. Say what was measured. +- **Failure must be loud.** The expensive faults are always the silent ones — work that + reported success and did nothing. Prefer failing to lying. +- **The mesh owns the truth.** State lives in the mesh and is derived onto nodes. A file + edited on a node is a bug with a delay on it. +- **Sovereignty.** The mesh runs on nodes it owns. External dependencies are + deliberate and few. +- **Dogfood everything.** The mesh's own components ship through the same machinery as + anything else it runs. If they need an exception, the machinery is not finished. diff --git a/01-RESEARCH/001-module-domain-decomposition/analysis.md b/01-RESEARCH/001-module-domain-decomposition/analysis.md new file mode 100644 index 0000000..aa8b1e9 --- /dev/null +++ b/01-RESEARCH/001-module-domain-decomposition/analysis.md @@ -0,0 +1,237 @@ +# Current state → ideal state + +## 1. The count is a symptom + +124 modules. The number is not the problem; the reason for it is. Modules are split +because splitting is the **only granularity lever the platform offers**: + +| symptom | module | why it exists | +|---|---|---| +| one daemon must run on one node | `hal/claude-licences` | no per-node feature opt-in | +| a library must not drag a 450 MB dep | `hal/claude` | no way to expose two npm packages | +| a domain needs a daemon *and* a library | `hal/claude-code` + `hal/claude` | one feature of each type per module | +| a handful of SDK verbs need a home | `infra` (10 lines) | tools must belong to *some* module | + +A feature is a **handler type**, discovered by directory presence, singleton per module +(`FEATURE_HANDLERS` in `modules/hal/sdk/src/feature-handlers/index.ts`). So "one more +deployable unit" always means "one more module". + +Splitting also fragments domains. Claude licence logic sits in three modules — +resolution in `hal/claude`, health in `hal/claude-code`, refresh in +`hal/claude-licences` — so following one token means reading three. + +## 2. Two things are conflated in one repo + +`noxflow` is 45 tables. Its original purpose — a HAL-native Jira — is 7 of them: + +| domain | tables | +|---|---:| +| tasks (the original concept) | 7 | +| agents / HR | 13 | +| knowledge (a Confluence) | 8 | +| meetings | 7 | +| scheduling / ops | 2 | +| other | 8 | + +Agents alone are nearly double the concept the module was built for. This is the direct +cause of concrete faults: per-agent Claude credentials had to be written by the noxflow +runtime purely because `agents.claude_account` lives in noxflow's database, while node +identity — the same kind of fact — lives in the mesh registry. + +Separately, 91 of the 124 modules are third-party software (`plex`, `sonarr`, +`firefox`, `postgres`). These **run on the mesh; they are not part of it**. They belong +outside this repository. The pipeline and provisioning system are deliberately generic, +so HAL's own modules dogfood exactly what external modules use — that property is what +makes the separation safe. + +## 3. Knowledge is split three ways + +| store | owner | size | search | +|---|---|---:|---| +| `mesh_docs` (+ revisions) | `hal/hippocampus`, in the mesh DB | 144 docs, 73 touched in 30d | yes | +| `knowledge_*` (8 tables) | `noxflow/runtime`, in the noxflow DB | 20 pages | yes | +| repo markdown | git | `VISION.md`, `FLAVOR-PROPOSAL.md`, `docs/` | **no** | + +The librarian is split across two of them: `librarian_file_document` writes via +hippocampus, `librarian_ask` and `librarian_retrieve` read via noxflow. Filing and +retrieval, different modules, different databases. + +Both DB stores have revisions; only noxflow has a review workflow (`_proposals`, +`_promotion_requests`). Neither has file access. Git has file access and history but no +search integration. **Every piece needed already exists — none of them together.** + +## 4. `hal/developer` and the missing local mesh + +`hal/developer` (≈2,600 lines) exposes `dev_stage`, `dev_diff`, `dev_validate`, +`dev_typecheck`, `dev_verify_flavor`, `dev_bootstrap`, `dev_build`, `dev_up`, `dev_down`, +`dev_env`. It is the tooling that both reflects and enforces the module model, so it +moves with any change to that model. + +`dev_up` builds a dev environment **for a single module**: resolves `requires:`, starts +providers "via systemctl (same as production)", provisions workspace-namespaced +databases. It borrows the host node's services. + +**There is no way to run a mesh locally.** Consequently anything that only manifests +*between* nodes — credential rotation reaching a running session, provision credentials +fanning out, cascade ordering, artifact packaging — is discoverable only in production. +Every fault fixed on 2026-08-22 was of that kind. + +A containerised mesh (N nodes, a broker, a registry, a coordinator) is therefore not a +convenience. It is the precondition for a refactor of this size being verifiable at all. + +## 5. `hal/sdk` is where domains go to hide + +155 files, 34,636 lines. It contains code from essentially every bounded context: + +| file | domain it belongs to | +|---|---| +| `installer-core.ts` (1,495) | runtime | +| `tools/provisions.ts` (1,311) | provisioning | +| `tools/deploy-service.ts` (1,081) | deploy | +| `workflow-engine.ts` (970), `task-commands.ts` (588) | **work** — noxflow's domain, inside HAL's SDK | +| `artifact-manager.ts` (948), `build-executor.ts` (760), `feature-handlers/*` | delivery | +| `env-generator.ts` (944) | config | +| `amqp-client.ts` (899) | comms | +| `meshware.ts` (799) | runtime | +| `module-registry.ts` (743) | mesh | +| `claude-credentials.ts` | **ai** | + +Nothing here was misplaced carelessly. Each landed in the SDK because the SDK is the one +package every module may import — so putting a shared type there is the way to avoid a +circular dependency between two modules that both need it. + +The result is that the dependency graph is trivially acyclic and completely +uninformative: everything depends on one 34k-line package, so a change anywhere in it +rebuilds everything. On 2026-08-22 a three-line fix in `feature-handlers/migrations.ts` +produced a 1,445-job, 26-minute mesh-wide cascade. + +The last row is from that same day, and is the clearest illustration: per-agent Claude +config-directory helpers were added to `@hal/sdk/claude-credentials.ts` because it was +the only place both `hal/claude` and the noxflow runtime could import from. The +alternative — `hal/ai` owning it and noxflow depending on `hal/ai` — was unavailable +without answering the ownership question this effort exists to answer. + +**Extraction is therefore blocked on the decomposition, not the reverse.** A domain +module can only own its shared code once the domain has a module. The SDK should retain +only what is genuinely cross-cutting: transport, manifest types, logging, process +helpers. + +## 6. The crux: two notions of "agent" + +The original vision was **a node is an agent with thinking abilities**. Later, noxflow +implemented actual agents — employees with skills, workspaces and tasks. Both survive, +and they collide. + +The collision is visible in the data. There are **two agent rows per node**: + +| agent | display name | skills | bound to | +|---|---|---|---| +| one named after each node | the node's name | `{deploy,verify,operate}` | its node | +| `hal-`, one per node | "HAL (\)" | `{}` | its node | + +One carries the **work**; the other carries only **identity** — no skills at all, existing +to hold a licence for HAL's own sessions (`resolveHalModuleEnv`). The node-as-agent idea +was implemented twice, half each. + +The same split appears in the schema: `nodes.hal_claude_account` and +`agents.claude_account` are the same fact in two tables in two databases. What was +documented as "three licence touchpoints" (NODE, HAL, AGENT) is **one concept modelled +three times**. + +### Resolution + +**Nodes and agents are decoupled.** The agents module owns agents; agents run on nodes. +A node is a place where an agent can run — that is the whole of the relationship. There +is no resident agent, no node-owned identity, no ownership in either direction. + +Agents named after a node still exist, but they are **ordinary agents defined through the +agents module** that happen to hold infra skills and be named after the node they usually run +on. Nothing about them is structural. + +Two consequences follow: + +1. **`nodes.node_license` and `nodes.hal_claude_account` should not exist.** A node does + not authenticate to a model provider — agents do. Both collapse into + `agents.claude_account`, and what was documented as three licence touchpoints becomes + one. +2. **The pair of agent rows per node merges.** `hal-` exists only to hold a licence + for HAL's own sessions; once licences belong to agents and any module may employ an + agent, it has no reason to be separate from ``. + +What remains unaccounted for is the operator's own interactive session: +`~/.claude/.credentials.json` on a node serves a **human at a prompt**, not an agent. +Under this decoupling that is not a node property either — it is the operator's +credential on whichever machine they are sitting at. Whether the operator is modelled as +a persona (soulstream treats humans and agents as peers with identical credentials) is an +open question, and the last remaining reason `nodes` carries a licence column today. + +## 7. The record, not "meetings" + +soulstream models work as **one signed log** in which humans and agents apply identical +changes, with work arriving by **mention**. Memory and decisions live in that single +auditable stream — no hidden state between components. Its stated economics, +*"cheap awareness and expensive thinking"*, is the same principle as the budget guard: +constant monitoring, selective deep processing. + +Under that model "meetings" is not a domain. A meeting is a **thread in the record**. +Which means six of today's modules and table groups are one thing seen from different +angles: + +`hal/axon` (DMs, rooms) · `hal/cortex` · `hal/synapse` · `hal/notifications` · +`conversations` · `meetings` + +A notification, in particular, is just a mention not yet read — which is why it needed +its own table, its own id column, and its own failure mode. + +## 8. Ideal state + +The mesh repository contains only the mesh, decomposed by domain rather than by +deployment accident: + +| context | what it is | absorbs today's | +|---|---|---| +| `hal/infra` | **the room** — nodes, registry, provisioning, pipeline, node runtime | `hal/mesh`, `hal/brain`, `hal/coordinator`, `hal/meshware`, `hal/developer`, `hal/bootstrap` | +| `hal/agents` | **the name and the thinking** — Agent owns identity, licence, runs, memory, thoughts | noxflow agents (13 tables), `hal/thoughts` | +| `hal/stream` | **the record** — mentions, threads, meetings, notifications | `hal/axon`, `hal/cortex`, `hal/synapse`, `hal/notifications`, meetings (7), conversations | +| `hal/knowledge` | documents, spaces, revisions, review, search | `hal/hippocampus` + noxflow `knowledge_*` | +| `hal/work` | tasks, workflows, boards — the original noxflow | noxflow tasks (7 tables) | +| `hal/ai` | provider integration, flavored per vendor | `hal/claude*`, `claude-code` | +| `hal/config` | env and config distribution, secrets, PKI | `hal/env-sync`, `hal/config-sync`, `hal/secrets`, `mesh-ca` | +| `hal/observability` | health, dashboards, logs | `hal/health`, `hal/meshboard` | + +Eight contexts instead of 33 platform modules, with the catalogue elsewhere. + +The naming corrects itself in the process: what is called "brain" today is +infrastructure, and the word survives only as a concept an Agent owns — never again as a +module name. + +This shape depends on **named features with per-node opt-in** — without it, every +independently-deployable unit inside a context becomes a module again and the count +returns. That mechanism is the subject of a separate ADR. + +## Decided since + +| question | decision | +|---|---| +| Meetings — own context or part of agents? | **Neither.** A meeting is a thread in the record; `hal/stream` absorbs it. The category was wrong. | +| Thoughts — autonomy or agent behaviour? | `hal/agents`. Any agent can have a thought. | +| `hal/axon` vs `hal/cortex` | `hal/cortex` is vestigial — no `daemon/` source, unit inactive on every node, description copy-pasted from axon. Messaging belongs to `hal/stream`. | +| `hal/hypothalamus` | Orphaned build output. A `dist/` directory with no `module.yml`, so not a module at all. | +| Who owns agent identity? | `hal/agents`. Nodes and agents are decoupled; a node is only a place an agent can run. | +| Where do third-party modules live? | Outside this repository. They run *on* the mesh, not *of* it. | + +## Open questions + +1. **`hal/scheduler`** — infrastructure (fire an event later, belongs to `hal/mesh`) or + part of `hal/work`? +2. **The executor.** Pulling and routing work items is `hal/work`; spawning a session is + `hal/agents`; providing the place to run is `hal/mesh`. It currently straddles all + three, which is the same kind of straddle that put credential-writing in the noxflow + runtime. +3. **Catalogue destination** — one repository, or per-application repositories, given the + pipeline resolves dependencies across the registry rather than the filesystem? +4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or + does that belong to `hal/stream`? Everything imports it, which argues both ways. +5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which + user, on which node, a human agent acts as. Where does it live — an attribute of the + agent, or of the agent-node binding? diff --git a/01-RESEARCH/001-module-domain-decomposition/status.md b/01-RESEARCH/001-module-domain-decomposition/status.md new file mode 100644 index 0000000..87597bd --- /dev/null +++ b/01-RESEARCH/001-module-domain-decomposition/status.md @@ -0,0 +1,46 @@ +# 001 — Module domain decomposition + +- **Status:** ONGOING — ADR 0001 accepted; graduates when `02-DESIGN` carries the per-context specifications +- **Initiated by:** jochen, 2026-08-22 +- **Areas touched:** every `hal/*` and `noxflow/*` module; the pipeline's dependency + graph; the knowledge base; agent identity and credentials. + +## Summary + +HAL has 124 modules. That number is not a maintenance problem in itself — it is the +**symptom of missing bounded contexts**. Modules are split not because they model +different domains, but because splitting is the only lever the platform offers: + +- no way to run one daemon on one node without making it a module + (`hal/claude-licences` — one daemon, single-node) +- no way to expose two of a kind from one module +- no namespace separating the mesh from the software it runs + +This effort establishes the **current state**, the **ideal state**, and the sequence +between them. + +## Trigger + +A night of debugging that produced four fixes and one conclusion. Every fault was a +boundary fault: + +- Per-agent Claude credentials had to be written by the *noxflow runtime*, because + `agents.claude_account` is in noxflow's database — even though agent identity is a + mesh concept and node identity already lives in the mesh registry. +- Whether `hal/brain` may depend on noxflow took three attempts to answer, twice + wrongly, because the ownership boundary was never stated. +- Authoritative documentation existed in `mesh_docs` and was not found, while a + proposal in repo markdown was invisible to search entirely. + +## Decisions taken (2026-08-22) + +| Question | Decision | +|---|---| +| What should noxflow become? | Decompose into `hal/*` modules; noxflow returns to tasks/workflows | +| Who owns agent identity? | `hal/agents` — a mesh concept, alongside nodes | +| Where do third-party apps live? | Out of this repo. They run *on* the mesh; they are not *of* it | +| Knowledge structure | Modelled on `papa-hq`; implementation choice left open | + +## Open questions + +Tracked in [`analysis.md`](analysis.md) under "Open questions". diff --git a/01-RESEARCH/002-local-mesh/analysis.md b/01-RESEARCH/002-local-mesh/analysis.md new file mode 100644 index 0000000..95fc1b9 --- /dev/null +++ b/01-RESEARCH/002-local-mesh/analysis.md @@ -0,0 +1,258 @@ +# A mesh that runs locally — current state and obstacles + +Evidence for Phase 0. Every claim here is either a file location or something measured on +2026-08-22; where a claim was checked and found false, that is recorded too. + +--- + +## 1. What exists today, and why neither is a mesh + +### `test/pipeline/` — the only containerised HAL node, and it is dead + +This harness builds `@hal/sdk` and the `hal/brain` daemon into a `node:23-alpine` image, +compiles `hal/coordinator`'s tools against it, and runs `brainstem.js` with +`HAL_MODE=daemon` against a containerised postgres and LavinMQ +(`test/pipeline/Dockerfile`, `test/pipeline/docker-compose.yml`). It proves the valuable +thing: **a node runtime needs no systemd, no `/services/`, and no nvm** — environment +variables and a seeded `nodes` row are sufficient. + +It also cannot build. Measured 2026-08-22: + +``` +Step 7/16 : RUN npm run build -w modules/hal/sdk +npm error No workspaces found: +npm error --workspace=modules/hal/sdk +``` + +The npm workspace it depends on was removed on 2026-06-04 (`21ef4a4e`, "kill npm +workspace"), and the root `package.json` now carries a comment explaining why it will not +come back. The harness has not been touched since before that commit. + +So the repository's only end-to-end pipeline test has been unrunnable for two and a half +months and nothing reported it — which is the same shape as the faults Phase 0 exists to +catch. **A test nobody runs is indistinguishable from a test that passes.** + +### `test/dev-mesh/` — a work plane, not a control plane + +Four `runtime-*` services, a noxflow API, a meshboard and a shared postgres + LavinMQ on +one network (`test/dev-mesh/docker-compose.yml`). It is genuinely useful and the topology +is a good precedent: one network, per-service `HAL_NODE`, `EXECUTOR_STUB=1` so dispatch and +run lifecycle are real while the model call is faked. + +It is the wrong layer for Phase 0, on three counts: + +| | dev-mesh does | Phase 0 needs | +|---|---|---| +| state | restores `pg_dump`s (`test/dev-mesh/init.sh`) | schema built by **migrations**, or fixture 3 is untestable by construction | +| code | bind-mounts pre-built `dist/` from the host | artifacts **built and delivered** by the pipeline | +| scope | noxflow runtime, API, meshboard | `hal/coordinator`, `hal/meshware`, MinIO, provisioning | + +No meshware, no coordinator, no artifact store, no provisioning, no build. It exercises +what the mesh *runs*; Phase 0 needs what the mesh *is*. + +### `dev_up` — confirmed host-coupled + +The claim in the work breakdown holds. `startDevProvider()` starts providers through the +host's user systemd instance and reads credentials from a host path +(`modules/hal/developer/tools/dev-env.ts:146-178`): + +```ts +const unitName = `hal-module@${provider}.service`; +execSync(`systemctl --user start ${unitName}`, { timeout: 60_000, stdio: "pipe" }); +const creds = resolveRunningProviderCreds(provider); // reads /services//.env +``` + +It borrows the host. There is no mesh to stand up. + +--- + +## 2. What is already portable + +- **Node identity is one environment variable.** `resolveNodeName()` is + `process.env.HAL_NODE || hostname()`, lowercased (`modules/hal/sdk/src/mesh-config.ts:215`). + No file, no registration handshake, no host coupling. +- **Registration is a plain upsert.** `mesh_node_register` inserts into `nodes` + (`modules/hal/mesh/tools/index.ts:490-540`); mandatory fields are `name` and `user_name`. +- **The daemons are AMQP loops.** `hal/meshware` consumes `cmd.feature.{build,install,configure,start}` + (`modules/hal/meshware/daemon/src/cerebellum.ts:780-857`); `hal/coordinator` consumes + `event.gitea.push` and drives the cascade. Both are configured entirely by environment. +- **`hal/brain` in daemon mode already runs in a plain node image**, per `test/pipeline/`. + +--- + +## 3. The obstacles, located + +Four couplings stand between the daemons and a container. All are in code, none are +mysterious. + +| # | coupling | location | +|---|---|---| +| 1 | meshware restarts itself via the host: `execFileSync("systemctl", ["--user","restart","hal-meshware.service"])` | `modules/hal/meshware/daemon/src/cerebellum.ts:826` | +| 2 | meshware restarts the broker via a hardcoded host path: `execFileSync("docker",["compose","-f","/services/lavinmq/docker-compose.yml","restart"])` | `cerebellum.ts:843-844` | +| 3 | **a module service *is* a systemd unit** — `hal-module@.service` runs `docker compose --project-directory /services/%i` | `modules/hal/meshware/systemd/hal-module@.service:10-11` | +| 4 | env generation writes `homedir()`-relative host paths — `~/.config/hal/env`, `~/.config/hal/modules/*.env`, `/services/*/.env` | `modules/hal/sdk/src/env-generator.ts:188-223`, `:538-568`, `:584-606` | + +1, 2 and 4 are small: a guard and a configurable root. **3 is the architectural one**, and +it is the first open question below — in a container, "systemd unit that runs docker +compose" has no natural translation. + +The bootstrap scripts add their own host assumptions — nvm for node resolution, `sudo +mkdir -p /services && chown`, and a `~/dotfiles` repository that is not in this repo +(`install.d/adopt.sh:165-170`) — but Phase 0 does not have to run them. A container image +can be built the way `test/pipeline/Dockerfile` builds one, bypassing the bootstrap path +entirely. That is a deliberate divergence to record: **the local mesh would not test the +bootstrap**, only the running mesh. + +--- + +## 4. External services, and whether they can be local + +| service | used for | local substitute | +|---|---|---| +| PostgreSQL | registry + pipeline state | yes — already the pattern in both harnesses | +| LavinMQ | all coordinator↔meshware messaging | yes — `cloudamqp/lavinmq`, already used | +| MinIO | per-feature artifact tarballs, `modules/{name}/{version}/{feature}.tar.gz` (`modules/hal/sdk/src/artifact-manager.ts:21,65`) | yes — `minio/minio`, needs only `REGISTRY_MINIO_*` | +| Gitea | source, push webhook, **and** the `@hal/*` npm registry | container exists, but heavy; see open question 2 | +| Docker registry | `docker push` for modules declaring `docker:` (`build-executor.ts:195-231`) | `registry:2`, or avoid modules that need it | +| Traefik | writes routing config at install; does not gate success | omit | + +Only Gitea is awkward, and only because it carries two roles at once. + +--- + +## 5. The fixtures + +Phase 0.5 asks for at least one reproducible known fault. All three are reachable; they +differ sharply in cost. + +### Fixture B — a provider deploy rotates a shared credential without fanning out + +**Cheapest, best understood, and the root cause is still open.** Documented at +`troubleshooting/provision-adoption-rotates-live-credential`. The mechanism is two lines: + +```ts +async provision(project, username, options) { + const password = generatePassword(); // ALWAYS a fresh password +``` + +and, in the adoption branch of `provisionDatabase`, an `ALTER ROLE ... WITH PASSWORD` that +rotates the live secret while updating only the provider's own `mesh_provisions` row +(`modules/postgres/tools/index.ts`). Every consumer sharing that role keeps a stale +password and fails permanently; the provider recovers alone, and that asymmetry is the +tell. + +Reproduction needs one provider node and two consumer nodes — the minimum interesting +mesh. It re-fires on every provider deploy, so it does not need to be provoked, only +observed. Confirmation is a timestamp comparison, not a hash: the stored values are +`enc:v1:` with a random IV, so identical plaintexts hash differently. + +Measured in production 2026-08-22: three nodes failing since 2026-08-20 14:15, 399 errors +each; the provider failed for 22 minutes and recovered by itself. + +### Fixture C — a migration ships nothing while the pipeline reports success + +Three independent silent-skip points, any one of which produces it: + +1. **Not-applicable and succeeded are the same status.** `MigrationsHandler.detect()` is + `existsSync(join(moduleDir, "migrations"))` against the freshly cloned workspace + (`modules/hal/sdk/src/feature-handlers/migrations.ts:32-41`). If the directory is not + there — uncommitted, ignored, or a `module_path` that does not line up — the build + reports `success` (`cerebellum.ts:634-638`). +2. **Upload failure is a warning.** `uploadFeatureArtifact()` returns `{sha256: ""}` with a + `warn` when none of the listed files exist (`artifact-manager.ts:41-44`), and callers + catch and warn rather than throw (`cerebellum.ts:672-673`). The symmetric download + failure on the target node is also non-fatal (`artifact-manager.ts:93-98`, + `cerebellum.ts:426-431`, commented "Non-fatal: feature may work without artifact"). +3. **Missing packaged files are skipped in silence.** `addFileEntries()` does + `if (!existsSync(abs)) continue;` (`module-builder-core.ts:134-135`) — an explicit + `package:` entry that is not on disk simply is not in the tarball. + +This is the same class as `troubleshooting/flavors-never-packaged` (`flavors/` was never +staged, so a node kept its first copy forever) and is catalogued in +`troubleshooting/deploy-reports-transport-not-effect`, which counted **0 of 119 modules +implementing `verify`** — the stage that exists, is dispatched, has a working handler, and +would turn every one of these into a red pipeline. + +### Fixture A — a credential rotation does not reach a running session + +The most valuable and the most work: it needs a *running session* to rotate underneath, +which means the local mesh must be able to start one. The mechanism is understood — an +`EnvironmentFile` is read once and `process.env` is a snapshot — and it is why credentials +became files that get re-read. Defer it behind B and C. + +**Recommendation:** B first. It needs three nodes and a provision, no build and no session, +and it is the one whose root cause is still open — so reproducing it locally has value +beyond proving the harness. + +--- + +## 6. A claim that was checked and found false + +The survey behind this document suspected that the live AMQP pipeline never writes +`deployments` or `node_modules.installed_version`, because the code that does so sits in +`installer-core.ts:1045-1126` on what is commented as the "CLI path". + +Measured against production, 2026-08-22 17:30 UTC — deployments in the preceding 24 hours: + +| node | deploys | newest | +|---|---|---| +| 1 | 24 | 17:26:17 | +| 2 | 27 | 17:22:22 | +| 3 | 47 | 17:25:54 | +| 4 | 26 | 17:22:23 | + +The record is live on all four nodes and minutes old. **The claim is refuted.** Recorded +because the reasoning was plausible and someone will retrace it. + +--- + +## 7. Questions + +### Settled (jochen, 2026-08-22) + +**Phase 0 is a development environment, not a fixture rig.** It is built as a supported +surface to work in daily, and it is what `dev_up` should become. The definition of done — +"demonstrated in the local mesh" — is therefore meant literally. + +**The trigger is a real Gitea container.** The webhook relay is part of what is under test, +including the changed-file enrichment that exists because the webhook truncates at 20 +commits (`modules/hal/gitea/tools/index.ts:136-195`). A synthetic `event.gitea.push` would +skip it. Gitea also carries the `@hal/*` npm registry, so the local mesh needs it twice +over either way. + +### Open — blocking + +1. **How does a node supervise a module service?** Today a module service is a systemd unit + running `docker compose` against `/services/` + (`modules/hal/meshware/systemd/hal-module@.service:10-11`), which has no direct + translation inside a container. The question raised in response is the better one, and + is broader than Phase 0: **what would it cost to stop using systemd altogether and have + the mesh supervise its own services?** That is a mesh-level architectural question, not + a local-mesh implementation detail — if the answer is that the mesh should own + supervision, Phase 0 should not build a container-only workaround first. Under + investigation; findings will land in [`003-service-supervision`](../003-service-supervision/) + and, if it goes ahead, ADR 0002. + +2. **What replaces `~/.config/hal/env` and `/services/` inside a container?** A configurable + root keeps one code path; container-specific targets keep the host paths untouched. This + decides whether env generation gets a seam or a conditional. + +3. **How faithful must the local mesh be to be trusted?** It will not run the bootstrap + scripts, and it will run one OS where the real mesh is heterogeneous by design + (`00-GENESIS/context.md`). Stating the divergence up front is what stops "it works + locally" from becoming its own class of silent failure. Now sharper, because a + development environment people use daily is trusted far more than a rig — and drifting + from production costs correspondingly more. + +--- + +## References + +- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this + phase unblocks +- [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) — Phase 0 tasks + and checkpoint +- `troubleshooting/provision-adoption-rotates-live-credential` — fixture B, root cause open +- `troubleshooting/deploy-reports-transport-not-effect` — the nine defects that shipped + green, and the unused `verify` stage +- `troubleshooting/flavors-never-packaged` — fixture C, previously seen in the wild diff --git a/01-RESEARCH/002-local-mesh/status.md b/01-RESEARCH/002-local-mesh/status.md new file mode 100644 index 0000000..b463852 --- /dev/null +++ b/01-RESEARCH/002-local-mesh/status.md @@ -0,0 +1,53 @@ +# 002 — A mesh that runs locally + +- **Status:** GRADUATED — the design is [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md) +- **Initiated by:** jochen, 2026-08-22 +- **Areas touched:** `install.d/`, `hal/meshware`, `hal/coordinator`, `hal/brain`, + `hal/developer` (`dev_up`), `hal/sdk` (env generation, feature handlers, artifact + manager), `test/pipeline/`, `test/dev-mesh/`, the provisioning path in + `modules/postgres/`. + +## Summary + +Phase 0 of [`02-DESIGN/00-work-breakdown.md`](../../02-DESIGN/00-work-breakdown.md) requires +a mesh that comes up in containers, runs its own pipeline, and reproduces known faults on +demand. Nothing else in the decomposition starts until it exists, because every fault the +decomposition addresses was found in production — there was nowhere else to find it. + +This effort establishes what already runs in a container, what is welded to the host, and +what it would take to close the gap. It does **not** choose an approach: the central +question — how a containerised node executes a module service, when a module service is +defined today as a systemd unit shelling to `docker compose` in `/services/` — is not +answered by ADR 0001 and is recorded below rather than decided. + +## What was established + +- The two existing container harnesses are **neither of them a mesh**, and one of them has + not been able to build since 2026-06-04. +- Node identity is already portable — a single environment variable, no host handshake. +- Four concrete host couplings block a containerised node, all with known locations. +- All three Phase 0 fixtures are reproducible; one of them is documented in the knowledge + base with an open root cause and is the cheapest place to start. + +Detail and evidence in [`analysis.md`](analysis.md). + +## Questions — all settled 2026-08-22 + +**Phase 0 is a development environment**, not a fixture rig — it is what the host-borrowing +dev tooling becomes. + +**The trigger is a real source-forge container**, because the webhook relay is part of what +is under test. + +**A node is a system container**, promotable to a virtual machine per node. This answered +the question the effort was stuck on, and dissolved it rather than solving it: against a +real node with a real init, the four host couplings catalogued in `analysis.md` §3 are not +couplings — they are how a node works. They were obstacles only to a node modelled as an +application container. + +Consequently [`003-service-supervision`](../003-service-supervision/) **no longer blocks +Phase 0**. It remains a live architecture question, on its own timeline. + +The remaining questions in `analysis.md` — what replaces host paths in a container, and how +faithful the lab must be — are answered in the design: nothing replaces them, because the +paths are real; and the divergences are enumerated rather than discovered. diff --git a/01-RESEARCH/003-service-supervision/analysis.md b/01-RESEARCH/003-service-supervision/analysis.md new file mode 100644 index 0000000..4c52bfe --- /dev/null +++ b/01-RESEARCH/003-service-supervision/analysis.md @@ -0,0 +1,225 @@ +# Who supervises a service — the cost of leaving systemd + +Measured 2026-08-22. Every claim is a file location or a count. + +--- + +## 1. What systemd actually does for HAL + +**14 modules** ship a `systemd/` directory. The units divide cleanly into three kinds: + +| kind | count | what it is | +|---|---|---| +| long-running daemons | 9 | `hal/brain`, `hal/meshware`, `hal/coordinator`, `hal/cortex`, `hal/env-sync`, `hal/file-share`, `hal/thoughts`, `noxflow/runtime`, `desktop_notifications` — all `Restart=always`, `RestartSec=10` | +| scheduled one-shots | 6 timers | docker-prune, docker-registry maintenance, claude-code sessions + usage, health, mailu cert-sync | +| the template | 1 | `hal-module@.service` — the per-module Docker lifecycle | + +Only `mailu` is system-scope. Everything else is a user unit under +`~/.config/systemd/user`, auto-detected from directory presence — the module author names +the file and that filename *is* the unit name +(`modules/hal/sdk/src/feature-handlers/systemd.ts:31-42`). + +The template is the piece `002` tripped over +(`modules/hal/meshware/systemd/hal-module@.service`): + +``` +ExecStart=docker compose --project-directory /services/%i --project-name %i up --remove-orphans --pull always +ExecStop=docker compose --project-directory /services/%i --project-name %i down +Restart=on-failure +StartLimitIntervalSec=120 +StartLimitBurst=5 +``` + +Systemd's roles, then: restart-on-failure, start at login, env-file loading, ordering, +per-module lifecycle, timers, and — via journald — **the only log store the 9 Node daemons +have** (`modules/hal/sdk/src/tools/log-tail.ts:50-53`). + +--- + +## 2. The question splits in two, and the halves disagree + +### For `docker compose` stacks, systemd is mostly redundant + +**44 of 44** module `docker-compose.yml` files declare a container-level `restart:` policy — +`always` or `unless-stopped`, none missing. Docker's own daemon already restarts crashed +containers, independently of systemd. + +`hal-module@.service` does not supervise the containers. It supervises the **`docker compose +up` foreground process**. It is a second layer on top of a restart policy that already +works. What it genuinely adds is narrower than it looks: + +- one uniform verb for every module type — `systemctl --user start hal-module@X` rather than + remembering each module's compose invocation, relied on across a dozen call sites in + `meshware.ts:124-154`, `dev-env.ts:142-163`, `installer-core.ts`, `infra.ts:166` +- recovery when the `docker compose up --pull always` process itself dies — a failed image + pull, not a crashed container +- rate-limited restart (`StartLimitBurst=5`) so a broken stack does not spin + +That is real, but it is a convenience layer, not a safety layer. **This half could go at +moderate cost.** + +### For the 9 Node daemons, systemd is load-bearing + +There is no alternative supervisor anywhere in the repo. No PM2, no forever, no nodemon, no +watchdog loop — all checked, zero hits. `Restart=always` is the only thing standing between +a crashed daemon and a dead node. + +**This half is the actual question.** + +--- + +## 3. The hard part is fate-sharing + +The difficulty is not systemd. It is that **a supervisor must not share fate with what it +supervises**, and the codebase already has a scar from exactly this. + +meshware cannot restart itself mid-request: killing the process before it ACKs the AMQP +message loses the message. So it defers its own restart by two seconds after closing the +connection (`modules/hal/meshware/daemon/src/cerebellum.ts:815-828`) — a commented +workaround for a problem that only exists because the thing being restarted is the thing +doing the restarting. + +Any mesh-native supervisor inherits this recursively. Something has to be the outermost +always-alive process, and if it is written by the mesh, the mesh must supervise it, and so +on. The recursion only terminates at a process the mesh does not own. Today that is +systemd. **A "more mesh" supervisor that is itself a mesh process is not a smaller problem; +it is the same problem with a new name.** + +This is the strongest argument for the status quo, and it is worth stating plainly before +looking at alternatives. + +--- + +## 4. The option neither of us named + +There is a third answer that terminates the recursion in something that is not systemd and +not written by us: **run HAL's own daemons as containers.** + +Docker is already the outermost supervisor for 44 of 44 module stacks. It does not share +fate with the mesh. It already has restart policies, backoff, and a log store that +`log_tail` already speaks (`log-tail.ts:43-47` reads `docker logs` for Docker modules +today). Extending it from "the things the mesh runs" to "the mesh itself" is not new +machinery — it is applying machinery the mesh already trusts to one more case. + +What that buys, beyond supervision: + +- **Phase 0 stops being a translation.** A containerised node becomes the same shape as a + production node, rather than a local approximation with a systemd-shaped hole in it. The + divergence that `002` open question 3 worries about largely disappears. +- **`/services/` and `~/.config/hal/` stop being special.** Mounts, not host paths. +- **The GENESIS "dogfood everything" value gets easier**, not harder: the mesh's own + components would ship and run exactly like everything else it carries. + +What it costs, honestly: + +- **journald → docker logs** for the 9 daemons. `log_tail` already handles both, but + `systemd_journal` and the health checks that read unit state + (`modules/hal/mesh/health.sh:34-40`, `modules/hal/health/hal-health.sh:335-395`) would + need a container-aware path. +- **Boot start** becomes Docker's `restart: always` plus the Docker daemon being enabled at + boot — which is still one systemd unit, but the OS's own, not ours. +- **A container needs the host to be reachable** for anything that touches the node itself. + Some of these daemons exist precisely to write host files. + +--- + +## 5. Why it cannot be all-or-nothing — and ADR 0001 already says so + +Some of what runs under systemd today **cannot** be containerised, and the reason is +already in the domain model. ADR 0001: + +> a non-human agent acts through a spawned session — a human agent acts through a shell or +> desktop + +`hal/brain` has two modes (`modules/hal/brain/daemon/src/brainstem.ts:7-13`): a daemon mode +that is an AMQP relay, and a **cortex mode that is an MCP server over stdio for an +interactive Claude Code session**. The second is a human agent's modality. It runs in the +human's shell, on the human's node, against the human's `~/.claude`. Containerising it is +not a hard engineering problem, it is a category error. + +The same holds for `desktop_notifications` and everything `hal/desktop-environment` touches. + +So the line is not "systemd or not". It is: + +| | belongs where | +|---|---| +| mesh daemons — meshware, coordinator, env-sync, thoughts, file-share, brain **in daemon mode** | supervisable by Docker; candidates to containerise | +| human-modality surfaces — brain **in cortex mode**, desktop notifications, desktop environment | on the host, by definition | +| module stacks | already Docker; systemd layer is the redundant part | + +This split is not a compromise between the options. It is what the domain model implies, +and it is a decent sign that the model is doing work. + +--- + +## 6. Options, with costs + +| | option | cost | what it buys | +|---|---|---|---| +| **A** | **Keep systemd; systemd-in-container for Phase 0** | privileged containers, heavy images, slow iteration; local mesh keeps a shape production does not have | nothing changes in production; smallest change to the mesh | +| **B** | **Drop the `hal-module@` layer only** — let Docker's restart policies supervise stacks directly | reimplement uniform start/stop across ~12 call sites; lose rate-limited restart and pull-failure recovery | removes the redundant layer; does **not** solve Phase 0 on its own, since the daemons still need supervising | +| **C** | **Containerise the mesh daemons; Docker supervises** | container-aware `systemd_journal`/health; host access for daemons that write host files; the human-modality surfaces stay on the host regardless | terminates the fate-sharing recursion without writing a supervisor; makes local and production the same shape; Phase 0 becomes much less of a special case | +| **D** | **Write a mesh-native supervisor** | the fate-sharing recursion (§3), plus matching systemd's maturity — backoff, resource limits, clean SIGTERM (`noxflow/runtime` already depends on `TimeoutStopSec=60`) — on machines that are somebody's daily driver | most "mesh"; least justified by the evidence | + +**D is the option the phrasing "more hal mesh approach" points at, and the evidence argues +against it.** Supervision is not a domain concern the mesh is better placed to solve than +the OS; the mesh's distinguishing feature is brokering capabilities, not restarting +processes. C gets the benefit D is reaching for — the mesh not depending on host-specific +init — without the recursion. + +**C and B compose.** C is the one that pays for Phase 0. + +Worth noting what is *not* in this table: the dozens of `systemctl` call sites that +configure the **host OS's own** units — NetworkManager, resolved, oomd, zram, docker.service, +fail2ban, sshd, zfs, across `modules/asusd`, `g14-power`, `wireguard`, `sshd`, `zfs` and +others. Those are not HAL supervising itself; they are HAL configuring an Arch box. They +are out of scope for every option above and do not go away under any of them. + +--- + +## 7. An incidental finding + +Documentation describes an automatic node rescue: `hal-rescue.sh:20` states it is "triggered +automatically by `hal-health.timer` when hal-meshware is failed". + +**It is not.** `hal-health.sh` contains no call to `hal-rescue.sh` (checked, zero matches), +and **no unit in the repository declares `OnFailure=`** (checked, zero matches). The only +real triggers are the manual `rescue_node` tool +(`modules/hal/sdk/src/tools/deploy-rescue.ts:77-92`) and running `install.d/rescue.sh` by +hand. + +This is worth recording for two reasons. It weakens any argument that systemd-adjacent +self-healing is already wired — it is not. And it is another instance of the pattern this +whole refactor is about: **a documented mechanism that does not exist, believed because it +was written down.** `00-GENESIS/how-we-build.md` calls this out as a rule; here it is again, +found by grep. + +--- + +## 8. Open question + +**Which supervision model does the mesh adopt?** A, B, C, D or a combination. + +**This no longer gates Phase 0.** When this was written, the local mesh was assumed to be +built from application containers, which forced the question — there is no natural way to +run an init system inside one. The decision of 2026-08-22 to build development nodes as +**system containers** (see [`02-DESIGN/01-end-to-end-testing.md`](../../02-DESIGN/01-end-to-end-testing.md)) +removes that pressure entirely: a system container runs a real init, so the existing model +works unmodified and the lab needs no answer here to exist. + +What remains is the question on its own merits, which is worth keeping open because the +evidence above still holds: option B removes a layer that 44 of 44 module stacks have already +made redundant, and option C would let the mesh stop depending on host-specific init. Neither +is urgent. Both are now cheap to *try*, because there is somewhere to try them. + +Option D — a mesh-written supervisor — remains the one the evidence argues against, for the +fate-sharing reason in §3. + +## References + +- [`002-local-mesh`](../002-local-mesh/analysis.md) — the effort this came out of +- [`adr/0001`](../../adr/0001-mesh-brokers-nodes-host-agents-think.md) — agent modality, which + decides what cannot leave the host +- `modules/hal/meshware/daemon/src/cerebellum.ts:815-828` — the self-restart workaround +- `modules/hal/meshware/systemd/hal-module@.service` — the per-module Docker lifecycle +- `modules/hal/sdk/src/feature-handlers/systemd.ts` — detection, install, start diff --git a/01-RESEARCH/003-service-supervision/status.md b/01-RESEARCH/003-service-supervision/status.md new file mode 100644 index 0000000..7ab178a --- /dev/null +++ b/01-RESEARCH/003-service-supervision/status.md @@ -0,0 +1,47 @@ +# 003 — Who supervises a service + +- **Status:** ONGOING — evidence gathered, options costed, decision open. + **No longer blocks Phase 0** (see below). +- **Initiated by:** jochen, 2026-08-22, in response to + [`002-local-mesh`](../002-local-mesh/analysis.md) open question 1 +- **Areas touched:** every module shipping a `systemd/` directory (14), the + `hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` / + `systemd_journal`, the bootstrap scripts. + +## The question + +`002` asked how a containerised node runs a module service, given that a module service is +defined today as a systemd unit running `docker compose` against `/services/`. The response +was the better question: + +> If it's possible to run systemd inside a container, that's the way to go I think. However, +> what would the cost be to step away from systemd to run our services and set it up in a +> different way? More hal mesh approach. + +This effort answers the cost half. It does not choose. + +## Summary of findings + +- **The question splits in two**, and the halves have opposite answers. Supervising + `docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose + files already declare a restart policy, so Docker is already the supervisor. Supervising + HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the + only thing between a crash and a dead node. +- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and + carries a documented workaround for it. Any mesh-native supervisor inherits that problem + recursively unless it sits outside the mesh's own process tree — at which point it is an + OS-level supervisor again, just reinvented. +- **There is a third option neither of us named**, and it is the one that also solves Phase 0: + run HAL's own daemons as containers, making Docker the supervisor for everything. Local and + production then have the same shape rather than a translation layer between them. +- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a + shell and a desktop. Those parts are on the host by definition. +- One incidental finding: the automatic node rescue that documentation describes **does not + exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer. + +Detail and costs in [`analysis.md`](analysis.md). + +## Decision needed + +Which supervision model the mesh adopts, recorded in ADR 0002 before Phase 0 builds +anything. The options and their costs are in `analysis.md` under "Options". diff --git a/01-RESEARCH/004-lab-network/analysis.md b/01-RESEARCH/004-lab-network/analysis.md new file mode 100644 index 0000000..fe601f9 --- /dev/null +++ b/01-RESEARCH/004-lab-network/analysis.md @@ -0,0 +1,198 @@ +# Reproducing the mesh network in a lab + +Established 2026-08-22 by reading the generating code and the live mesh DB. Every claim is a +file location or a queried row. + +--- + +## 1. The network is data, not configuration + +`install.d/mesh-init.sh` and `install.d/adopt.sh` perform **no network configuration at +all** — no WireGuard, no DNS, no firewall, no `/etc/hosts`. Every part of the network layer is +generated by module hooks from mesh-DB rows: + +| layer | generated by | from | +|---|---|---| +| WireGuard interface + peers | `modules/wireguard/hooks/index.ts` (`postConfigure`) | `node_wg_keys`, `nodes.site`, `nodes.underlay_addr`, `module_env.WG_ADDRESS` | +| `.internal` name resolution | `modules/dnsmasq-app/hooks/index.ts` | mesh config peers → `internal_domain` + `wg_address` | +| public routing / vhosts | `modules/hal/sdk/src/vhost-gen.ts`, `feature-handlers/vhost.ts` | `vhosts:` manifests + `node_accessors` | +| internal TLS | `modules/mesh-ca/hooks/index.ts` | a singleton CA row in the mesh DB | + +**Consequence:** a faithful lab is mostly a matter of writing the right rows. The network that +results is produced by the same code production runs, which is the difference between testing +the network and testing a model of it. + +--- + +## 2. The constraint that decides whether the lab works + +`modules/wireguard/hooks/index.ts:225-240` decides, per pair, whether to write an `Endpoint`: + +```js +const isPrivate = (a) => /^(10\.|127\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.)/.test(a); + +if (coLocated && underlay) Endpoint = `${underlay}:${port}` // same LAN +else if (underlay && !isPrivate(underlay)) Endpoint = `${underlay}:${port}` // public +// else: no Endpoint — the peer must initiate, and we learn its endpoint from the handshake +``` + +A simulated public segment addressed out of RFC1918 space — `10.200.0.0/24`, say — makes the +hub's underlay test as **private**. No spoke writes an `Endpoint` for the hub. Nothing can +initiate. **No handshake ever occurs and the mesh silently never forms**, presenting as a +WireGuard fault rather than an addressing choice. + +**The simulated public segment must therefore be `203.0.113.0/24`** — TEST-NET-3, reserved by +RFC 5737 for documentation, guaranteed never to route on the real internet, and not matched by +that regex. The code then treats it exactly as it treats a real hosting provider address. + +This is the single most important fact in this document. + +--- + +## 3. The topology being reproduced + +The shape below is what a mesh of this kind looks like: one node with a routable address, one +publicly named but behind a household NAT, one stationary workstation, one that roams. +Addresses use the documentation ranges of RFC 5737 and RFC 1918 throughout. + +| node | profile | site | underlay | WG | accessors | +|---|---|---|---|---|---| +| `anchor` | server | `dc` | `203.0.113.10` (routable) | `10.10.0.1/24` | `anchor.example` (public, primary) + `anchor.internal` (lan) | +| `home-server` | server | `home` | `192.168.1.135` | `10.10.0.2/24` | `home-server.example` (public, primary) + `home-server.internal` (lan) | +| `workstation` | workstation | `home` | `192.168.1.250` | `10.10.0.3/24` | `workstation.internal` (lan, primary) | +| `laptop` | workstation | `NULL` | `NULL` | `10.10.0.4/24` | `laptop.internal` (lan, primary) | + +**Hub election is by convention, not by flag.** The hub is the node whose `profile='server'` +*and* whose `WG_ADDRESS` begins `10.10.0.1` (`hooks/index.ts:188`). A lab must assign +`10.10.0.1` to the node it intends as hub or there will be no hub. + +**`site` drives direct peering** (`hooks/index.ts:197-203`). Two nodes with the same non-null +`site` peer directly with a `/32`; everything else routes through the hub's `/24`. A `NULL` +site means roaming and hub-only — deliberately, because WireGuard has no failover and a more +specific `/32` route to a dead endpoint blackholes rather than falling back. + +The four interesting pairs, all of which the lab must reproduce: + +| pair | behaviour | branch taken | +|---|---|---| +| anything → `anchor` | `Endpoint` written | underlay non-private | +| `home-server` ↔ `workstation` | direct peer, LAN endpoints, keepalive | co-located | +| **`anchor` → `home-server`** | **no `Endpoint`; learned from handshake** | not co-located, underlay private | +| `laptop` → anything | hub only, always initiates | `site` is `NULL` | + +The third is the one worth building the lab for. The code comments at `hooks/index.ts:206-224` +record what it cost to get right: testing `profile === "server"` was tried and was wrong, +because a home-hosted node **is** a server yet is not publicly reachable — *"role does not +imply reachability; the address does."* An earlier version aimed the hub at that node's public +name, which hairpinned off the household NAT: 1.77 MiB sent, 0 B received, no handshake. + +--- + +## 4. The lab + +``` + br-wan 203.0.113.0/24 TEST-NET-3 — non-private, so the code treats it as public + │ + ├── hub 203.0.113.10 profile=server site=dc WG 10.10.0.1 + │ + └── router VM 203.0.113.1 / 192.168.1.1 + │ NAT, plus one forwarded port to reproduce a published-but-NATed node + │ + br-lan 192.168.1.0/24 identical to production, same host addresses + ├── a 192.168.1.135 profile=server site=home WG 10.10.0.2 + └── b 192.168.1.250 profile=workstation site=home WG 10.10.0.3 + + c — attach to br-lan, or br-wan ("away"), or detach ("asleep") + underlay NULL profile=workstation site=NULL WG 10.10.0.4 +``` + +Kept **byte-identical** to production: the LAN subnet and its host addresses, and the entire +WireGuard plan. Only the public segment is substituted, and only because it must be. + +**The router earns its own VM.** It is what makes the published-but-NATed case real: that node +is reachable from outside only through a forwarded port, and the hub must learn its endpoint. +It also gives somewhere to break things — drop the forward and observe whether the mesh +notices or whether the public name simply stops working. + +**Names.** A resolver on the wan side is authoritative for the public zone. `.internal` names +need nothing extra: `dnsmasq-app` generates them from mesh config on each node, and writes an +`/etc/hosts` block as a floor underneath, because a node must reach the mesh DB before its own +DNS exists. + +--- + +## 5. What a node needs before any of this works + +Rows in the mesh DB — `nodes` (`name`, `user_name`, `profile`, `site`, `underlay_addr`), +`node_accessors`, `module_env.WG_ADDRESS` (the hook hard-fails without it, +`hooks/index.ts:114-116`), and `node_modules` assigning at least `wireguard`, `dnsmasq-app`, +`mesh-ca`, and `traefik` where it serves. + +On disk beforehand, because the node must reach the mesh DB before it can read any of the +above: registry database and object-store host and credentials, plus an npm token. This +ordering — contact the mesh before the mesh has configured you — is itself worth reproducing, +and is why the `/etc/hosts` floor exists. + +Everything else is generated: the WireGuard keypair locally (the private key never leaves the +node; the public key is published to `node_wg_keys`), the peer list, the DNS records, the TLS +leaf. + +--- + +## 6. Certificates — the lab issues its own + +**Settled 2026-08-22: the lab runs its own ACME issuer.** + +Public certificates use ACME **HTTP-01** via the reverse proxy, which requires genuine public +reachability, so an isolated lab cannot use the real issuer. Rather than forgo certificate +testing, the lab stands up an ACME server on its wan segment. + +**The lab keeps production's two-CA split rather than collapsing it.** Production issues +public names from a public authority and internal names from the mesh CA; a lab with one CA +would hide any bug living in that split. So: + +| | production | lab | +|---|---|---| +| public names | a public ACME authority | a test ACME server on the wan segment | +| `.internal` names | `mesh-ca` | `mesh-ca`, unchanged | + +**A test issuer is the right shape, not a shortcut.** Purpose-built ACME test servers +deliberately vary their behaviour — validation timing, nonce handling, chain composition — to +expose assumptions a well-behaved authority would let pass. A lab CA that is *too* polite +tests less than the real thing, not more. + +### It also exercises the port forward + +HTTP-01 means the issuer must reach the node being certified on port 80. In the lab: + +- the hub is directly reachable on the wan segment — straightforward +- **the published-but-NATed node is reachable only through the router's forwarded port** + +So certificate issuance for that node passes only if the forward is correct. That is exactly +why its certificate works in production, and it makes "the forward is missing" a reproducible +failure rather than a mystery. + +### Required change: `caServer` must be configurable + +`modules/traefik/docker-compose.yml:17-19` sets the challenge entrypoint, the contact address +and the storage path — but **no `caServer`**, so Traefik defaults to the public authority's +*production* endpoint. Pointing the lab at its own issuer requires adding a `caServer` flag +fed by an environment value, defaulting to production so real nodes are unaffected and the lab +overrides it per node. + +Worth noting independently of the lab: aiming at the production endpoint rather than a staging +one means every certificate experiment on a real node consumes production issuance quota, and +a retry loop can exhaust it for a week. The lab issuer removes that exposure. + +--- + +## 7. Incidental finding: `scope:` is read by nothing + +Several manifests declare `scope: public` on firewall rules — `wireguard`, `traefik`, `gitea`, +`mailu`, `qbittorrent`. It is **not part of the rule type** (`module-registry.ts:17-29`) and is +**referenced by no code** in the firewall path. Real scoping is done with `from:`, as +`modules/unifi/module.yml:52-93` does deliberately. + +So a manifest can appear to restrict a port to the public scope and in fact restrict nothing. +This is the same shape as the rule in `00-GENESIS/how-we-build.md` — *an unenforced rule is +indistinguishable from a wrong one, and costs more, because people believe it.* diff --git a/01-RESEARCH/004-lab-network/status.md b/01-RESEARCH/004-lab-network/status.md new file mode 100644 index 0000000..ac92569 --- /dev/null +++ b/01-RESEARCH/004-lab-network/status.md @@ -0,0 +1,44 @@ +# 004 — Reproducing the mesh network in a lab + +- **Status:** ONGOING — topology established and mapped; not yet stood up +- **Initiated by:** jochen, 2026-08-22 — *"the most difficult part of our VM setup will be + the networking part"* +- **Areas touched:** `modules/wireguard`, `modules/dnsmasq-app`, `modules/traefik`, + `modules/mesh-ca`, `node_accessors`, `nodes.site` / `nodes.underlay_addr`. + +## Summary + +The network is **entirely generated from mesh-DB rows by module hooks**. `install.d` performs +no network configuration whatsoever — no WireGuard, no DNS, no firewall. That makes a faithful +lab primarily a *data* problem rather than a networking problem, and means the lab exercises +the real code path instead of a reimplementation of it. + +One constraint decides whether the lab works at all: the WireGuard endpoint rule tests the +underlay address against an RFC1918 regex to decide reachability. **A simulated public segment +addressed from RFC1918 space silently prevents the mesh from forming** — no endpoint is written +for the hub, so nothing can ever initiate. The simulated public segment must therefore use +TEST-NET-3 (`203.0.113.0/24`). + +With that one substitution the lab reproduces the production topology exactly, including the +case that is hardest to get right: a node that is publicly *named* but sits behind NAT, whose +endpoint the hub can only learn from a handshake. + +Detail in [`analysis.md`](analysis.md). + +## Settled + +**The lab issues its own certificates.** Public names are certified by an ACME server on the +lab's wan segment; `.internal` names keep the mesh CA. The lab preserves production's two-CA +split rather than collapsing it, because a single-CA lab would hide any bug living in that +split. It also makes the router's port forward load-bearing — HTTP-01 must reach the +published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate +failure instead of a mystery. + +Requires one change: `caServer` is not set on the reverse proxy today, so it defaults to the +public authority's **production** endpoint. It must become configurable, defaulting to +production so real nodes are unaffected. + +## Open + +- Not yet stood up. `incus` is declared in `modules/hal/developer/module.yml` and merged + (PR #944); the lab itself is unbuilt. diff --git a/01-RESEARCH/README.md b/01-RESEARCH/README.md new file mode 100644 index 0000000..bc0284e --- /dev/null +++ b/01-RESEARCH/README.md @@ -0,0 +1,16 @@ +# 01-RESEARCH + +Investigations that have not yet hardened into design. + +## Structure + +Each effort lives in `NNN-descriptive-name/` and **must** contain `status.md` with: + +- a short summary of the effort +- who initiated it +- the areas it touches +- current status: `ONGOING`, `GRADUATED`, or `ABANDONED` + +An effort graduates by producing an ADR and a `02-DESIGN` entry. It is abandoned in +place — never deleted. What was rejected, and why, is the more expensive half to +rediscover. diff --git a/02-DESIGN/00-work-breakdown.md b/02-DESIGN/00-work-breakdown.md new file mode 100644 index 0000000..cdec042 --- /dev/null +++ b/02-DESIGN/00-work-breakdown.md @@ -0,0 +1,150 @@ +# Work breakdown — the decomposition + +How ADR 0001 gets built, in what order, and where a human must look. + +Ordering is not preference. Each phase removes a constraint the next one needs gone. + +--- + +## Rules of engagement + +These exist so the work can run largely unattended without accumulating the kind of +damage this refactor is meant to remove. + +### Autonomous by default + +An agent may, without asking: + +- read anything, measure anything, query any database read-only +- create branches, write code and tests, open pull requests +- run the test suite and typechecks +- write and update `hq/` documents + +### Always stop and ask + +- **destroying or overwriting data** — dropping a table, deleting a provision, rotating a + live credential, removing a module from a node +- **merging anything** — every merge is a human checkpoint, without exception +- **a decision the ADRs do not already answer** — record the question in the relevant + research effort rather than picking and moving on +- **any change to `hq/00-GENESIS`** — it is stable by nature + +### Definition of done for every task + +1. tests written **and failing first**, then passing +2. typecheck clean in every package the change touches +3. the local mesh (Phase 0) comes up, and the behaviour is demonstrated in it +4. `hq/` updated if the task changed or answered anything documented +5. deployed, and **delivery verified on every node** — not "the pipeline was green" + +### Non-negotiables carried from the current system + +- **Never edit mesh-managed files on disk.** Use the owning tool. +- **Never write to production databases directly.** Migrations for schema, application + code for data. +- **Every schema change ships twice** — consolidated schema *and* an incremental + migration. +- **Expand, then contract.** Add the new shape, migrate, verify, and only then remove the + old one — never in a single step. +- **A green pipeline proves transport, not effect.** Verify the effect. + +--- + +## Phase 0 — A mesh that runs locally *(prerequisite)* + +Nothing else starts until this exists. Every fault this refactor addresses was found in +production because there was nowhere else to find it. + +| # | task | done when | +|---|---|---| +| 0.1 | Container image for a node runtime | a node process starts in a container and registers | +| 0.2 | Compose topology: broker, registry DB, object store, *n* nodes | `up` yields a mesh that elects a provider node and settles | +| 0.3 | Seed a minimal mesh: nodes, one module, one provision | a module deploys end-to-end with no external service | +| 0.4 | Run the pipeline inside it | a push-equivalent produces a cascade and a deployed artifact | +| 0.5 | Fixtures for the failure modes already known | credential rotation reaching a running session; a provider deploy rotating a shared credential; a migration that ships nothing — each reproducible on demand | + +**Checkpoint:** a human confirms the local mesh reproduces at least one bug from +2026-08-22 before any decomposition begins. + +--- + +## Phase 1 — Make the model expressible + +The decomposition is impossible while a feature is a singleton per module. + +| # | task | done when | +|---|---|---| +| 1.1 | ADR 0002 — named features, per-node opt-in | accepted | +| 1.2 | Manifest: declared `features:` with type + directory | a module declares two of one kind and both build | +| 1.3 | Selection: `always` / flavor-selected / `optional` | a node installs a subset; artifacts stay flavor-blind | +| 1.4 | `requires:` moves onto the feature | a schema feature's database is not provisioned where the feature is not installed | +| 1.5 | Assignment carries the opted-in feature set | opting a node in requires no rebuild | + +**Checkpoint:** one existing module converted to declared features, deployed, verified — +before any others follow. + +--- + +## Phase 2 — Draw the boundary the domain already has + +Cheapest first, and each one proves the extraction pattern before the expensive ones. + +| # | task | extracted from | risk | +|---|---|---|---| +| 2.1 | `hal/knowledge` — one store, review workflow ported | hippocampus + noxflow `knowledge_*` | low — additive | +| 2.2 | `hal/stream` — the record; notifications and messaging as views | axon, synapse, notifications, meetings, conversations | medium | +| 2.3 | `hal/agents` — identity, licence, runs, memory, thoughts | noxflow agents, `hal/thoughts` | **high** — touches credentials | +| 2.4 | `hal/work` — what remains of noxflow | noxflow tasks | medium | +| 2.5 | `hal/ai` — provider integration, flavored | `hal/claude*` | medium | + +Each extraction is expand-then-contract: new context alongside, dual-write, verify, cut +over, remove. **Never a move commit.** + +**Checkpoint:** after 2.1, a human confirms the extraction pattern before 2.2 begins. +After 2.3, a human confirms credentials still reach every agent on every node. + +--- + +## Phase 3 — Reclaim the kernel + +Only possible once domains have modules to own their code. + +| # | task | done when | +|---|---|---| +| 3.1 | Move work-domain code out of `hal/sdk` | `workflow-engine.ts`, `task-commands.ts` live in `hal/work` | +| 3.2 | Move provider code out | `claude-credentials.ts` lives in `hal/ai` | +| 3.3 | Move delivery code out | feature handlers, artifact manager, build executor live in `hal/delivery` | +| 3.4 | Decide the residue | ADR: what `hal/sdk` keeps (open question 4) | + +**Measure:** `hal/sdk` line count, tracked per task. Today: **34,636** across **155** +files. + +--- + +## Phase 4 — Separate what the mesh runs from the mesh + +| # | task | done when | +|---|---|---| +| 4.1 | Decide the destination (open question 3) | ADR accepted | +| 4.2 | Cross-repository dependency resolution proven | a catalogue module builds against a published `@hal/*` | +| 4.3 | Move the 91 catalogue modules | this repository contains only mesh contexts | + +**Checkpoint:** move one application first and run it for a week before the rest follow. + +--- + +## Sequencing constraints + +- **0 before everything.** Unverifiable refactors are how this list got long. +- **1 before 2.** Extracting into contexts without per-node features recreates the module + count inside the new names. +- **2 before 3.** A domain can only own its shared code once the domain has a module. +- **2.3 after 2.1 and 2.2.** Agents touch credentials; do it once the pattern is proven on + cheaper contexts. +- **4 last.** It is the only phase that is pure movement, so it is the only one safe to + defer indefinitely. + +## What "done" looks like + +Eight contexts. `hal/sdk` holding only what is genuinely cross-cutting. A mesh that stands +up on a laptop. A module count that grows only when the domain does. diff --git a/02-DESIGN/01-end-to-end-testing.md b/02-DESIGN/01-end-to-end-testing.md new file mode 100644 index 0000000..e1e028f --- /dev/null +++ b/02-DESIGN/01-end-to-end-testing.md @@ -0,0 +1,355 @@ +# End-to-end testing + +**What is under test is a module.** The mesh is the harness. + +You change a module or write a new one, run it end to end, and get a verdict before it goes +anywhere near production. That loop is the product of this design; everything else exists to +make it fast and honest. + +This is the Phase 0 prerequisite from [`00-work-breakdown.md`](00-work-breakdown.md). + +--- + +## Where this sits in the way work happens + +Work reaches the mesh along one path today: + +``` + a change is made a human in a session, or an agent given work + │ + ▼ + a pull request appears + │ + ▼ + a human reads the diff and merges ← the gate + │ + ▼ + the coordinator delivers to the real nodes + │ + ▼ + production reports whether it worked ← the test +``` + +**The gate is a human reading a diff, and the test is production.** That is workable at a +change a day and it is the constraint at ten. For autonomous work it is worse than a +constraint: an agent's output arrives as a diff that *looks* right, carrying no evidence +that it runs, and the only reviewer is the condition `00-GENESIS/context.md` calls mandatory +— *human agents are few, often one, and usually asleep.* + +The missing step goes between the pull request and the merge: + +``` + a pull request appears + │ + ▼ + the coordinator delivers the branch to a SCENARIO mesh ← the missing step + and runs the same stages, ending in verify + │ + ▼ + the verdict is attached to the pull request + │ + ▼ + a human merges evidence rather than hope + │ + ▼ + the coordinator delivers to the real nodes — same verify, now loud +``` + +### One pipeline, two targets + +| target | triggered by | what a failure means | +|---|---|---| +| a scenario mesh | a branch, or a pull request | the change is not finished; it should not merge | +| the real mesh | a merge | a red delivery, loudly, before anything is built on it | + +Same coordinator, same cascade, same stages, same verification. **Only the target mesh +differs.** This is the through-line of the whole design — one verification statement with two +jobs, one runner with two callers, one pipeline with two targets. Nothing forks, so nothing +drifts. + +### What it demands + +- **The coordinator must accept a target mesh.** Today a pipeline's targets are derived: the + nodes that have the module assigned. It needs to be able to run the same pipeline against + a mesh named by the request instead. +- **Scenarios must be concurrent and cheap.** Several agents working means several scenarios + at once, each needing its own network and nodes. This is affordable with system containers + and would not be with virtual machines — the unit choice is what makes the gate possible + at all. +- **The gate is only as good as the verification behind it.** A module with no assertions + gets a weak gate: delivery succeeded, nothing checked. So **verification coverage becomes + the number that matters**, and it starts at approximately zero. +- **It has to be fast enough to wait for.** A gate an agent cannot wait on is a report nobody + reads. + +--- + +## The coordinator drives it + +The temptation is to build a framework that delivers a module and checks it. That would be a +**second delivery path**, and a second delivery path is worthless — the faults worth catching +live in the real one. + +So the rule is not "no new components". It is: **nothing new drives delivery.** A scenario is +a complete mesh with its own coordinator. Push the working tree to that mesh's forge; its +coordinator does exactly what a coordinator does — works out the cascade, dispatches build, +install, configure, start, and then **verify** — and its meshware executes on its nodes. The +result of that pipeline *is* the verdict. + +A **test runner** is a legitimate component within that, used *by* the coordinator rather +than instead of it. Two jobs plausibly belong to it, and their boundary is worth settling +before either is built: + +- **scenario lifecycle** — materialise the mesh a test needs, restore it to a snapshot, tear + it down. Something must do this before a coordinator exists to drive anything. +- **assertion execution** — give verification more than "run a script and check the exit + code": setup and teardown, timeouts, retry-until-true for things that settle, and results + structured enough to report rather than grep. + +The line to hold is the pipeline itself. A runner that stands up a mesh and executes +assertions is a component. A runner that decides what to build, in what order, and ships it +to a node is a fork of the coordinator. + +### It has two callers, and they want different things + +The runner serves **the coordinator** and **a person developing the mesh**, and its interface +has to suit both: + +| caller | wants | +|---|---| +| the coordinator | non-interactive, structured results it can record against a pipeline, a clean teardown, no prompts and no colour | +| someone working on HAL | readable output, the failing mesh **left standing** to open a shell into, and a way to re-run one assertion without repeating the whole delivery | + +Hence at least two verbs: one that runs to a verdict and tears down, and one that stands a +scenario up and leaves it there. The second is how a developer works *inside* a mesh — +which is the thing the current host-borrowing tooling is really for, and the reason this +replaces it rather than sitting beside it. + +That the same runner serves both is deliberate, and it is the same argument as the module's +own assertions serving both development and delivery: **one statement, two jobs.** Anything +that only the developer path can do is a divergence, and it will drift. + +``` + edit a module in the working tree + │ + ▼ + push to the scenario's forge + │ + ▼ + the scenario's COORDINATOR runs a pipeline ← existing machinery, unchanged + │ + ├── cascade: which modules are affected + ├── build → install → configure → start + └── verify: the module's own assertions ← existing stage, dispatched today + │ + ▼ + the pipeline result is the verdict +``` + +This is the same property `00-GENESIS/mission.md` asks for: *the mesh's own components ship +through the same machinery as anything else it carries — if they need an exception, the +machinery is not finished.* A test that needed its own delivery path would be that exception. + +**It tests the working tree**, because the forge is inside the scenario. A loop that requires +pushing to production and waiting is not a loop. Real path, local code, nothing shared with +production. + +--- + +## What it catches + +This list is the specification. These are the ways a module change fails today, and each one +currently reaches production or wastes a pipeline run: + +| failure | why it survives today | +|---|---| +| a file never reached the artifact | absence and "declares nothing here" are indistinguishable, so it ships green | +| a migration compiled to nothing, or never ran | the stage reports success when there is nothing to run | +| the manifest is wrong — bad package list, wrong paths | validated shallowly, if at all | +| a capability was never provisioned, or its credential never arrived | delivery reports transport, not effect | +| an environment value was not generated | the module starts and reads a default | +| the service did not come up, or came up and crashed | nothing asserts it is still running a minute later | +| the dependency cascade did not include the module | a green pipeline that rebuilt the wrong set | +| it works on a fresh install but breaks on upgrade | almost never exercised — see below | +| it works on one node and not another | only one node is ever tried | + +A test that only proves "the pipeline went green" reproduces the exact blindness this is +meant to remove. + +--- + +## Fresh install and upgrade are different tests + +The most common shape of a module bug is: works from scratch, breaks on the machine that +already had the previous version. Existing state outranks new state, a file is added but +never removed, a migration assumes a column that an older node lacks. + +Snapshots make both cheap, so both are default: + +- **fresh** — restore a mesh that has never seen the module, deliver, assert +- **upgrade** — restore a mesh running the *previous released version*, deliver the working + tree over it, assert + +Same assertions, different starting state. A module that passes one and fails the other is +the normal case, not an edge case. + +--- + +## A module carries its own assertions + +**A module states what must be true about it, and it states it once.** That statement is the +module's verification — the stage the coordinator already dispatches at the end of every +delivery, with a working handler, implemented today by essentially nothing. + +It asserts **outcomes**, never that a step ran: the unit is active and still active shortly +after, the schema has the column, the name resolves, the credential authenticates, the file +on the node holds what the mesh believes it holds, the endpoint answers. + +The same statement serves both places, which is the point: + +| where it runs | what a failure means | +|---|---| +| in a scenario, during development | your change is not finished — cheap, fast, nobody affected | +| on delivery to production | the deploy is red, loudly, before anyone builds on it | + +**This is what finally makes verification worth writing.** Today it can only ever cost you a +deploy, which is precisely why almost no module has one. Give it a second job — telling a +developer whether their change works — and writing it stops being an act of discipline and +starts being the fastest way to get an answer. + +Where a module has no verification yet, the coordinator asserts the generic invariants it can +know on its own: the artifact contained what the manifest declared, the migrations that were +pending ran, the declared capabilities were provisioned, the declared services are up. + +--- + +## The mesh shape is a parameter + +A test declares the mesh it needs, and the default is the smallest one that can exercise the +module: + +```yaml +module: a-web-service +mesh: + my-cool-node: { role: published, publishes: my-cool-node.com } +assert: + - https://my-cool-node.com answers 200 + - the certificate presented is valid for that name +``` + +One node, because one node is enough to answer that question. Standing up a mesh to test one +module is the same mistake as starting the application to test a function. + +More nodes when the module's behaviour is *between* nodes: + +```yaml +module: a-module-requiring-a-database +mesh: + store: { role: anchor } + consumer-a: { role: resident } + consumer-b: { role: mobile } +assert: + - both consumers authenticate against the database + - after redeploying the provider, both still do +``` + +Node names and domains in a test are **invented**. The mesh under test is whatever the test +says it is — which is also how this document stays free of any particular installation. + +### Scale, when scale is the question + +Size is chosen by what is being asked, and ranges from one node to twenty or more. + +| size | what only this size answers | +|---|---| +| one | does it install, migrate, provision and run at all — the fastest loop | +| two–three | anything *between* nodes: delivery, provisioning, rotation, absence | +| ten–twenty | whether a fan-out reaches *every* node, whether the cascade converges, whether something is quietly quadratic | + +A fan-out reaching three of four nodes reads as a flake; at twenty it is a diagnosis, and +which nodes were missed tells you why. **This is what the unit choice below bought** — twenty +virtual machines do not fit on a workstation, and twenty system containers do. + +--- + +## Two kinds of test + +**Module tests** are the daily case and the reason this exists: does my change work, end to +end. + +**Mesh tests** use the same machinery to ask whether the mesh itself behaves — that a +credential rotation reaches every consumer, that delivery to an absent node is reported as +pending rather than done, that a returning node catches up. These are fewer and change +rarely, but they are where the known production faults get encoded so they stay fixed. + +The known faults become mesh tests that fail today. That is the Phase 0 checkpoint. + +--- + +## The substrate + +### A node is a system container + +An OS userspace with its own init, its own network interface, its own filesystem, sharing +the host kernel. + +**A node's job is to run containers**, so modelling a node *as* an application container +inverts the thing being modelled: it forces nested containers through a privileged daemon or +a shared socket, and a shared socket makes isolation between nodes cosmetic. A system +container has no such problem — init runs as PID 1, so units and timers work as written; +containers nest properly, so module stacks run as they do anywhere. **Module code and mesh +code both run unmodified**, which is the property that makes a verdict trustworthy. Anything +needing a special case locally is a divergence that will hide a fault. + +Boot is around a second and snapshots are cheap, which is what makes this an inner loop +rather than an errand. + +**Any node can be a full virtual machine instead**, through the same tooling and the same +test file — when a question needs a kernel to answer it, or when the point is that nodes are +*not* identical. + +### The network + +Two segments and an overlay, because some module behaviour is only visible across a real +network boundary: + +- **wan** — a published node holds an address here, and an authoritative resolver maps its + name to it, so a public touchpoint is real enough to exercise routing, virtual hosts and + certificates +- **local** — behind translation, as a home network is +- **the overlay** — the mechanism production uses; the mesh addresses peers by mesh name and + never learns which segment anyone is on + +A node can be moved between segments or detached entirely, mid-test. + +### What is not real + +- **the model provider** — thinking is stubbed, so tests cost nothing to run +- **the public internet** — a bridge, with an authoritative resolver rather than delegation +- **the public certificate authority** — the lab runs **its own ACME issuer** on the public + segment, so issuance, challenge and renewal are genuinely exercised rather than stubbed. + Internal names keep the mesh CA, so the lab preserves production's two-authority split + rather than collapsing it into one. + +Everything a node itself does is real, because a node is a real machine. + +--- + +## Consequences + +**Bringing a node into being is part of the framework.** A test creates its own nodes — one +for the smallest, twenty for the largest — repeatably, unattended, and cheaply enough to do +it twenty times in a row. Those are the constraints a real mesh wants, so the mechanism the +runner needs is the one the mesh should keep. + +**Module verification becomes worth writing**, because it is the thing that gives a +developer a verdict, not just a stricter deploy. + +**Host-borrowing ends.** Today's tooling starts providers on the host's own init system and +reads credentials from host paths, because there is nowhere else to put a mesh. Once there +is, a workstation stops being collateral. + +**The supervision question stops gating anything.** +[`01-RESEARCH/003-service-supervision`](../01-RESEARCH/003-service-supervision/analysis.md) +remains open on its own merits — and once this exists, its options are cheap to try rather +than expensive to argue about. diff --git a/02-DESIGN/README.md b/02-DESIGN/README.md new file mode 100644 index 0000000..2eef4c8 --- /dev/null +++ b/02-DESIGN/README.md @@ -0,0 +1,10 @@ +# 02-DESIGN + +The authoritative specification. Implementation is built against what is written here. + +A document enters this folder only after the decision behind it is recorded in +[`adr/`](../adr/) and the research that produced it is marked `GRADUATED`. + +Empty for now: the decomposition in ADR 0001 is decided but not yet specified. The first +entries will be the per-context designs — `hal/mesh` brokering, the `hal/stream` record, +and the feature model that `hal/delivery` owns. diff --git a/DECISIONS.md b/DECISIONS.md new file mode 100644 index 0000000..45e3316 --- /dev/null +++ b/DECISIONS.md @@ -0,0 +1,91 @@ +# Decision ledger + +Every decision, in the order it was taken. Append-only — a decision that stops being true is +marked superseded and left in place, because the reasoning that was rejected is the expensive +half to rediscover. + +An entry here is a **record**, not the reasoning. Anything architecturally significant carries +its full context, options and consequences in an [ADR](adr/); anything still being worked out +lives in [`01-RESEARCH`](01-RESEARCH/). This file is the index that makes both findable, and +the place small decisions live that never warrant a document of their own. + +**Columns.** *Decided* is who made the call. *Where* points at the reasoning. A decision with +no pointer is one small enough that this line is the whole record. + +--- + +## 2026-08-22 — decomposition + +| # | Decision | Decided | Where | +|---|---|---|---| +| 1 | The mesh brokers capabilities; nodes host; agents think. Eight bounded contexts replace 33 platform modules. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) | +| 2 | Nodes and agents decouple — a node holds no licence; an agent holds credentials and delivery follows its bindings and modality. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) | +| 3 | noxflow dissolves; `hal/work` inherits tasks and workflows. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) | +| 4 | Third-party modules leave this repository — they run *on* the mesh, not *of* it. | jochen | [ADR 0001](adr/0001-mesh-brokers-nodes-host-agents-think.md) | +| 5 | ~~Documentation lives inside the code repository under `hq/`.~~ | jochen | **Superseded by #27** | + +## 2026-08-22 — the lab + +| # | Decision | Decided | Where | +|---|---|---|---| +| 6 | Phase 0 is a **development environment**, not a fixture rig — it is what the host-borrowing dev tooling becomes. | jochen | [002](01-RESEARCH/002-local-mesh/status.md) | +| 7 | The delivery trigger is a **real forge** inside the lab, not a synthetic event — the webhook relay is part of what is under test. | jochen | [002](01-RESEARCH/002-local-mesh/status.md) | +| 8 | ~~A lab node is a **system container**, promotable to a virtual machine.~~ | jochen | **Superseded by #10** | +| 9 | Adoption is a legacy path and is out of scope. Bringing nodes into being is the lab runner's job instead. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 10 | A lab node is a **virtual machine** running the real install. Supersedes #8: the scale argument for system containers was invented rather than required, and a virtual machine dissolves the fidelity question instead of answering it. | jochen | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md) | +| 11 | The environment is called **the lab**. | jochen | — | +| 12 | The lab is driven by `incus` — for virtual machines, snapshots and bridges through one interface, and because it also runs system containers if a scale run is ever needed. | jochen | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md) | +| 13 | The simulated public segment uses **TEST-NET-3** (`203.0.113.0/24`). Not cosmetic: an RFC1918 public segment makes the hub test as unreachable and the mesh silently never forms. | — | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md), [004](01-RESEARCH/004-lab-network/analysis.md) | +| 14 | The lab **issues its own certificates**, keeping production's two-authority split rather than collapsing it. | jochen | [ADR 0002](adr/0002-a-lab-node-is-a-virtual-machine.md) | + +## 2026-08-22 — what the lab is for + +| # | Decision | Decided | Where | +|---|---|---|---| +| 15 | **The module is what is under test; the mesh is the harness.** The loop is: change a module, run it end to end, get a verdict. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 16 | **Nothing new drives delivery.** A lab mesh has its own coordinator; the pipeline that runs is the real one. A second delivery path would be blind to exactly the faults worth catching. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 17 | A **test runner** is a legitimate component *used by* the coordinator — lab lifecycle and assertion execution. The line is the pipeline: a runner that decides what to build is a fork of the coordinator. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 18 | The runner has **two callers** — the coordinator, and a person developing the mesh — so it needs both a run-to-verdict verb and a leave-it-standing verb. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 19 | **A module carries its own assertions**, stated once, in the verification stage the coordinator already dispatches. Running them where failing is free is what makes writing them worth doing. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 20 | A test declares the mesh it needs; **size ranges from one node upward**, chosen by the question rather than by what the mesh happens to have. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | +| 21 | The real topology is **one test among many**, not the baseline. Anything only testable there is a gap in the vocabulary. | jochen | [design](02-DESIGN/01-end-to-end-testing.md) | + +## 2026-08-22 — working agreements + +| # | Decision | Decided | Where | +|---|---|---|---| +| 22 | **Never install packages by hand.** A package is declared in the module manifest and arrives the way every other package does. | jochen | — | +| 23 | **Never open a pull request unprompted.** A permissions list saying it is allowed is not a request. | jochen | — | +| 24 | **Every merge is a human checkpoint**, without exception. | jochen | [`02-DESIGN/00-work-breakdown.md`](02-DESIGN/00-work-breakdown.md) | +| 25 | Work in an **isolated worktree**, never a shared checkout. | jochen | `CLAUDE.md` | +| 26 | Decisions are recorded **here**, thoroughly, as they are taken. | jochen | this file | +| 27 | HQ is **its own repository**, `hal-hq`. Supersedes #5: the original objection was to a fourth knowledge *system*, which indexing answers rather than location. Cadence, reviewers, and a scope wider than one repository all argue for separation. | jochen | [`README`](README.md) | +| 28 | **This repository is public.** Written for a reader who is not its author and has no access to the mesh it describes. No routable addresses, real domains, hosting providers, node names, absolute paths, or operational detail useful only to an attacker. Supersedes the previous rule that research may name instances. | jochen | [`README`](README.md) | + +--- + +## Observations — not decisions, but they should not be lost + +Things established by measurement that no decision has yet answered. + +| Observed | What it means | +|---|---| +| **A failed package install does not fail the job.** Declaring `incus` produced `error: failed retrieving file … 404` from every mirror, then `-> error installing repo packages`, and the prepare job reported **success**. The package is absent; the pipeline is green. | The first thing the lab was asked to install demonstrated the exact fault the lab exists to catch. A fix is already written and open as PR #848, unmerged since 2026-08-20. | +| **The package database is stale on at least one node.** The install ran without a sync, so it requested a version the mirrors had already superseded — 404 from every mirror. | A declared package can fail purely because the node's index is old, and today that failure is silent. | +| **`scope:` on firewall rules is read by no code.** Declared in five manifests; not part of the rule type. Real scoping is `from:`. | A manifest can appear to restrict a port and restrict nothing. | +| **The reverse proxy sets no `caServer`**, so certificate issuance targets the public authority's production endpoint rather than staging. | Every certificate experiment on a real node consumes production issuance quota. | +| **`test/pipeline/` has been unbuildable since 2026-06-04**, when the npm workspace it depends on was removed. Nothing runs it. | The repository's only end-to-end pipeline test has been silently dead for two and a half months. | + +## Deliberately not decided + +Recorded so they are not mistaken for oversights. + +| Question | Status | +|---|---| +| Which supervision model the mesh adopts — keep the host init system, drop the redundant per-module layer, containerise the daemons, or write a supervisor. | Open. Options costed in [003](01-RESEARCH/003-service-supervision/analysis.md). **No longer gates the lab.** | +| Whether the lab verdict is a workflow guard or an advisory check on the pull request. | Open, and deliberately trivial — a policy detail, changeable in an afternoon, not an architectural choice. | +| `hal/scheduler` — infrastructure or part of `hal/work`. | Open, from ADR 0001. | +| Which context owns the executor. | Open, from ADR 0001. | +| Catalogue destination — one repository or many. | Open, from ADR 0001. Phase 4. | +| What `hal/sdk` keeps after extraction. | Open, from ADR 0001. Phase 3. | +| Where human agent modality is recorded — which user, on which node, a human agent acts as. | Open, from ADR 0001. Required by the model; not yet stored. | diff --git a/README.md b/README.md new file mode 100644 index 0000000..d0f2853 --- /dev/null +++ b/README.md @@ -0,0 +1,77 @@ +# HAL — HQ + +The single source of truth for what the HAL mesh **is**, what it is **becoming**, and +why. Implementation lives in `modules/`; the reasoning behind it lives here. + +## Structure + +| Folder | Purpose | +|--------|---------| +| [`00-GENESIS`](00-GENESIS/) | Mission and foundational context. The northern star for every decision. | +| [`01-RESEARCH`](01-RESEARCH/) | Active and historical investigations, before they harden into design. | +| [`02-DESIGN`](02-DESIGN/) | The authoritative specification. Implementation is built against this. | +| [`adr`](adr/) | Numbered architecture decisions — what was chosen, and what was rejected. | + +## Rules + +- Markdown only. +- No new top-level folders without explicit confirmation. +- Knowledge flows `GENESIS → RESEARCH → DESIGN`. Research graduates into design only + after analysis against GENESIS confirms alignment. +- **GENESIS and DESIGN are instance-agnostic.** They describe the mesh as a concept — no + machine names, no counts, no topology. A reader must not be able to tell how many nodes + the author happened to have. +- **RESEARCH describes real observations, but never identifies the mesh it observed.** + Evidence is what makes research worth reading, and the shape of a finding survives + anonymisation intact — *a node publicly named but behind a household NAT* carries the whole + lesson without naming anything. +- A document that states a rule about the mesh should say how that rule is **checked**. + This repo has a rule requiring every tools module to declare `brain` as a dependency; + zero modules do. An unenforced rule is indistinguishable from a wrong one. + +### This repository is public + +Written for a reader who is not its author and has no access to the mesh it describes. +Concretely, nothing here may contain: + +- **routable addresses, real domain names, hosting providers, or node names** — use the + documentation ranges (RFC 5737 `203.0.113.0/24`, RFC 1918) and role names such as `anchor`, + `home-server`, `workstation`, `laptop` +- **absolute paths** from anyone's machine, usernames, home directories, or email addresses +- **credentials in any form**, including lengths or hashes of live secrets +- **operational detail that is only useful to an attacker** — which host is the VPN hub, on + which port, which node is reachable only through a forwarded port + +Private-range addresses and the overlay plan are fine: they describe a pattern, not a target. + +The test is whether a paragraph would still teach something to a stranger running an entirely +different mesh. If it would, it belongs. If it only makes sense to someone who knows this +particular installation, it is either a note in the wrong place or a disclosure. + +## Why this is its own repository + +It began inside the code repository, on the reasoning that HAL already has a mesh-native +knowledge store and that adding a fourth knowledge system would repeat the mistake this +folder was created to fix. + +**That objection was about a fourth knowledge *system*, and it is answered by indexing, not +by location.** These documents are still indexed into the knowledge base, so +`recall_search` returns them beside everything else. One source, many surfaces — which was +always the actual requirement. Where the source is authored is a separate question. + +Answered separately, a repository of its own is the better home: + +- **The cadence is different.** A decision changes when thinking changes, not when code + changes. Tying documents to a code branch means they merge on the code's schedule. +- **The reviewers are different.** A design argument is not reviewed the way an + implementation is, and it should not queue behind a build. +- **The scope is wider than one repository.** ADR 0001 sends most modules out of the + monorepo entirely. Documentation that governs several repositories cannot live inside one + of them. + +The trade is real and worth naming: a change to a document and the change to the code it +describes can no longer land in one commit. Keeping them honest is a discipline now rather +than a mechanism — which is why [`DECISIONS.md`](DECISIONS.md) records decisions as they are +taken, and why a document that states a rule should say how the rule is checked. + +Recorded as decision 27 in [`DECISIONS.md`](DECISIONS.md). diff --git a/adr/0001-mesh-brokers-nodes-host-agents-think.md b/adr/0001-mesh-brokers-nodes-host-agents-think.md new file mode 100644 index 0000000..5ada9d9 --- /dev/null +++ b/adr/0001-mesh-brokers-nodes-host-agents-think.md @@ -0,0 +1,185 @@ +# 1. The mesh brokers capabilities; nodes host; agents think + +- **Status:** Accepted +- **Date:** 2026-08-22 +- **Deciders:** jochen + +## Context + +HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split +because splitting is the only granularity the platform offers, and domains are merged +because a shared database is the only integration it offers. Both pressures push in the +same direction: boundaries end up drawn by deployment accident rather than by domain. + +Three observations establish the state. + +**The word "agent" means two different things.** The original design treated a node as an +agent with thinking abilities. Later, noxflow implemented agents as employees with +skills, workspaces and tasks. Both survive. The collision is visible in the data — there +are **two agent rows per node**: + +| agent | skills | +|---|---| +| one named after the node | `{deploy,verify,operate}` | +| one named `hal-` | `{}` | + +One carries the work; the other carries only identity, existing to hold a licence for +HAL's own sessions. The same split appears in the schema: `nodes.hal_claude_account` and +`agents.claude_account` are one fact in two tables in two databases, and what was +documented as "three licence touchpoints" is one concept modelled three times. + +**Domains integrate by sharing a schema.** `noxflow` is 45 tables spanning five domains — +tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose +is 7 of 45. Because agent identity lives in noxflow's database, work that belongs +elsewhere must be implemented there: per-agent Claude credentials had to be written by +the noxflow runtime, even though node identity — the same kind of fact — lives in the +mesh registry. + +**The core domain has no context to live in.** The mesh's distinguishing feature is that +a module declares `requires: postgres/database` and never learns where the database +lives, who owns the credential, or how it rotates. That is capability brokering, and it +is what makes this a mesh rather than four machines with a configuration manager. Yet the +logic implementing it sits in `modules/postgres/tools/index.ts` — a provider module, +where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its +invariants were found violated simultaneously (see Consequences). + +## Considered Options + +1. **Keep the current layout; fix bugs as they surface.** Rejected. The faults are not + independent. Every incident on 2026-08-22 — credentials written by the wrong module, a + dependency question answered wrongly twice, an invariant enforced nowhere — traced to + a boundary that was never stated. Fixing them individually leaves the generator intact. + +2. **Merge aggressively into few large modules.** Fewer names, same problem: a shared + schema across domains is what produced noxflow, and doing it deliberately would + produce it again at larger scale. + +3. **Decompose by bounded context, with the mesh as a broker.** Name contexts after their + aggregates, integrate through a published record rather than a shared schema, and let + deployment granularity be a feature-level concern rather than a reason to create a + module. **Adopted.** + +## Decision + +### The domain, in one sentence + +**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas +that think and act.** Everything else supports one of those three. + +### Nodes and agents are decoupled + +A node is a place where an agent can run — that is the entire relationship. There is no +resident agent, no node-owned identity, no ownership in either direction. Agents named +after a node remain, as **ordinary agents** that happen to hold infra skills. + +Consequently `nodes.node_license` and `nodes.hal_claude_account` cease to exist: a node +does not authenticate to a model provider, agents do. The two agent rows per node merge. + +### There is one kind of participant, and some are human + +Human and non-human participants are both **agents**. Both hold identity and credentials; +both act, remember and coordinate. What differs is **modality** — how an agent acts: + +| modality | credential is delivered to | +|---|---| +| spawned session | that agent's own config directory | +| shell or desktop | that agent's home on the node it acts from | + +A node holds no licence. **An agent holds credentials, and delivery follows that agent's +node bindings and modality.** A human agent's grant arrives in the home directory of the +user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's +grant in its config directory, with a different target. + +This requires one fact the mesh does not record today: which user, on which node, a given +human agent acts as. Adding it is what removes the node licence — the node is currently +standing in for an identity the mesh cannot name. It is also what makes a second human +agent require no new mechanism. + +### Contexts + +| context | aggregate | subdomain | +|---|---|---| +| `hal/mesh` | **Provision**, Node, Module — brokering and its bookkeeping | core | +| `hal/agents` | **Agent** — identity, licence, runs, memory, thoughts | core | +| `hal/work` | **Task** — workflows, bindings | core | +| `hal/stream` | **Thread** — mentions, messages, meetings, notifications | core | +| `hal/delivery` | **Pipeline** — jobs, artifacts, features | supporting | +| `hal/knowledge` | **Document** — spaces, revisions, review | supporting | +| `hal/ai` | **Licence** — provider grants and rotation | supporting | +| `hal/observability` | **Check** | supporting | +| `hal/config` | **Setting** — env, secrets, PKI | generic | + +A *brain* — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is +not a module. Anatomy makes attractive names and poor boundaries; today `hal/brain` names +infrastructure and `hal/cortex` describes itself as messaging while running nowhere. + +### Provisioning is the core domain, not plumbing + +`hal/mesh` is a broker; the registry is its bookkeeping. Its invariants are explicit and +owned: + +- one rotation source per resource +- a credential change fans out to every consumer +- a consumer never holds a credential the provider does not know about + +### Contexts integrate through the record, never a shared schema + +`hal/stream` is the published language. A context publishes; it does not join across a +boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a +notification is a mention not yet read. + +### Third-party software leaves the repository + +`plex`, `sonarr`, `postgres`, `verdaccio`, `docker-registry` and 88 others run **on** the +mesh; they are not **of** it. The pipeline and provisioning are deliberately +module-agnostic, so HAL's own modules dogfood exactly what external modules use — which +is what makes the separation safe rather than merely tidy. + +## Consequences + +**The invariants now have an owner, and were measurably unowned before.** On 2026-08-22, +`provision_ensure` — documented as "NEVER rotates an existing secret" — was found to mint +a new password on every adoption and update only the provider's row. `hal_notifications` +consumers on three nodes held dead credentials for two days; `hal_transcripts` had two +rows written 216 ms apart, so at most one could match the live role. + +**noxflow dissolves.** `hal/work` inherits tasks and workflows — the concept it was built +for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes +away. + +**`hal/sdk` shrinks.** 155 files, 34,636 lines, containing code from every context — +including `workflow-engine.ts` and `task-commands.ts`, work-domain logic in the kernel +every module imports. Each landed there to avoid a cycle between modules that both needed +it; a domain module can only own its shared code once the domain has a module. Extraction +is therefore downstream of this decision, not independent of it. + +**Two mechanisms are prerequisites, not follow-ups.** + +- *Named features with per-node opt-in.* Without it, every independently deployable unit + inside a context becomes a module again and the count returns. `hal/claude-licences` + exists solely because one daemon must run on one node. +- *A local mesh in containers.* `dev_up` starts providers "via systemctl (same as + production)" — it borrows the host, and no mesh can be stood up locally. Everything that + manifests **between** nodes is therefore discoverable only in production, which is where + every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable. + +**Credential delivery becomes uniform.** Per-agent credential directories, built for +spawned sessions, extend to human agents unchanged. A file previously scoped to a node +becomes scoped to an agent — which would have prevented the class of failure where a +rotation reached one node of four while the mesh reported success. + +**Migration is incremental and long.** Contexts can be extracted one at a time behind the +existing pipeline. Nothing here requires a flag day, and nothing here is cheap. + +## References + +- [`01-RESEARCH/001-module-domain-decomposition`](../01-RESEARCH/001-module-domain-decomposition/analysis.md) + — current-state evidence, table counts, open questions +- [`00-GENESIS/how-we-build.md`](../00-GENESIS/how-we-build.md) — naming and integration rules +- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler + array that makes a feature a singleton per module +- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential +- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the + constraints that make a unit independently shippable, applicable unchanged to features +- impire.io / soulstream — *the record* as integration substrate, personas over services, + and "cheap awareness and expensive thinking" diff --git a/adr/0002-a-lab-node-is-a-virtual-machine.md b/adr/0002-a-lab-node-is-a-virtual-machine.md new file mode 100644 index 0000000..e0b1bdb --- /dev/null +++ b/adr/0002-a-lab-node-is-a-virtual-machine.md @@ -0,0 +1,135 @@ +# 2. A lab node is a virtual machine running the real install + +- **Status:** Accepted +- **Date:** 2026-08-22 +- **Deciders:** jochen + +## Context + +[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite +rather than a convenience: *"everything that manifests between nodes is discoverable only in +production, which is where every fault of 2026-08-22 was found."* + +Two things were measured while establishing what exists +([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)): + +- **There is no local mesh.** The dev tooling starts providers through the host's own init + system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`). + It borrows the machine because there is nowhere else to put a mesh. +- **The one containerised node in the repository has been unable to build since 2026-06-04**, + when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it. + +So the question is not how to improve a local mesh. It is what a node *is* when it is not a +physical machine. Every subsequent question — how faithful is faithful enough, what may be +mocked, which failures remain reachable — follows from that one answer. + +The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and +hardware virtualisation present. + +## Considered Options + +1. **An application container.** Rejected. **A node's job is to run containers**, so modelling + a node as one inverts the thing being modelled: module service stacks then require nested + containers through a privileged daemon, or a shared socket that makes isolation between + nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start + and the least like a node. + +2. **A system container.** Rejected, after first being recommended. It is genuinely good — + real init, properly nested containers, roughly a second to boot, cheap snapshots — and it + is the only option that makes a twenty-node run affordable. It was rejected because **the + scale requirement that justified it was invented rather than required**: the stated goal is + to run the real mesh, which is four nodes, on one computer. And a system container still + forces the question a virtual machine dissolves — *how faithful must a node be?* — which + then has to be answered again for every capability under test. + +3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive: + no storage pools, no snapshot management, no network management, no virtual machines. + Snapshots are what make the loop fast, so the saving is not worth what it costs. + +4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script + turns into a node. + +## Decision + +**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image, +runs the real install, and becomes a node. It is not a model of a node, so no question arises +about how good the model is. + +The environment is called **the lab**. + +Three things follow directly and are decided here: + +### The lab is driven by `incus` + +Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the +bridges between them, through one interface. It also manages system containers, so if a run +ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite. + +Declared in `modules/hal/developer/module.yml`, so it installs the way every other package +does. + +### The simulated public segment uses TEST-NET-3 + +`203.0.113.0/24`, reserved by RFC 5737, never routable. + +This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the +peer's underlay address against an RFC1918 regex +(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from +private space makes the hub test as unreachable, so no spoke writes an endpoint for it, +nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault +rather than an addressing mistake. + +The production LAN subnet and the entire overlay address plan are reproduced unchanged. + +### The lab issues its own certificates + +Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh +CA. **The lab keeps production's two-authority split rather than collapsing it**, because a +single-authority lab would hide any fault living in that split. + +This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a +published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate +failure rather than a mystery. + +## Consequences + +**The fidelity question disappears, and with it a class of argument.** There is no "how real +is this node" to litigate per capability, because the node is real. What remains not-real is a +short, enumerable list: the model provider, the public internet, and the public certificate +authority. + +**The install becomes the thing under test.** A container-shaped lab would have had to skip +the bootstrap entirely. Here it runs, so it is exercised on every fresh lab. + +**Reproducing the network is mostly a data problem.** The bootstrap performs no network +configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks +from mesh-DB rows. The lab therefore exercises the same code production runs rather than a +reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)). + +**Scale runs get expensive, and this is the real cost.** Four virtual machines are +comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out +reaching most consumers rather than all, a cascade that stalls with many modules — stay hard +to reproduce. The mitigation is that the same tooling runs system containers, so a scale run +remains possible at lower fidelity if one is ever genuinely needed. + +**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run +includes a full delivery — build, publish, install, migrate — measured in minutes, so boot +time is noise. + +**One change is required before the lab can issue certificates.** The reverse proxy sets no +`caServer`, so it defaults to the public authority's *production* endpoint +(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to +production so real nodes are unaffected. Worth noting on its own: aiming at production rather +than staging means every certificate experiment on a real node consumes issuance quota. + +## References + +- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and + the four host couplings that only obstruct a container-shaped node +- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology + being reproduced and the endpoint constraint +- [`02-DESIGN/01-end-to-end-testing.md`](../02-DESIGN/01-end-to-end-testing.md) — what the lab + is for +- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments + recording what it cost to get right +- RFC 5737 — reserved documentation address blocks diff --git a/adr/README.md b/adr/README.md new file mode 100644 index 0000000..e6df8dd --- /dev/null +++ b/adr/README.md @@ -0,0 +1,30 @@ +# Architecture Decision Records + +One file per decision, numbered, never deleted. A superseded ADR gets its status changed +and a pointer to what replaced it — the reasoning that was rejected is the expensive half +to rediscover. + +## Format + +``` +# N. Title in plain language + +- **Status:** Proposed | Accepted | Superseded by ADR-XXXX +- **Date:** YYYY-MM-DD +- **Deciders:** + +## Context what is true today, with evidence +## Considered Options numbered, each with why it was rejected +## Decision what we are doing +## Consequences what follows, including what gets harder +## References code, data, prior art +``` + +State evidence, not assertion. "Zero of 124 modules declare `brain` as a dependency" +outranks "the dependency rule is not followed". + +## Index + +| ADR | Title | Status | +|-----|-------|--------| +| [0001](0001-mesh-brokers-nodes-host-agents-think.md) | The mesh brokers capabilities; nodes host; agents think | Accepted |