Files
hq/01-RESEARCH/001-module-domain-decomposition/analysis.md
T
jschoubben cf9357e8e9 HQ — the mesh's own documentation
What the mesh is, what it is becoming, and why. Implementation lives in the
code repositories; the reasoning lives here.

  00-GENESIS   mission, engineering context, effect, and the rules that hold
  01-RESEARCH  investigations, before they harden into design
  02-DESIGN    the authoritative specification
  adr          numbered decisions — what was chosen, and what was rejected
  DECISIONS.md the ledger: every decision, in the order it was taken

Written for a reader who is not its author and has no access to the mesh it
describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918;
nodes are named by role.

Single initial commit by intent. The prior history came from a private
repository and carried operational detail — a routable address identified as a
VPN hub, real domain names, a hosting provider — which sanitising a tip commit
would not have removed from the log.
2026-08-22 22:01:32 +02:00

238 lines
12 KiB
Markdown

# Current state → ideal state
## 1. The count is a symptom
124 modules. The number is not the problem; the reason for it is. Modules are split
because splitting is the **only granularity lever the platform offers**:
| symptom | module | why it exists |
|---|---|---|
| one daemon must run on one node | `hal/claude-licences` | no per-node feature opt-in |
| a library must not drag a 450 MB dep | `hal/claude` | no way to expose two npm packages |
| a domain needs a daemon *and* a library | `hal/claude-code` + `hal/claude` | one feature of each type per module |
| a handful of SDK verbs need a home | `infra` (10 lines) | tools must belong to *some* module |
A feature is a **handler type**, discovered by directory presence, singleton per module
(`FEATURE_HANDLERS` in `modules/hal/sdk/src/feature-handlers/index.ts`). So "one more
deployable unit" always means "one more module".
Splitting also fragments domains. Claude licence logic sits in three modules —
resolution in `hal/claude`, health in `hal/claude-code`, refresh in
`hal/claude-licences` — so following one token means reading three.
## 2. Two things are conflated in one repo
`noxflow` is 45 tables. Its original purpose — a HAL-native Jira — is 7 of them:
| domain | tables |
|---|---:|
| tasks (the original concept) | 7 |
| agents / HR | 13 |
| knowledge (a Confluence) | 8 |
| meetings | 7 |
| scheduling / ops | 2 |
| other | 8 |
Agents alone are nearly double the concept the module was built for. This is the direct
cause of concrete faults: per-agent Claude credentials had to be written by the noxflow
runtime purely because `agents.claude_account` lives in noxflow's database, while node
identity — the same kind of fact — lives in the mesh registry.
Separately, 91 of the 124 modules are third-party software (`plex`, `sonarr`,
`firefox`, `postgres`). These **run on the mesh; they are not part of it**. They belong
outside this repository. The pipeline and provisioning system are deliberately generic,
so HAL's own modules dogfood exactly what external modules use — that property is what
makes the separation safe.
## 3. Knowledge is split three ways
| store | owner | size | search |
|---|---|---:|---|
| `mesh_docs` (+ revisions) | `hal/hippocampus`, in the mesh DB | 144 docs, 73 touched in 30d | yes |
| `knowledge_*` (8 tables) | `noxflow/runtime`, in the noxflow DB | 20 pages | yes |
| repo markdown | git | `VISION.md`, `FLAVOR-PROPOSAL.md`, `docs/` | **no** |
The librarian is split across two of them: `librarian_file_document` writes via
hippocampus, `librarian_ask` and `librarian_retrieve` read via noxflow. Filing and
retrieval, different modules, different databases.
Both DB stores have revisions; only noxflow has a review workflow (`_proposals`,
`_promotion_requests`). Neither has file access. Git has file access and history but no
search integration. **Every piece needed already exists — none of them together.**
## 4. `hal/developer` and the missing local mesh
`hal/developer` (≈2,600 lines) exposes `dev_stage`, `dev_diff`, `dev_validate`,
`dev_typecheck`, `dev_verify_flavor`, `dev_bootstrap`, `dev_build`, `dev_up`, `dev_down`,
`dev_env`. It is the tooling that both reflects and enforces the module model, so it
moves with any change to that model.
`dev_up` builds a dev environment **for a single module**: resolves `requires:`, starts
providers "via systemctl (same as production)", provisions workspace-namespaced
databases. It borrows the host node's services.
**There is no way to run a mesh locally.** Consequently anything that only manifests
*between* nodes — credential rotation reaching a running session, provision credentials
fanning out, cascade ordering, artifact packaging — is discoverable only in production.
Every fault fixed on 2026-08-22 was of that kind.
A containerised mesh (N nodes, a broker, a registry, a coordinator) is therefore not a
convenience. It is the precondition for a refactor of this size being verifiable at all.
## 5. `hal/sdk` is where domains go to hide
155 files, 34,636 lines. It contains code from essentially every bounded context:
| file | domain it belongs to |
|---|---|
| `installer-core.ts` (1,495) | runtime |
| `tools/provisions.ts` (1,311) | provisioning |
| `tools/deploy-service.ts` (1,081) | deploy |
| `workflow-engine.ts` (970), `task-commands.ts` (588) | **work** — noxflow's domain, inside HAL's SDK |
| `artifact-manager.ts` (948), `build-executor.ts` (760), `feature-handlers/*` | delivery |
| `env-generator.ts` (944) | config |
| `amqp-client.ts` (899) | comms |
| `meshware.ts` (799) | runtime |
| `module-registry.ts` (743) | mesh |
| `claude-credentials.ts` | **ai** |
Nothing here was misplaced carelessly. Each landed in the SDK because the SDK is the one
package every module may import — so putting a shared type there is the way to avoid a
circular dependency between two modules that both need it.
The result is that the dependency graph is trivially acyclic and completely
uninformative: everything depends on one 34k-line package, so a change anywhere in it
rebuilds everything. On 2026-08-22 a three-line fix in `feature-handlers/migrations.ts`
produced a 1,445-job, 26-minute mesh-wide cascade.
The last row is from that same day, and is the clearest illustration: per-agent Claude
config-directory helpers were added to `@hal/sdk/claude-credentials.ts` because it was
the only place both `hal/claude` and the noxflow runtime could import from. The
alternative — `hal/ai` owning it and noxflow depending on `hal/ai` — was unavailable
without answering the ownership question this effort exists to answer.
**Extraction is therefore blocked on the decomposition, not the reverse.** A domain
module can only own its shared code once the domain has a module. The SDK should retain
only what is genuinely cross-cutting: transport, manifest types, logging, process
helpers.
## 6. The crux: two notions of "agent"
The original vision was **a node is an agent with thinking abilities**. Later, noxflow
implemented actual agents — employees with skills, workspaces and tasks. Both survive,
and they collide.
The collision is visible in the data. There are **two agent rows per node**:
| agent | display name | skills | bound to |
|---|---|---|---|
| one named after each node | the node's name | `{deploy,verify,operate}` | its node |
| `hal-<node>`, one per node | "HAL (\<node\>)" | `{}` | its node |
One carries the **work**; the other carries only **identity** — no skills at all, existing
to hold a licence for HAL's own sessions (`resolveHalModuleEnv`). The node-as-agent idea
was implemented twice, half each.
The same split appears in the schema: `nodes.hal_claude_account` and
`agents.claude_account` are the same fact in two tables in two databases. What was
documented as "three licence touchpoints" (NODE, HAL, AGENT) is **one concept modelled
three times**.
### Resolution
**Nodes and agents are decoupled.** The agents module owns agents; agents run on nodes.
A node is a place where an agent can run — that is the whole of the relationship. There
is no resident agent, no node-owned identity, no ownership in either direction.
Agents named after a node still exist, but they are **ordinary agents defined through the
agents module** that happen to hold infra skills and be named after the node they usually run
on. Nothing about them is structural.
Two consequences follow:
1. **`nodes.node_license` and `nodes.hal_claude_account` should not exist.** A node does
not authenticate to a model provider — agents do. Both collapse into
`agents.claude_account`, and what was documented as three licence touchpoints becomes
one.
2. **The pair of agent rows per node merges.** `hal-<node>` exists only to hold a licence
for HAL's own sessions; once licences belong to agents and any module may employ an
agent, it has no reason to be separate from `<node>`.
What remains unaccounted for is the operator's own interactive session:
`~/.claude/.credentials.json` on a node serves a **human at a prompt**, not an agent.
Under this decoupling that is not a node property either — it is the operator's
credential on whichever machine they are sitting at. Whether the operator is modelled as
a persona (soulstream treats humans and agents as peers with identical credentials) is an
open question, and the last remaining reason `nodes` carries a licence column today.
## 7. The record, not "meetings"
soulstream models work as **one signed log** in which humans and agents apply identical
changes, with work arriving by **mention**. Memory and decisions live in that single
auditable stream — no hidden state between components. Its stated economics,
*"cheap awareness and expensive thinking"*, is the same principle as the budget guard:
constant monitoring, selective deep processing.
Under that model "meetings" is not a domain. A meeting is a **thread in the record**.
Which means six of today's modules and table groups are one thing seen from different
angles:
`hal/axon` (DMs, rooms) · `hal/cortex` · `hal/synapse` · `hal/notifications` ·
`conversations` · `meetings`
A notification, in particular, is just a mention not yet read — which is why it needed
its own table, its own id column, and its own failure mode.
## 8. Ideal state
The mesh repository contains only the mesh, decomposed by domain rather than by
deployment accident:
| context | what it is | absorbs today's |
|---|---|---|
| `hal/infra` | **the room** — nodes, registry, provisioning, pipeline, node runtime | `hal/mesh`, `hal/brain`, `hal/coordinator`, `hal/meshware`, `hal/developer`, `hal/bootstrap` |
| `hal/agents` | **the name and the thinking** — Agent owns identity, licence, runs, memory, thoughts | noxflow agents (13 tables), `hal/thoughts` |
| `hal/stream` | **the record** — mentions, threads, meetings, notifications | `hal/axon`, `hal/cortex`, `hal/synapse`, `hal/notifications`, meetings (7), conversations |
| `hal/knowledge` | documents, spaces, revisions, review, search | `hal/hippocampus` + noxflow `knowledge_*` |
| `hal/work` | tasks, workflows, boards — the original noxflow | noxflow tasks (7 tables) |
| `hal/ai` | provider integration, flavored per vendor | `hal/claude*`, `claude-code` |
| `hal/config` | env and config distribution, secrets, PKI | `hal/env-sync`, `hal/config-sync`, `hal/secrets`, `mesh-ca` |
| `hal/observability` | health, dashboards, logs | `hal/health`, `hal/meshboard` |
Eight contexts instead of 33 platform modules, with the catalogue elsewhere.
The naming corrects itself in the process: what is called "brain" today is
infrastructure, and the word survives only as a concept an Agent owns — never again as a
module name.
This shape depends on **named features with per-node opt-in** — without it, every
independently-deployable unit inside a context becomes a module again and the count
returns. That mechanism is the subject of a separate ADR.
## Decided since
| question | decision |
|---|---|
| Meetings — own context or part of agents? | **Neither.** A meeting is a thread in the record; `hal/stream` absorbs it. The category was wrong. |
| Thoughts — autonomy or agent behaviour? | `hal/agents`. Any agent can have a thought. |
| `hal/axon` vs `hal/cortex` | `hal/cortex` is vestigial — no `daemon/` source, unit inactive on every node, description copy-pasted from axon. Messaging belongs to `hal/stream`. |
| `hal/hypothalamus` | Orphaned build output. A `dist/` directory with no `module.yml`, so not a module at all. |
| Who owns agent identity? | `hal/agents`. Nodes and agents are decoupled; a node is only a place an agent can run. |
| Where do third-party modules live? | Outside this repository. They run *on* the mesh, not *of* it. |
## Open questions
1. **`hal/scheduler`** — infrastructure (fire an event later, belongs to `hal/mesh`) or
part of `hal/work`?
2. **The executor.** Pulling and routing work items is `hal/work`; spawning a session is
`hal/agents`; providing the place to run is `hal/mesh`. It currently straddles all
three, which is the same kind of straddle that put credential-writing in the noxflow
runtime.
3. **Catalogue destination** — one repository, or per-application repositories, given the
pipeline resolves dependencies across the registry rather than the filesystem?
4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or
does that belong to `hal/stream`? Everything imports it, which argues both ways.
5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which
user, on which node, a human agent acts as. Where does it live — an attribute of the
agent, or of the agent-node binding?