Files
hq/01-RESEARCH/001-module-domain-decomposition/analysis.md
T
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00

243 lines
12 KiB
Markdown

---
effort: 001-module-domain-decomposition
updated: 2026-08-22
---
# Current state → ideal state
## 1. The count is a symptom
124 modules. The number is not the problem; the reason for it is. Modules are split
because splitting is the **only granularity lever the platform offers**:
| symptom | module | why it exists |
|---|---|---|
| one daemon must run on one node | `hal/claude-licences` | no per-node feature opt-in |
| a library must not drag a 450 MB dep | `hal/claude` | no way to expose two npm packages |
| a domain needs a daemon *and* a library | `hal/claude-code` + `hal/claude` | one feature of each type per module |
| a handful of SDK verbs need a home | `infra` (10 lines) | tools must belong to *some* module |
A feature is a **handler type**, discovered by directory presence, singleton per module
(`FEATURE_HANDLERS` in `modules/hal/sdk/src/feature-handlers/index.ts`). So "one more
deployable unit" always means "one more module".
Splitting also fragments domains. Claude licence logic sits in three modules —
resolution in `hal/claude`, health in `hal/claude-code`, refresh in
`hal/claude-licences` — so following one token means reading three.
## 2. Two things are conflated in one repo
`noxflow` is 45 tables. Its original purpose — a HAL-native Jira — is 7 of them:
| domain | tables |
|---|---:|
| tasks (the original concept) | 7 |
| agents / HR | 13 |
| knowledge (a Confluence) | 8 |
| meetings | 7 |
| scheduling / ops | 2 |
| other | 8 |
Agents alone are nearly double the concept the module was built for. This is the direct
cause of concrete faults: per-agent Claude credentials had to be written by the noxflow
runtime purely because `agents.claude_account` lives in noxflow's database, while node
identity — the same kind of fact — lives in the mesh registry.
Separately, 91 of the 124 modules are third-party software (`plex`, `sonarr`,
`firefox`, `postgres`). These **run on the mesh; they are not part of it**. They belong
outside this repository. The pipeline and provisioning system are deliberately generic,
so HAL's own modules dogfood exactly what external modules use — that property is what
makes the separation safe.
## 3. Knowledge is split three ways
| store | owner | size | search |
|---|---|---:|---|
| `mesh_docs` (+ revisions) | `hal/hippocampus`, in the mesh DB | 144 docs, 73 touched in 30d | yes |
| `knowledge_*` (8 tables) | `noxflow/runtime`, in the noxflow DB | 20 pages | yes |
| repo markdown | git | `VISION.md`, `FLAVOR-PROPOSAL.md`, `docs/` | **no** |
The librarian is split across two of them: `librarian_file_document` writes via
hippocampus, `librarian_ask` and `librarian_retrieve` read via noxflow. Filing and
retrieval, different modules, different databases.
Both DB stores have revisions; only noxflow has a review workflow (`_proposals`,
`_promotion_requests`). Neither has file access. Git has file access and history but no
search integration. **Every piece needed already exists — none of them together.**
## 4. `hal/developer` and the missing local mesh
`hal/developer` (≈2,600 lines) exposes `dev_stage`, `dev_diff`, `dev_validate`,
`dev_typecheck`, `dev_verify_flavor`, `dev_bootstrap`, `dev_build`, `dev_up`, `dev_down`,
`dev_env`. It is the tooling that both reflects and enforces the module model, so it
moves with any change to that model.
`dev_up` builds a dev environment **for a single module**: resolves `requires:`, starts
providers "via systemctl (same as production)", provisions workspace-namespaced
databases. It borrows the host node's services.
**There is no way to run a mesh locally.** Consequently anything that only manifests
*between* nodes — credential rotation reaching a running session, provision credentials
fanning out, cascade ordering, artifact packaging — is discoverable only in production.
Every fault fixed on 2026-08-22 was of that kind.
A containerised mesh (N nodes, a broker, a registry, a coordinator) is therefore not a
convenience. It is the precondition for a refactor of this size being verifiable at all.
## 5. `hal/sdk` is where domains go to hide
155 files, 34,636 lines. It contains code from essentially every bounded context:
| file | domain it belongs to |
|---|---|
| `installer-core.ts` (1,495) | runtime |
| `tools/provisions.ts` (1,311) | provisioning |
| `tools/deploy-service.ts` (1,081) | deploy |
| `workflow-engine.ts` (970), `task-commands.ts` (588) | **work** — noxflow's domain, inside HAL's SDK |
| `artifact-manager.ts` (948), `build-executor.ts` (760), `feature-handlers/*` | delivery |
| `env-generator.ts` (944) | config |
| `amqp-client.ts` (899) | comms |
| `meshware.ts` (799) | runtime |
| `module-registry.ts` (743) | mesh |
| `claude-credentials.ts` | **ai** |
Nothing here was misplaced carelessly. Each landed in the SDK because the SDK is the one
package every module may import — so putting a shared type there is the way to avoid a
circular dependency between two modules that both need it.
The result is that the dependency graph is trivially acyclic and completely
uninformative: everything depends on one 34k-line package, so a change anywhere in it
rebuilds everything. On 2026-08-22 a three-line fix in `feature-handlers/migrations.ts`
produced a 1,445-job, 26-minute mesh-wide cascade.
The last row is from that same day, and is the clearest illustration: per-agent Claude
config-directory helpers were added to `@hal/sdk/claude-credentials.ts` because it was
the only place both `hal/claude` and the noxflow runtime could import from. The
alternative — `hal/ai` owning it and noxflow depending on `hal/ai` — was unavailable
without answering the ownership question this effort exists to answer.
**Extraction is therefore blocked on the decomposition, not the reverse.** A domain
module can only own its shared code once the domain has a module. The SDK should retain
only what is genuinely cross-cutting: transport, manifest types, logging, process
helpers.
## 6. The crux: two notions of "agent"
The original vision was **a node is an agent with thinking abilities**. Later, noxflow
implemented actual agents — employees with skills, workspaces and tasks. Both survive,
and they collide.
The collision is visible in the data. There are **two agent rows per node**:
| agent | display name | skills | bound to |
|---|---|---|---|
| one named after each node | the node's name | `{deploy,verify,operate}` | its node |
| `hal-<node>`, one per node | "HAL (\<node\>)" | `{}` | its node |
One carries the **work**; the other carries only **identity** — no skills at all, existing
to hold a licence for HAL's own sessions (`resolveHalModuleEnv`). The node-as-agent idea
was implemented twice, half each.
The same split appears in the schema: `nodes.hal_claude_account` and
`agents.claude_account` are the same fact in two tables in two databases. What was
documented as "three licence touchpoints" (NODE, HAL, AGENT) is **one concept modelled
three times**.
### Resolution
**Nodes and agents are decoupled.** The agents module owns agents; agents run on nodes.
A node is a place where an agent can run — that is the whole of the relationship. There
is no resident agent, no node-owned identity, no ownership in either direction.
Agents named after a node still exist, but they are **ordinary agents defined through the
agents module** that happen to hold infra skills and be named after the node they usually run
on. Nothing about them is structural.
Two consequences follow:
1. **`nodes.node_license` and `nodes.hal_claude_account` should not exist.** A node does
not authenticate to a model provider — agents do. Both collapse into
`agents.claude_account`, and what was documented as three licence touchpoints becomes
one.
2. **The pair of agent rows per node merges.** `hal-<node>` exists only to hold a licence
for HAL's own sessions; once licences belong to agents and any module may employ an
agent, it has no reason to be separate from `<node>`.
What remains unaccounted for is the operator's own interactive session:
`~/.claude/.credentials.json` on a node serves a **human at a prompt**, not an agent.
Under this decoupling that is not a node property either — it is the operator's
credential on whichever machine they are sitting at. Whether the operator is modelled as
a persona (soulstream treats humans and agents as peers with identical credentials) is an
open question, and the last remaining reason `nodes` carries a licence column today.
## 7. The record, not "meetings"
soulstream models work as **one signed log** in which humans and agents apply identical
changes, with work arriving by **mention**. Memory and decisions live in that single
auditable stream — no hidden state between components. Its stated economics,
*"cheap awareness and expensive thinking"*, is the same principle as the budget guard:
constant monitoring, selective deep processing.
Under that model "meetings" is not a domain. A meeting is a **thread in the record**.
Which means six of today's modules and table groups are one thing seen from different
angles:
`hal/axon` (DMs, rooms) · `hal/cortex` · `hal/synapse` · `hal/notifications` ·
`conversations` · `meetings`
A notification, in particular, is just a mention not yet read — which is why it needed
its own table, its own id column, and its own failure mode.
## 8. Ideal state
The mesh repository contains only the mesh, decomposed by domain rather than by
deployment accident:
| context | what it is | absorbs today's |
|---|---|---|
| `hal/infra` | **the room** — nodes, registry, provisioning, pipeline, node runtime | `hal/mesh`, `hal/brain`, `hal/coordinator`, `hal/meshware`, `hal/developer`, `hal/bootstrap` |
| `hal/agents` | **the name and the thinking** — Agent owns identity, licence, runs, memory, thoughts | noxflow agents (13 tables), `hal/thoughts` |
| `hal/stream` | **the record** — mentions, threads, meetings, notifications | `hal/axon`, `hal/cortex`, `hal/synapse`, `hal/notifications`, meetings (7), conversations |
| `hal/knowledge` | documents, spaces, revisions, review, search | `hal/hippocampus` + noxflow `knowledge_*` |
| `hal/work` | tasks, workflows, boards — the original noxflow | noxflow tasks (7 tables) |
| `hal/ai` | provider integration, flavored per vendor | `hal/claude*`, `claude-code` |
| `hal/config` | env and config distribution, secrets, PKI | `hal/env-sync`, `hal/config-sync`, `hal/secrets`, `mesh-ca` |
| `hal/observability` | health, dashboards, logs | `hal/health`, `hal/meshboard` |
Eight contexts instead of 33 platform modules, with the catalogue elsewhere.
The naming corrects itself in the process: what is called "brain" today is
infrastructure, and the word survives only as a concept an Agent owns — never again as a
module name.
This shape depends on **named features with per-node opt-in** — without it, every
independently-deployable unit inside a context becomes a module again and the count
returns. That mechanism is the subject of a separate ADR.
## Decided since
| question | decision |
|---|---|
| Meetings — own context or part of agents? | **Neither.** A meeting is a thread in the record; `hal/stream` absorbs it. The category was wrong. |
| Thoughts — autonomy or agent behaviour? | `hal/agents`. Any agent can have a thought. |
| `hal/axon` vs `hal/cortex` | `hal/cortex` is vestigial — no `daemon/` source, unit inactive on every node, description copy-pasted from axon. Messaging belongs to `hal/stream`. |
| `hal/hypothalamus` | Orphaned build output. A `dist/` directory with no `module.yml`, so not a module at all. |
| Who owns agent identity? | `hal/agents`. Nodes and agents are decoupled; a node is only a place an agent can run. |
| Where do third-party modules live? | Outside this repository. They run *on* the mesh, not *of* it. |
## Open questions
1. **`hal/scheduler`** — infrastructure (fire an event later, belongs to `hal/mesh`) or
part of `hal/work`?
2. **The executor.** Pulling and routing work items is `hal/work`; spawning a session is
`hal/agents`; providing the place to run is `hal/mesh`. It currently straddles all
three, which is the same kind of straddle that put credential-writing in the noxflow
runtime.
3. **Catalogue destination** — one repository, or per-application repositories, given the
pipeline resolves dependencies across the registry rather than the filesystem?
4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or
does that belong to `hal/stream`? Everything imports it, which argues both ways.
5. **Human agent modality.** ADR 0008 requires a fact the mesh does not record: which
user, on which node, a human agent acts as. Where does it live — an attribute of the
agent, or of the agent-node binding?