Four things settled by talking them through, all of which had been true in somebody's head and written nowhere. It is not a mesh in the peer-to-peer sense and will not become one. 0001 now says what it is instead: machines linked by a private network, one node holding knowledge of all of them, modules as the way anything is built and delivered, and agents hired onto nodes to do the work. The word describes what machines can reach, not how they are governed. "Master" overstates it the other way -- nothing needs that node to keep running, only to change. 0006 gains the option that would make it a real mesh, recorded as considered rather than rejected by silence: every node holding the whole inventory, a replication process, an elected master with promotion on failure. What settles it is not the complexity but that it still would not deliver the name, because application databases are not replicated -- so a genuine peer-to-peer mesh means becoming a replicated database system for every consumer's data too. That is a larger product than the thing it would support. Also in 0006: three central roles, not one. Losing the control plane costs change, losing the broker costs being told anything, and losing the hub costs nodes in different places reaching each other at all -- which is operation, not administration. Whether they are one node is not decided. And SSH access is identity's. It appeared three times as something that uses the overlay and never as something the mesh provides, which reads as settled when nothing decided it. Nobody else could: the mesh is the only thing that knows which humans and agents exist and which nodes they may reach. Node to node SSH stays out -- the host has no inbound control surface by decision, and nodes reaching each other that way is a second control path through the back door. 0007 gains the requirement underneath all of it. Reachability was recorded as a fact to track and never as a thing some node must have. The broker's node and the hub must be dialable by every node at a stable address, or nothing can join and a disconnected node cannot return. A mesh entirely behind NAT cannot be raised. That is a precondition and it belongs with the others. The link staying on the underlay is also argued now rather than asserted. At join time it is forced; afterwards it is a choice, and the reason is that a repair channel carried over the thing being repaired is not one. Moving it onto the overlay, with fallback, is recorded as open with what it would have to get right -- a WireGuard interface has no link state to test, and a silent fallback is this repository's recurring fault in a new place. 0010 says in one line what was the intention throughout: the module system is the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are one reconciliation seen at four points, which is why a thing that cannot be a module cannot be delivered.
11 KiB
topic, status, date, deciders, reconstructed
| topic | status | date | deciders | reconstructed |
|---|---|---|---|---|
| the mesh | accepted | 2026-08-22 | jochen | false |
1. The mesh brokers capabilities; nodes host; agents think
Context
HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split because splitting is the only granularity the platform offers, and domains are merged because a shared database is the only integration it offers. Both pressures push in the same direction: boundaries end up drawn by deployment accident rather than by domain.
Three observations establish the state.
The word "agent" means two different things. The original design treated a node as an agent with thinking abilities. Later, noxflow implemented agents as employees with skills, workspaces and tasks. Both survive. The collision is visible in the data — there are two agent rows per node:
| agent | skills |
|---|---|
| one named after the node | {deploy,verify,operate} |
one named hal-<node> |
{} |
One carries the work; the other carries only identity, existing to hold a licence for
HAL's own sessions. The same split appears in the schema: nodes.hal_claude_account and
agents.claude_account are one fact in two tables in two databases, and what was
documented as "three licence touchpoints" is one concept modelled three times.
Domains integrate by sharing a schema. noxflow is 45 tables spanning five domains —
tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose
is 7 of 45. Because agent identity lives in noxflow's database, work that belongs
elsewhere must be implemented there: per-agent Claude credentials had to be written by
the noxflow runtime, even though node identity — the same kind of fact — lives in the
mesh registry.
The core domain has no context to live in. The mesh's distinguishing feature is that
a module declares requires: postgres/database and never learns where the database
lives, who owns the credential, or how it rotates. That is capability brokering, and it
is what makes this a mesh rather than four machines with a configuration manager. Yet the
logic implementing it sits in modules/postgres/tools/index.ts — a provider module,
where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its
invariants were found violated simultaneously (see Consequences).
Considered Options
-
Keep the current layout; fix bugs as they surface. Rejected. The faults are not independent. Every incident on 2026-08-22 — credentials written by the wrong module, a dependency question answered wrongly twice, an invariant enforced nowhere — traced to a boundary that was never stated. Fixing them individually leaves the generator intact.
-
Merge aggressively into few large modules. Fewer names, same problem: a shared schema across domains is what produced noxflow, and doing it deliberately would produce it again at larger scale.
-
Decompose by bounded context, with the mesh as a broker. Name contexts after their aggregates, integrate through a published record rather than a shared schema, and let deployment granularity be a feature-level concern rather than a reason to create a module. Adopted.
Decision
The domain, in one sentence
The mesh brokers capabilities. Nodes are places where work runs. Agents are personas that think and act. Everything else supports one of those three.
What this is, plainly — and what "mesh" does not mean
Written 2026-08-29, from working through connectivity and asking whether the word still fits.
Four layers. Naming them honestly is worth more than the word on the tin:
| machines are linked by a private network | and every machine reaches every other over it |
| one node holds knowledge of all of them | the control plane, and only it |
| modules are how anything is built and delivered | this is the CI/CD, not something beside it (ADR 0010) |
| agents are hired onto nodes and do the work | the layer the other three exist to carry |
This is not a mesh in the peer-to-peer sense and will not become one. The word describes what machines can reach, not how they are governed:
| a mesh? | |
|---|---|
| what a machine can reach | yes — genuinely any to any |
| how the traffic travels | no — anything crossing sites transits the hub |
| who decides | no. One node, declared |
And master overstates it in the other direction. A master implies the others need it in order to function. They do not: every node holds what it was last told and runs from that copy always — not as a fallback, as the only mode it has. So the control plane being gone is every node in the ordinary disconnected situation at once, and what is lost is change, not operation.
The accurate phrase is one authority, no failover, and both halves are deliberate (ADR 0006).
Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no resident agent, no node-owned identity, no ownership in either direction. Agents named after a node remain, as ordinary agents that happen to hold infra skills.
Consequently nodes.node_license and nodes.hal_claude_account cease to exist: a node
does not authenticate to a model provider, agents do. The two agent rows per node merge.
There is one kind of participant, and some are human
Human and non-human participants are both agents. Both hold identity and credentials; both act, remember and coordinate. What differs is modality — how an agent acts:
| modality | credential is delivered to |
|---|---|
| spawned session | that agent's own config directory |
| shell or desktop | that agent's home on the node it acts from |
A node holds no licence. An agent holds credentials, and delivery follows that agent's node bindings and modality. A human agent's grant arrives in the home directory of the user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's grant in its config directory, with a different target.
This requires one fact the mesh does not record today: which user, on which node, a given human agent acts as. Adding it is what removes the node licence — the node is currently standing in for an identity the mesh cannot name. It is also what makes a second human agent require no new mechanism.
Contexts
| context | aggregate | subdomain |
|---|---|---|
hal/mesh |
Provision, Node, Module — brokering and its bookkeeping | core |
hal/agents |
Agent — identity, licence, runs, memory, thoughts | core |
hal/work |
Task — workflows, bindings | core |
hal/stream |
Thread — mentions, messages, meetings, notifications | core |
hal/delivery |
Pipeline — jobs, artifacts, features | supporting |
hal/knowledge |
Document — spaces, revisions, review | supporting |
hal/ai |
Licence — provider grants and rotation | supporting |
hal/observability |
Check | supporting |
hal/config |
Setting — env, secrets, PKI | generic |
A brain — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is
not a module. Anatomy makes attractive names and poor boundaries; today hal/brain names
infrastructure and hal/cortex describes itself as messaging while running nowhere.
Provisioning is the core domain, not plumbing
hal/mesh is a broker; the registry is its bookkeeping. Its invariants are explicit and
owned:
- one rotation source per resource
- a credential change fans out to every consumer
- a consumer never holds a credential the provider does not know about
Contexts integrate through the record, never a shared schema
hal/stream is the published language. A context publishes; it does not join across a
boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a
notification is a mention not yet read.
Third-party software leaves the repository
plex, sonarr, postgres, verdaccio, docker-registry and 88 others run on the
mesh; they are not of it. The pipeline and provisioning are deliberately
module-agnostic, so HAL's own modules dogfood exactly what external modules use — which
is what makes the separation safe rather than merely tidy.
Consequences
The invariants now have an owner, and were measurably unowned before. On 2026-08-22,
provision_ensure — documented as "NEVER rotates an existing secret" — was found to mint
a new password on every adoption and update only the provider's row. hal_notifications
consumers on three nodes held dead credentials for two days; hal_transcripts had two
rows written 216 ms apart, so at most one could match the live role.
noxflow dissolves. hal/work inherits tasks and workflows — the concept it was built
for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes
away.
hal/sdk shrinks. 155 files, 34,636 lines, containing code from every context —
including workflow-engine.ts and task-commands.ts, work-domain logic in the kernel
every module imports. Each landed there to avoid a cycle between modules that both needed
it; a domain module can only own its shared code once the domain has a module. Extraction
is therefore downstream of this decision, not independent of it.
Two mechanisms are prerequisites, not follow-ups.
- Named features with per-node opt-in. Without it, every independently deployable unit
inside a context becomes a module again and the count returns.
hal/claude-licencesexists solely because one daemon must run on one node. - A local mesh in containers.
dev_upstarts providers "via systemctl (same as production)" — it borrows the host, and no mesh can be stood up locally. Everything that manifests between nodes is therefore discoverable only in production, which is where every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable.
Credential delivery becomes uniform. Per-agent credential directories, built for spawned sessions, extend to human agents unchanged. A file previously scoped to a node becomes scoped to an agent — which would have prevented the class of failure where a rotation reached one node of four while the mesh reported success.
Migration is incremental and long. Contexts can be extracted one at a time behind the existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
References
01-RESEARCH/001-module-domain-decomposition— current-state evidence, table counts, open questions00-META/how-we-build.md— naming and integration rulesmodules/hal/sdk/src/feature-handlers/index.ts—FEATURE_HANDLERS, the fixed handler array that makes a feature a singleton per modulemodules/postgres/tools/index.ts— the adoption path that rotates a shared credential- Mediahuis
papa-hq, ADR 0020 Composable, independently-shippable modules — the constraints that make a unit independently shippable, applicable unchanged to features - impire.io / soulstream — the record as integration substrate, personas over services, and "cheap awareness and expensive thinking"