Files
hq/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
T
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00

11 KiB

topic, status, date, deciders, reconstructed
topic status date deciders reconstructed
the mesh accepted 2026-08-22 jochen false

1. The mesh brokers capabilities; nodes host; agents think

Context

HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split because splitting is the only granularity the platform offers, and domains are merged because a shared database is the only integration it offers. Both pressures push in the same direction: boundaries end up drawn by deployment accident rather than by domain.

Three observations establish the state.

The word "agent" means two different things. The original design treated a node as an agent with thinking abilities. Later, noxflow implemented agents as employees with skills, workspaces and tasks. Both survive. The collision is visible in the data — there are two agent rows per node:

agent skills
one named after the node {deploy,verify,operate}
one named hal-<node> {}

One carries the work; the other carries only identity, existing to hold a licence for HAL's own sessions. The same split appears in the schema: nodes.hal_claude_account and agents.claude_account are one fact in two tables in two databases, and what was documented as "three licence touchpoints" is one concept modelled three times.

Domains integrate by sharing a schema. noxflow is 45 tables spanning five domains — tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose is 7 of 45. Because agent identity lives in noxflow's database, work that belongs elsewhere must be implemented there: per-agent Claude credentials had to be written by the noxflow runtime, even though node identity — the same kind of fact — lives in the mesh registry.

The core domain has no context to live in. The mesh's distinguishing feature is that a module declares requires: postgres/database and never learns where the database lives, who owns the credential, or how it rotates. That is capability brokering, and it is what makes this a mesh rather than four machines with a configuration manager. Yet the logic implementing it sits in modules/postgres/tools/index.ts — a provider module, where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its invariants were found violated simultaneously (see Consequences).

Considered Options

  1. Keep the current layout; fix bugs as they surface. Rejected. The faults are not independent. Every incident on 2026-08-22 — credentials written by the wrong module, a dependency question answered wrongly twice, an invariant enforced nowhere — traced to a boundary that was never stated. Fixing them individually leaves the generator intact.

  2. Merge aggressively into few large modules. Fewer names, same problem: a shared schema across domains is what produced noxflow, and doing it deliberately would produce it again at larger scale.

  3. Decompose by bounded context, with the mesh as a broker. Name contexts after their aggregates, integrate through a published record rather than a shared schema, and let deployment granularity be a feature-level concern rather than a reason to create a module. Adopted.

Decision

The domain, in one sentence

The mesh brokers capabilities. Nodes are places where work runs. Agents are personas that think and act. Everything else supports one of those three.

What this is, plainly — and what "mesh" does not mean

Written 2026-08-29, from working through connectivity and asking whether the word still fits.

Four layers. Naming them honestly is worth more than the word on the tin:

machines are linked by a private network and every machine reaches every other over it
one node holds knowledge of all of them the control plane, and only it
modules are how anything is built and delivered this is the CI/CD, not something beside it (ADR 0010)
agents are hired onto nodes and do the work the layer the other three exist to carry

This is not a mesh in the peer-to-peer sense and will not become one. The word describes what machines can reach, not how they are governed:

a mesh?
what a machine can reach yes — genuinely any to any
how the traffic travels no — anything crossing sites transits the hub
who decides no. One node, declared

And master overstates it in the other direction. A master implies the others need it in order to function. They do not: every node holds what it was last told and runs from that copy always — not as a fallback, as the only mode it has. So the control plane being gone is every node in the ordinary disconnected situation at once, and what is lost is change, not operation.

The accurate phrase is one authority, no failover, and both halves are deliberate (ADR 0006).

Nodes and agents are decoupled

A node is a place where an agent can run — that is the entire relationship. There is no resident agent, no node-owned identity, no ownership in either direction. Agents named after a node remain, as ordinary agents that happen to hold infra skills.

Consequently nodes.node_license and nodes.hal_claude_account cease to exist: a node does not authenticate to a model provider, agents do. The two agent rows per node merge.

There is one kind of participant, and some are human

Human and non-human participants are both agents. Both hold identity and credentials; both act, remember and coordinate. What differs is modality — how an agent acts:

modality credential is delivered to
spawned session that agent's own config directory
shell or desktop that agent's home on the node it acts from

A node holds no licence. An agent holds credentials, and delivery follows that agent's node bindings and modality. A human agent's grant arrives in the home directory of the user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's grant in its config directory, with a different target.

This requires one fact the mesh does not record today: which user, on which node, a given human agent acts as. Adding it is what removes the node licence — the node is currently standing in for an identity the mesh cannot name. It is also what makes a second human agent require no new mechanism.

Contexts

context aggregate subdomain
hal/mesh Provision, Node, Module — brokering and its bookkeeping core
hal/agents Agent — identity, licence, runs, memory, thoughts core
hal/work Task — workflows, bindings core
hal/stream Thread — mentions, messages, meetings, notifications core
hal/delivery Pipeline — jobs, artifacts, features supporting
hal/knowledge Document — spaces, revisions, review supporting
hal/ai Licence — provider grants and rotation supporting
hal/observability Check supporting
hal/config Setting — env, secrets, PKI generic

A brain — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is not a module. Anatomy makes attractive names and poor boundaries; today hal/brain names infrastructure and hal/cortex describes itself as messaging while running nowhere.

Provisioning is the core domain, not plumbing

hal/mesh is a broker; the registry is its bookkeeping. Its invariants are explicit and owned:

  • one rotation source per resource
  • a credential change fans out to every consumer
  • a consumer never holds a credential the provider does not know about

Contexts integrate through the record, never a shared schema

hal/stream is the published language. A context publishes; it does not join across a boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a notification is a mention not yet read.

Third-party software leaves the repository

plex, sonarr, postgres, verdaccio, docker-registry and 88 others run on the mesh; they are not of it. The pipeline and provisioning are deliberately module-agnostic, so HAL's own modules dogfood exactly what external modules use — which is what makes the separation safe rather than merely tidy.

Consequences

The invariants now have an owner, and were measurably unowned before. On 2026-08-22, provision_ensure — documented as "NEVER rotates an existing secret" — was found to mint a new password on every adoption and update only the provider's row. hal_notifications consumers on three nodes held dead credentials for two days; hal_transcripts had two rows written 216 ms apart, so at most one could match the live role.

noxflow dissolves. hal/work inherits tasks and workflows — the concept it was built for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes away.

hal/sdk shrinks. 155 files, 34,636 lines, containing code from every context — including workflow-engine.ts and task-commands.ts, work-domain logic in the kernel every module imports. Each landed there to avoid a cycle between modules that both needed it; a domain module can only own its shared code once the domain has a module. Extraction is therefore downstream of this decision, not independent of it.

Two mechanisms are prerequisites, not follow-ups.

  • Named features with per-node opt-in. Without it, every independently deployable unit inside a context becomes a module again and the count returns. hal/claude-licences exists solely because one daemon must run on one node.
  • A local mesh in containers. dev_up starts providers "via systemctl (same as production)" — it borrows the host, and no mesh can be stood up locally. Everything that manifests between nodes is therefore discoverable only in production, which is where every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable.

Credential delivery becomes uniform. Per-agent credential directories, built for spawned sessions, extend to human agents unchanged. A file previously scoped to a node becomes scoped to an agent — which would have prevented the class of failure where a rotation reached one node of four while the mesh reported success.

Migration is incremental and long. Contexts can be extracted one at a time behind the existing pipeline. Nothing here requires a flag day, and nothing here is cheap.

References

  • 01-RESEARCH/001-module-domain-decomposition — current-state evidence, table counts, open questions
  • 00-META/how-we-build.md — naming and integration rules
  • modules/hal/sdk/src/feature-handlers/index.ts — FEATURE_HANDLERS, the fixed handler array that makes a feature a singleton per module
  • modules/postgres/tools/index.ts — the adoption path that rotates a shared credential
  • Mediahuis papa-hq, ADR 0020 Composable, independently-shippable modules — the constraints that make a unit independently shippable, applicable unchanged to features
  • impire.io / soulstream — the record as integration substrate, personas over services, and "cheap awareness and expensive thinking"