Files
hq/02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
T
jschoubben 88ba81e9c1 Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a
mesh function" and "a second control path through the back door", reasoning
from ADR 0004's rule that the host has no inbound control surface. That
conflated two different things and got the product backwards.

There is no node-to-node SSH to forbid. The actor is always an agent; a node is
only where it happens to be running -- ADR 0001 already says a node is a place
where an agent can run and that is the entire relationship. An agent hired onto
one node reaching another to do work is the capability the whole arrangement
exists to provide.

The credential is the agent's, in its own credential directory, which ADR 0001
already established. So a node's authorized_keys lists agents and never nodes,
and three things follow: no node holds a key reaching another node, so 0004's
"a node holds its own identity and nothing else" stays literally true; a
compromised node costs the credentials of the agents that were on it rather
than a way into everything; and who may reach what stays a mesh-wide fact,
which is why it is identity's.

The rule I misapplied is about how a node's declared state changes -- over the
broker, never by being dialled. An agent with a shell is not the mesh
reconfiguring a machine, it is what a person with a terminal has always been,
and this design already depends on that working: the overlay is the way back in
when a declaration breaks something. What such a session leaves behind is
drift, and drift is what reconciliation is for.

0001 also stops underselling the fourth layer. It read as "the layer the other
three exist to carry", which is true and flat. The value is that an agent can
work across a set of machines as though they were one -- centrally configurable
machines are ordinary; that is not.
2026-08-29 13:06:40 +02:00

231 lines
12 KiB
Markdown

---
topic: the mesh
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 1. The mesh brokers capabilities; nodes host; agents think
## Context
HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split
because splitting is the only granularity the platform offers, and domains are merged
because a shared database is the only integration it offers. Both pressures push in the
same direction: boundaries end up drawn by deployment accident rather than by domain.
Three observations establish the state.
**The word "agent" means two different things.** The original design treated a node as an
agent with thinking abilities. Later, noxflow implemented agents as employees with
skills, workspaces and tasks. Both survive. The collision is visible in the data — there
are **two agent rows per node**:
| agent | skills |
|---|---|
| one named after the node | `{deploy,verify,operate}` |
| one named `hal-<node>` | `{}` |
One carries the work; the other carries only identity, existing to hold a licence for
HAL's own sessions. The same split appears in the schema: `nodes.hal_claude_account` and
`agents.claude_account` are one fact in two tables in two databases, and what was
documented as "three licence touchpoints" is one concept modelled three times.
**Domains integrate by sharing a schema.** `noxflow` is 45 tables spanning five domains —
tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose
is 7 of 45. Because agent identity lives in noxflow's database, work that belongs
elsewhere must be implemented there: per-agent Claude credentials had to be written by
the noxflow runtime, even though node identity — the same kind of fact — lives in the
mesh registry.
**The core domain has no context to live in.** The mesh's distinguishing feature is that
a module declares `requires: postgres/database` and never learns where the database
lives, who owns the credential, or how it rotates. That is capability brokering, and it
is what makes this a mesh rather than four machines with a configuration manager. Yet the
logic implementing it sits in `modules/postgres/tools/index.ts` — a provider module,
where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its
invariants were found violated simultaneously (see Consequences).
## Considered Options
1. **Keep the current layout; fix bugs as they surface.** Rejected. The faults are not
independent. Every incident on 2026-08-22 — credentials written by the wrong module, a
dependency question answered wrongly twice, an invariant enforced nowhere — traced to
a boundary that was never stated. Fixing them individually leaves the generator intact.
2. **Merge aggressively into few large modules.** Fewer names, same problem: a shared
schema across domains is what produced noxflow, and doing it deliberately would
produce it again at larger scale.
3. **Decompose by bounded context, with the mesh as a broker.** Name contexts after their
aggregates, integrate through a published record rather than a shared schema, and let
deployment granularity be a feature-level concern rather than a reason to create a
module. **Adopted.**
## Decision
### The domain, in one sentence
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### What this is, plainly — and what "mesh" does not mean
*Written 2026-08-29, from working through connectivity and asking whether the word still fits.*
Four layers. Naming them honestly is worth more than the word on the tin:
| | |
|---|---|
| **machines are linked by a private network** | and every machine reaches every other over it |
| **one node holds knowledge of all of them** | the control plane, and only it |
| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) |
| **agents are hired onto nodes and do the work** | the layer the other three exist to carry |
**The fourth row is where the value is, and the first three are what make it possible.** An agent
hired onto one node can reach any other — a shell, a service, a file — because the private network
makes every node reachable and `identity` decides which agents may reach which. **That is the
capability being built**: not machines that can be configured centrally, which is ordinary, but a
set of machines an agent can work across as though they were one.
The credential belongs to the **agent**, never to the node it is sitting on
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), which is the same rule as *nodes and
agents are decoupled* below, applied to access.
**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what
machines can reach, not how they are governed:
| | a mesh? |
|---|---|
| what a machine can reach | **yes** — genuinely any to any |
| how the traffic travels | no — anything crossing sites transits the hub |
| who decides | no. One node, declared |
**And *master* overstates it in the other direction.** A master implies the others need it in
order to function. They do not: every node holds what it was last told and runs from that copy
*always* — not as a fallback, as the only mode it has. So the control plane being gone is every
node in the ordinary disconnected situation at once, and **what is lost is change, not
operation.**
The accurate phrase is **one authority, no failover**, and both halves are deliberate
([ADR 0006](0006-the-substrate-and-the-control-plane.md)).
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
resident agent, no node-owned identity, no ownership in either direction. Agents named
after a node remain, as **ordinary agents** that happen to hold infra skills.
Consequently `nodes.node_license` and `nodes.hal_claude_account` cease to exist: a node
does not authenticate to a model provider, agents do. The two agent rows per node merge.
### There is one kind of participant, and some are human
Human and non-human participants are both **agents**. Both hold identity and credentials;
both act, remember and coordinate. What differs is **modality** — how an agent acts:
| modality | credential is delivered to |
|---|---|
| spawned session | that agent's own config directory |
| shell or desktop | that agent's home on the node it acts from |
A node holds no licence. **An agent holds credentials, and delivery follows that agent's
node bindings and modality.** A human agent's grant arrives in the home directory of the
user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's
grant in its config directory, with a different target.
This requires one fact the mesh does not record today: which user, on which node, a given
human agent acts as. Adding it is what removes the node licence — the node is currently
standing in for an identity the mesh cannot name. It is also what makes a second human
agent require no new mechanism.
### Contexts
| context | aggregate | subdomain |
|---|---|---|
| `hal/mesh` | **Provision**, Node, Module — brokering and its bookkeeping | core |
| `hal/agents` | **Agent** — identity, licence, runs, memory, thoughts | core |
| `hal/work` | **Task** — workflows, bindings | core |
| `hal/stream` | **Thread** — mentions, messages, meetings, notifications | core |
| `hal/delivery` | **Pipeline** — jobs, artifacts, features | supporting |
| `hal/knowledge` | **Document** — spaces, revisions, review | supporting |
| `hal/ai` | **Licence** — provider grants and rotation | supporting |
| `hal/observability` | **Check** | supporting |
| `hal/config` | **Setting** — env, secrets, PKI | generic |
A *brain* — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is
not a module. Anatomy makes attractive names and poor boundaries; today `hal/brain` names
infrastructure and `hal/cortex` describes itself as messaging while running nowhere.
### Provisioning is the core domain, not plumbing
`hal/mesh` is a broker; the registry is its bookkeeping. Its invariants are explicit and
owned:
- one rotation source per resource
- a credential change fans out to every consumer
- a consumer never holds a credential the provider does not know about
### Contexts integrate through the record, never a shared schema
`hal/stream` is the published language. A context publishes; it does not join across a
boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a
notification is a mention not yet read.
### Third-party software leaves the repository
`plex`, `sonarr`, `postgres`, `verdaccio`, `docker-registry` and 88 others run **on** the
mesh; they are not **of** it. The pipeline and provisioning are deliberately
module-agnostic, so HAL's own modules dogfood exactly what external modules use — which
is what makes the separation safe rather than merely tidy.
## Consequences
**The invariants now have an owner, and were measurably unowned before.** On 2026-08-22,
`provision_ensure` — documented as "NEVER rotates an existing secret" — was found to mint
a new password on every adoption and update only the provider's row. `hal_notifications`
consumers on three nodes held dead credentials for two days; `hal_transcripts` had two
rows written 216 ms apart, so at most one could match the live role.
**noxflow dissolves.** `hal/work` inherits tasks and workflows — the concept it was built
for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes
away.
**`hal/sdk` shrinks.** 155 files, 34,636 lines, containing code from every context —
including `workflow-engine.ts` and `task-commands.ts`, work-domain logic in the kernel
every module imports. Each landed there to avoid a cycle between modules that both needed
it; a domain module can only own its shared code once the domain has a module. Extraction
is therefore downstream of this decision, not independent of it.
**Two mechanisms are prerequisites, not follow-ups.**
- *Named features with per-node opt-in.* Without it, every independently deployable unit
inside a context becomes a module again and the count returns. `hal/claude-licences`
exists solely because one daemon must run on one node.
- *A local mesh in containers.* `dev_up` starts providers "via systemctl (same as
production)" — it borrows the host, and no mesh can be stood up locally. Everything that
manifests **between** nodes is therefore discoverable only in production, which is where
every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable.
**Credential delivery becomes uniform.** Per-agent credential directories, built for
spawned sessions, extend to human agents unchanged. A file previously scoped to a node
becomes scoped to an agent — which would have prevented the class of failure where a
rotation reached one node of four while the mesh reported success.
**Migration is incremental and long.** Contexts can be extracted one at a time behind the
existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
## References
- [`01-RESEARCH/001-module-domain-decomposition`](../01-RESEARCH/001-module-domain-decomposition/analysis.md)
— current-state evidence, table counts, open questions
- [`00-META/how-we-build.md`](../00-META/how-we-build.md) — naming and integration rules
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
array that makes a feature a singleton per module
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
- Mediahuis `papa-hq`, ADR 0020 *Composable, independently-shippable modules* — the
constraints that make a unit independently shippable, applicable unchanged to features
- impire.io / soulstream — *the record* as integration substrate, personas over services,
and "cheap awareness and expensive thinking"