HQ — the mesh's own documentation

What the mesh is, what it is becoming, and why. Implementation lives in the
code repositories; the reasoning lives here.

  00-GENESIS   mission, engineering context, effect, and the rules that hold
  01-RESEARCH  investigations, before they harden into design
  02-DESIGN    the authoritative specification
  adr          numbered decisions — what was chosen, and what was rejected
  DECISIONS.md the ledger: every decision, in the order it was taken

Written for a reader who is not its author and has no access to the mesh it
describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918;
nodes are named by role.

Single initial commit by intent. The prior history came from a private
repository and carried operational detail — a routable address identified as a
VPN hub, real domain names, a hosting provider — which sanitising a tip commit
would not have removed from the log.
This commit is contained in:
2026-08-22 22:01:32 +02:00
commit cf9357e8e9
22 changed files with 2376 additions and 0 deletions
@@ -0,0 +1,185 @@
# 1. The mesh brokers capabilities; nodes host; agents think
- **Status:** Accepted
- **Date:** 2026-08-22
- **Deciders:** jochen
## Context
HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split
because splitting is the only granularity the platform offers, and domains are merged
because a shared database is the only integration it offers. Both pressures push in the
same direction: boundaries end up drawn by deployment accident rather than by domain.
Three observations establish the state.
**The word "agent" means two different things.** The original design treated a node as an
agent with thinking abilities. Later, noxflow implemented agents as employees with
skills, workspaces and tasks. Both survive. The collision is visible in the data — there
are **two agent rows per node**:
| agent | skills |
|---|---|
| one named after the node | `{deploy,verify,operate}` |
| one named `hal-<node>` | `{}` |
One carries the work; the other carries only identity, existing to hold a licence for
HAL's own sessions. The same split appears in the schema: `nodes.hal_claude_account` and
`agents.claude_account` are one fact in two tables in two databases, and what was
documented as "three licence touchpoints" is one concept modelled three times.
**Domains integrate by sharing a schema.** `noxflow` is 45 tables spanning five domains —
tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose
is 7 of 45. Because agent identity lives in noxflow's database, work that belongs
elsewhere must be implemented there: per-agent Claude credentials had to be written by
the noxflow runtime, even though node identity — the same kind of fact — lives in the
mesh registry.
**The core domain has no context to live in.** The mesh's distinguishing feature is that
a module declares `requires: postgres/database` and never learns where the database
lives, who owns the credential, or how it rotates. That is capability brokering, and it
is what makes this a mesh rather than four machines with a configuration manager. Yet the
logic implementing it sits in `modules/postgres/tools/index.ts` — a provider module,
where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its
invariants were found violated simultaneously (see Consequences).
## Considered Options
1. **Keep the current layout; fix bugs as they surface.** Rejected. The faults are not
independent. Every incident on 2026-08-22 — credentials written by the wrong module, a
dependency question answered wrongly twice, an invariant enforced nowhere — traced to
a boundary that was never stated. Fixing them individually leaves the generator intact.
2. **Merge aggressively into few large modules.** Fewer names, same problem: a shared
schema across domains is what produced noxflow, and doing it deliberately would
produce it again at larger scale.
3. **Decompose by bounded context, with the mesh as a broker.** Name contexts after their
aggregates, integrate through a published record rather than a shared schema, and let
deployment granularity be a feature-level concern rather than a reason to create a
module. **Adopted.**
## Decision
### The domain, in one sentence
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
resident agent, no node-owned identity, no ownership in either direction. Agents named
after a node remain, as **ordinary agents** that happen to hold infra skills.
Consequently `nodes.node_license` and `nodes.hal_claude_account` cease to exist: a node
does not authenticate to a model provider, agents do. The two agent rows per node merge.
### There is one kind of participant, and some are human
Human and non-human participants are both **agents**. Both hold identity and credentials;
both act, remember and coordinate. What differs is **modality** — how an agent acts:
| modality | credential is delivered to |
|---|---|
| spawned session | that agent's own config directory |
| shell or desktop | that agent's home on the node it acts from |
A node holds no licence. **An agent holds credentials, and delivery follows that agent's
node bindings and modality.** A human agent's grant arrives in the home directory of the
user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's
grant in its config directory, with a different target.
This requires one fact the mesh does not record today: which user, on which node, a given
human agent acts as. Adding it is what removes the node licence — the node is currently
standing in for an identity the mesh cannot name. It is also what makes a second human
agent require no new mechanism.
### Contexts
| context | aggregate | subdomain |
|---|---|---|
| `hal/mesh` | **Provision**, Node, Module — brokering and its bookkeeping | core |
| `hal/agents` | **Agent** — identity, licence, runs, memory, thoughts | core |
| `hal/work` | **Task** — workflows, bindings | core |
| `hal/stream` | **Thread** — mentions, messages, meetings, notifications | core |
| `hal/delivery` | **Pipeline** — jobs, artifacts, features | supporting |
| `hal/knowledge` | **Document** — spaces, revisions, review | supporting |
| `hal/ai` | **Licence** — provider grants and rotation | supporting |
| `hal/observability` | **Check** | supporting |
| `hal/config` | **Setting** — env, secrets, PKI | generic |
A *brain* — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is
not a module. Anatomy makes attractive names and poor boundaries; today `hal/brain` names
infrastructure and `hal/cortex` describes itself as messaging while running nowhere.
### Provisioning is the core domain, not plumbing
`hal/mesh` is a broker; the registry is its bookkeeping. Its invariants are explicit and
owned:
- one rotation source per resource
- a credential change fans out to every consumer
- a consumer never holds a credential the provider does not know about
### Contexts integrate through the record, never a shared schema
`hal/stream` is the published language. A context publishes; it does not join across a
boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a
notification is a mention not yet read.
### Third-party software leaves the repository
`plex`, `sonarr`, `postgres`, `verdaccio`, `docker-registry` and 88 others run **on** the
mesh; they are not **of** it. The pipeline and provisioning are deliberately
module-agnostic, so HAL's own modules dogfood exactly what external modules use — which
is what makes the separation safe rather than merely tidy.
## Consequences
**The invariants now have an owner, and were measurably unowned before.** On 2026-08-22,
`provision_ensure` — documented as "NEVER rotates an existing secret" — was found to mint
a new password on every adoption and update only the provider's row. `hal_notifications`
consumers on three nodes held dead credentials for two days; `hal_transcripts` had two
rows written 216 ms apart, so at most one could match the live role.
**noxflow dissolves.** `hal/work` inherits tasks and workflows — the concept it was built
for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes
away.
**`hal/sdk` shrinks.** 155 files, 34,636 lines, containing code from every context —
including `workflow-engine.ts` and `task-commands.ts`, work-domain logic in the kernel
every module imports. Each landed there to avoid a cycle between modules that both needed
it; a domain module can only own its shared code once the domain has a module. Extraction
is therefore downstream of this decision, not independent of it.
**Two mechanisms are prerequisites, not follow-ups.**
- *Named features with per-node opt-in.* Without it, every independently deployable unit
inside a context becomes a module again and the count returns. `hal/claude-licences`
exists solely because one daemon must run on one node.
- *A local mesh in containers.* `dev_up` starts providers "via systemctl (same as
production)" — it borrows the host, and no mesh can be stood up locally. Everything that
manifests **between** nodes is therefore discoverable only in production, which is where
every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable.
**Credential delivery becomes uniform.** Per-agent credential directories, built for
spawned sessions, extend to human agents unchanged. A file previously scoped to a node
becomes scoped to an agent — which would have prevented the class of failure where a
rotation reached one node of four while the mesh reported success.
**Migration is incremental and long.** Contexts can be extracted one at a time behind the
existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
## References
- [`01-RESEARCH/001-module-domain-decomposition`](../01-RESEARCH/001-module-domain-decomposition/analysis.md)
— current-state evidence, table counts, open questions
- [`00-GENESIS/how-we-build.md`](../00-GENESIS/how-we-build.md) — naming and integration rules
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
array that makes a feature a singleton per module
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the
constraints that make a unit independently shippable, applicable unchanged to features
- impire.io / soulstream — *the record* as integration substrate, personas over services,
and "cheap awareness and expensive thinking"
+135
View File
@@ -0,0 +1,135 @@
# 2. A lab node is a virtual machine running the real install
- **Status:** Accepted
- **Date:** 2026-08-22
- **Deciders:** jochen
## Context
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
Two things were measured while establishing what exists
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
- **There is no local mesh.** The dev tooling starts providers through the host's own init
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
It borrows the machine because there is nowhere else to put a mesh.
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
physical machine. Every subsequent question — how faithful is faithful enough, what may be
mocked, which failures remain reachable — follows from that one answer.
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
hardware virtualisation present.
## Considered Options
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
a node as one inverts the thing being modelled: module service stacks then require nested
containers through a privileged daemon, or a shared socket that makes isolation between
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
and the least like a node.
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
is the only option that makes a twenty-node run affordable. It was rejected because **the
scale requirement that justified it was invented rather than required**: the stated goal is
to run the real mesh, which is four nodes, on one computer. And a system container still
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
then has to be answered again for every capability under test.
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
no storage pools, no snapshot management, no network management, no virtual machines.
Snapshots are what make the loop fast, so the saving is not worth what it costs.
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
turns into a node.
## Decision
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
runs the real install, and becomes a node. It is not a model of a node, so no question arises
about how good the model is.
The environment is called **the lab**.
Three things follow directly and are decided here:
### The lab is driven by `incus`
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
bridges between them, through one interface. It also manages system containers, so if a run
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
does.
### The simulated public segment uses TEST-NET-3
`203.0.113.0/24`, reserved by RFC 5737, never routable.
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
peer's underlay address against an RFC1918 regex
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
rather than an addressing mistake.
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
### The lab issues its own certificates
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
single-authority lab would hide any fault living in that split.
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure rather than a mystery.
## Consequences
**The fidelity question disappears, and with it a class of argument.** There is no "how real
is this node" to litigate per capability, because the node is real. What remains not-real is a
short, enumerable list: the model provider, the public internet, and the public certificate
authority.
**The install becomes the thing under test.** A container-shaped lab would have had to skip
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
remains possible at lower fidelity if one is ever genuinely needed.
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
time is noise.
**One change is required before the lab can issue certificates.** The reverse proxy sets no
`caServer`, so it defaults to the public authority's *production* endpoint
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
than staging means every certificate experiment on a real node consumes issuance quota.
## References
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`02-DESIGN/01-end-to-end-testing.md`](../02-DESIGN/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
- RFC 5737 — reserved documentation address blocks
+30
View File
@@ -0,0 +1,30 @@
# Architecture Decision Records
One file per decision, numbered, never deleted. A superseded ADR gets its status changed
and a pointer to what replaced it — the reasoning that was rejected is the expensive half
to rediscover.
## Format
```
# N. Title in plain language
- **Status:** Proposed | Accepted | Superseded by ADR-XXXX
- **Date:** YYYY-MM-DD
- **Deciders:**
## Context what is true today, with evidence
## Considered Options numbered, each with why it was rejected
## Decision what we are doing
## Consequences what follows, including what gets harder
## References code, data, prior art
```
State evidence, not assertion. "Zero of 124 modules declare `brain` as a dependency"
outranks "the dependency rule is not followed".
## Index
| ADR | Title | Status |
|-----|-------|--------|
| [0001](0001-mesh-brokers-nodes-host-agents-think.md) | The mesh brokers capabilities; nodes host; agents think | Accepted |