The numbering is the flow: decisions are 02, design is 03

papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
This commit is contained in:
2026-08-23 18:05:11 +02:00
parent f05e4a0dce
commit c0b35652d0
72 changed files with 217 additions and 204 deletions
@@ -0,0 +1,68 @@
---
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 1. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
> deliberation that happened.
## Context
The mesh is a set of machines that must call each other's capabilities. On the day the
repository was founded there was no inter-node transport at all — each node was configured
independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
## Considered options
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
participant. It also makes node liveness a caller's problem: a request to a sleeping node
is an error rather than a wait.
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
independent of demand, and request/reply has to be built on top of it by hand.
3. **A central message broker with per-node exchanges.** Chosen.
## Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange
named for itself and a request queue; a shared mesh exchange carries commands and events that
are not addressed to one node.
Three message shapes, and only three:
- **RPC** — request/reply, for calling a capability that lives on another node.
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
- **Events** — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
## Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker
connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually
right and occasionally the wrong thing entirely — a queued command for a node that never
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
where the capability is.
## References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.
@@ -0,0 +1,73 @@
---
status: accepted
date: 2026-03-14
deciders: jochen
reconstructed: true
---
# 2. Everything is a module, and one manifest describes all of them
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh carries several kinds of thing: containerised services with data and ports, pure
capability providers with no service at all, and bare markers whose only content is that a
node has them. Before this decision these were separate concepts with separate handling —
the earlier vocabulary was *capabilities*, and services were installed by a different path
than tools.
Every distinct kind of thing needs its own install path, its own change detection, its own
place in the delivery pipeline, and its own documentation. Three kinds means three of each,
and every new feature has to be built three times or, more commonly, once — leaving two kinds
quietly unsupported.
## Considered options
1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag
list. Rejected: it is what existed, and the cost was paid in every cross-cutting change.
2. **One manifest, kind inferred from directory contents.** Chosen.
3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service
still declares itself — but the general rule became inference, because a declared list and
the directory it describes drift, and the directory is the one that is true.
## Decision
Everything the mesh installs is a **module**: a directory with a manifest. The manifest
declares identity, environment variables, what the module provides, what it requires, and how
it is exposed. What kind of module it is follows from what the directory contains:
| Contains | Is |
|---|---|
| a compose definition | a service |
| a tools directory | a capability provider |
| a daemon or unit directory | a long-running process |
| a configs directory | a source of managed files |
| nothing but a manifest | a flag — presence is the whole content |
A module may be several of these at once. Each is a **feature**, and the delivery pipeline
addresses features, not modules.
The mesh's own components are modules on exactly these terms. They get no privileged install
path, no separate registry, and no exemption from the pipeline.
## Consequences
- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches
everything the mesh carries.
- Dogfooding stops being a discipline and becomes structural: if the mesh's own components
need an exception, the machinery is unfinished, and that is visible immediately.
- Feature detection from directory contents means a directory rename silently changes what a
module *is*. This has bitten repeatedly — a hook named for a feature the module does not
have is skipped without complaint.
- The manifest becomes load-bearing and grows. It is now the largest single point of
coupling in the mesh.
## References
- `Rename capabilities → modules across the entire codebase`, 2026-03-14.
- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive
by conversion rather than by creation.
- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`.
- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`,
`troubleshooting/health-check-tools-index-false-positive`.
@@ -0,0 +1,65 @@
---
status: accepted
date: 2026-04-02
deciders: jochen
reconstructed: true
---
# 3. The mesh database is the source of truth; the repository is node-agnostic
> Reconstructed after the fact from the evidence cited below.
## Context
Two things must be known to run the mesh: **what exists** — which modules there are, what each
declares, how each is built — and **what runs where** — which node hosts which module, with
which settings, at which version.
The repository is the natural home of the first. It was initially also the home of the second:
per-node directories held that node's configuration, and adopting a machine meant committing
its files. That has three costs. A node cannot be changed without a commit, so runtime state
and source share a review cadence they do not share a rhythm with. Two nodes cannot be
reconciled, because nothing holds both. And the repository becomes an inventory of the
installation, which is exactly the content that cannot be made public.
## Considered options
1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding
change is a commit and a deploy, and the repository accumulates an inventory of one
particular mesh.
2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no
authority: two nodes disagreeing have no arbiter, and drift is invisible until something
breaks.
3. **A mesh database as the single authority, cached locally for resilience.** Chosen.
## Decision
A single database holds every binding: which node hosts which module, at which selection, with
which environment overrides, plus mesh-level settings that all nodes read. The runtime loads
its configuration from that database at startup and falls back to a local cache when the
database is unreachable.
**The repository defines what exists. The database defines what runs where.** No node-to-module
mapping is ever committed.
A node is therefore not described anywhere in source. Bringing one into the mesh is a database
operation.
## Consequences
- The repository becomes node-agnostic, and can be published without disclosing an
installation. This repository's public stance rests on that property.
- A binding changes without a commit, a build, or a deploy.
- The local cache means a node survives losing the database, but a node running from cache is
running from a snapshot with no indication of its age. Divergence is silent by construction.
- The database is the hardest dependency in the mesh. It is also a module, provisioned like
any other, which makes its bootstrap circular — resolved by the first-node initialisation
script, and the reason such a script exists.
- Nothing on a node is authoritative. That is what makes the next decision necessary.
## References
- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration
tables that landed with it.
- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments."
- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`.
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-04-03
deciders: jochen
reconstructed: true
---
# 4. Managed files are generated onto nodes and never edited there
> Reconstructed after the fact from the evidence cited below.
## Context
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh
database. But the things that consume those bindings — environment files, service
definitions, daemon configuration, firewall rules — are files on a node's disk, because that
is what the software reading them requires.
So the same value exists twice: authoritatively in the database, and materialised in a file.
Any edit to the file is a change to a copy. Before this decision, environment values could be
pushed from a node back into the database, which made the direction ambiguous in both
directions at once.
## Considered options
1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed.
Two writers and no arbiter: whichever synced last wins, and neither is authority.
2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts
ADR 0003 and returns to state that cannot be reconciled across nodes.
3. **Strictly one-directional: the database is written, files are generated.** Chosen.
## Decision
Every managed file is **derived**. A synchroniser regenerates it from the mesh database
whenever the underlying values change. The write path is the mesh tool that owns the value;
the file is an output.
This applies to generated environment files, service definitions, managed configuration, and
anything else a synchroniser lists as its own.
An edit to a managed file survives until the next synchronisation and is then overwritten,
without a warning, taking whatever it was fixing with it.
## Consequences
- **A file edited on a node is a bug with a delay on it.** This is now one of the mesh's core
values, and it is a consequence of this decision rather than a stance taken independently.
- To change a value you must know which tool owns it. That is a real cost, paid every time,
and the reason the mesh provides a way to ask whether a given file is managed.
- Debugging by editing a file no longer works, and fails in the most confusing way available:
it works, and then stops working later for no locally visible reason.
- Recovery is cheap. A node's entire managed surface can be regenerated from the database.
- Values resolve by precedence — database override, then existing file value, then generated,
then manifest default — which means an unset override does not clobber a generated
password. The subtlety is real and has caused its own confusion.
## References
- `Extract hal/env-sync module, remove syncEnvToDb`, 2026-04-03 — the commit that removed the
node-to-database direction.
- Knowledge base: `conventions/no-direct-file-mutation`, `provisioning` (resolution priority).
- Regeneration gaps: `troubleshooting/changed-manifest-default-not-rerendered`,
`troubleshooting/config-removed-from-manifest-not-pruned`.
@@ -0,0 +1,67 @@
---
status: accepted
date: 2026-04-06
deciders: jochen
reconstructed: true
---
# 5. Capabilities are provisioned on declaration, not configured by hand
> Reconstructed after the fact from the evidence cited below.
## Context
Most modules need something another module holds — a database, a cache, a bucket, a message
vhost, an identity client. Wiring that by hand means creating the resource, creating a user,
generating a credential, putting it in the consumer's configuration, and repeating all of it
on every node the consumer runs on.
Every step is a place to make a mistake that surfaces much later, and the credential ends up
written somewhere it can be read.
## Considered options
1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a
description of the mistakes people will make.
2. **A shared credential per resource type**, distributed to all consumers. Rejected: no
isolation, and rotation becomes a mesh-wide outage.
3. **Declared requirements, satisfied by the provider module.** Chosen.
## Decision
A module declares what it **provides** and what it **requires**. A requirement names the
provider, the resource type, optionally a name and a target node, and a mapping from the
resource's connection fields to the consumer's environment variables.
The mesh satisfies it: a provisioner belonging to the provider creates the resource and its
credential, records the grant, and writes the mapped values as database overrides. The
synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then
materialises them. Neither the credential nor the topology is ever written by hand.
A requirement may name a provider on another node. The grant records consumer and provider
nodes separately, so cross-node wiring is the same declaration.
## Consequences
- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a
capability; where it lives is the mesh's problem. This is the property
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model
around.
- Credentials are never authored, so they are never authored badly, and they are never in the
repository.
- Each consumer gets its own credential, so revocation is per-consumer.
- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the
peers holding the old one, and this has taken the mesh down. The declaration model makes
granting easy and says nothing about fan-out.
- A module with no requirements skips the stage entirely, which is correct and also means the
absence of provisioning is indistinguishable from provisioning that did not run.
## References
- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the
provisioner daemon becomes its own component.
- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then-
environment-then-start sequence becomes the coordinator's.
- Knowledge base: `provisioning`, `provisioning/requires`.
- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`,
`troubleshooting/provision-adoption-rotates-live-credential`.
@@ -0,0 +1,68 @@
---
status: accepted
date: 2026-05-14
deciders: jochen
reconstructed: true
---
# 6. Schema and state changes are numbered migrations, in the same language as the code
> Reconstructed after the fact from the evidence cited below.
## Context
Modules own persistent state. That state has to change as they change, across nodes that are
at different versions, some of which have data that predates the change.
Two things were being done that do not survive contact with a second node. Schema was created
at startup, so what a table looked like depended on which version last started. And migrations
were shell scripts, so they could not use the types, connection handling or helpers the module
already had, and were not compiled or checked with it.
## Considered options
1. **Startup SQL / create-if-missing.** Rejected. It converges only for a node that started
with the newest version. A node that never restarts never migrates; a node that restarts on
an old version can undo a change.
2. **Shell migrations.** Rejected. Unchecked, untyped, and a separate dialect from the module
they belong to. Also, as later discovered, packaged differently — and therefore
occasionally not packaged at all.
3. **Numbered migrations in the module's own language, compiled with it.** Chosen.
## Decision
Every schema or state change is a numbered migration file, written in the same language as the
module and compiled with it. There is no startup schema creation and no ad-hoc statement.
Rules that come with it:
- The initial migration is **frozen** once it has run anywhere. It is never modified; a change
is a new number.
- Every statement is **idempotent** — guarded so that re-running is safe.
- A change needs **both** a baseline for a fresh installation and an incremental migration for
installations that already exist. Code referencing a column requires that the migration
creating it exists.
- Migration numbers are unique. A duplicate prefix is a defect, not a style issue.
## Consequences
- A node at any version converges to the current schema by running the migrations it has not
run.
- Migrations are checked by the same compiler as the code, and a migration that does not
compile fails the build rather than the deployment.
- The rules are enforced unevenly. Duplicate prefixes have shipped repeatedly and been fixed by
renumbering afterwards; the mesh now validates for them, which is the check this rule needed
in order to be real.
- A migration directory is a feature like any other, which means it is packaged like any other
— and when packaging is wrong, migrations silently do not ship. This has happened.
- Freezing the initial migration means a fresh installation replays the entire history. That
cost grows and nothing currently bounds it.
## References
- `fix(agents): convert workflow migrations to TypeScript` (#50), 2026-05-14.
- `feat(dev_validate): guard against duplicate migration numeric prefixes` (#379), 2026-06-26 —
the rule acquiring a check.
- Knowledge base: `migrations/schema-drift`, `noxflow/troubleshooting/migration-number-collision`.
- Packaging failures: `troubleshooting/shell-migrations-never-packaged`,
`troubleshooting/provision-migration-never-applied`.
+65
View File
@@ -0,0 +1,65 @@
---
status: accepted
date: 2026-06-04
deciders: jochen
reconstructed: true
---
# 7. No workspace — each module is a standalone package consuming published dependencies
> Reconstructed after the fact from the evidence cited below.
## Context
Modules depend on each other, above all on the shared library every module builds against.
A workspace was the obvious way to express that: sibling packages, resolved locally, one
install at the root.
It produced a divergence that is worth stating precisely, because it is not obvious. In
development, a workspace member importing a sibling resolves to that sibling's **local source**.
In the pipeline, each module is built alone, from a clone, without its siblings present — so
the same import resolves to the **published version**. The two environments were therefore
building different code from identical source, and the failure appeared only in the pipeline,
in a module that had not been touched.
## Considered options
1. **Keep the workspace and make the pipeline replicate it** — clone every module, build the
graph. Rejected: it makes every build a whole-repository build, which is the cost the
per-module pipeline exists to avoid, and it does not extend to modules in their own
repositories.
2. **Keep the workspace and pin siblings to published versions.** Rejected as the worst of
both: the workspace's local resolution silently overrides the pin, so the divergence
remains while looking solved.
3. **No workspace. Every module is standalone and consumes published dependencies.** Chosen.
## Decision
There is no workspace. Each module is an independent package that declares its dependencies
and consumes them from the private registry, including the mesh's own shared library.
A cross-package change is therefore two steps: publish the producer, then consume it. The
pipeline does the first on push and resolves the levels so that a module always builds against
its dependencies' freshly published versions.
## Consequences
- Development and the pipeline resolve imports identically. The divergence is gone by
construction rather than by discipline.
- A module in its own repository is not a special case. It builds exactly as a module in the
monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md)
cheap.
- A cross-package change costs a publish-and-consume round trip. This is the real price, paid
on every shared-library change.
- There is no repository-wide install and no repository-wide build. Anything that assumed one
broke, and one thing that assumed one has stayed broken: the end-to-end pipeline harness has
not built since this decision landed. See
[`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).
## References
- `fix(noxflow): kill npm workspace, restore encryption inside PgAdminRepo` (#240),
2026-06-04. The reason is recorded in the root package manifest, which still carries the
note explaining why no workspace exists.
- The divergence it fixed is named there: workspace members importing each other resolved to
local unbuilt source in the pipeline.
@@ -0,0 +1,71 @@
---
status: accepted
date: 2026-06-05
deciders: jochen
reconstructed: true
---
# 8. A step that fails must fail the job
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh's expensive faults are not crashes. They are the operations that reported success and
did nothing: an artifact that partially downloaded and was extracted anyway, a package that
404ed from every mirror while the job went green, a hook that never ran because it was named
for a feature the module does not declare, a deploy that reported the transport succeeded
rather than that the effect happened.
Each of these was found long after it happened, by someone investigating an unrelated symptom.
The cost is not the failure; it is the interval between the failure and anyone learning of it,
during which decisions are made on the assumption that the thing worked.
## Considered options
1. **Continue on error and report at the end.** Rejected — it is largely what existed. A
summary nobody reads is not a report, and later steps run against the state the failed step
should have produced.
2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a
precise, located failure into a vague one discovered elsewhere, and requires a health check
for every possible partial state.
3. **Fail the step, fail the job, say which step.** Chosen.
## Decision
A step that fails stops the sequence it is part of, and the failure is surfaced where the work
was requested — not only in a log.
Concretely, and these are the forms it takes:
- A scripted sequence gates each step on the previous one. A directory change that fails must
stop the commands that assumed it.
- An artifact that does not fully download is not extracted.
- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must
mean the thing is running, not that a command returned.
- A template that cannot resolve a variable is not written half-rendered.
**Prefer failing to lying.** A green result that is not true costs more than a red one.
## Consequences
- Failures are noisier and land earlier, on the person who caused them.
- Some jobs that used to complete now stop. In every case examined so far, that job was
producing a partial result that something downstream trusted.
- This is a rule the mesh has adopted repeatedly rather than once, because each instance is
written in a different place — a shell hook, a download path, a deploy stage. It is not
enforced by a mechanism, and cannot currently be checked in general. New instances are still
being found; the package-install case remains open as
[`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md).
## References
- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05.
- `A flavor template with an unresolved variable is written to disk instead of failing`
(#710), 2026-08-08.
- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`,
`troubleshooting/service-started-is-not-ready`,
`troubleshooting/green-pipeline-means-transport-not-effect`,
`troubleshooting/silent-failures-and-stale-state`.
- The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must
be loud."
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 9. The mesh is governed by a constitution, injected where work is decided
> Reconstructed after the fact from the evidence cited below.
## Context
By mid-2026 the mesh was doing a large share of its own design and implementation work through
agents. The rules those agents were expected to follow existed — in operating instructions, in
convention documents, in the knowledge base — but they were **retrieved**: an agent had to know
a rule existed in order to look it up.
Rules that must be looked up are followed by whoever already knows them, which is precisely the
population that does not need them. The rules being violated were the ones nobody thought to
search for.
## Considered options
1. **Documentation plus review.** Rejected — it is what existed. Review catches a violation
after the work is done, and only if the reviewer knows the rule.
2. **Lint and automated checks only.** Rejected as insufficient, not wrong. A check catches
what can be expressed mechanically; most of these rules are about judgement — what belongs
in a repository, when a criterion counts as verified.
3. **A canonical rule set, injected into context wherever work is decided, with a check phase
before output is accepted.** Chosen.
## Decision
A single canonical document states the mesh's non-negotiable rules. It is **injected
proactively** into every eligible design and analysis session — agents do not fetch it, it
arrives — and a check phase verifies the session's output against it before the work proceeds.
It is a governed document, not a page. Changing it requires a proposal, sign-off by reviewers
who are not the proposer, and a recorded decision. Drive-by edits are reverted.
Scoped override pages may **tighten** it for a team or product. They may never relax it.
## Consequences
- A rule reaches the work whether or not anyone remembered it existed.
- The check phase makes a violation a blocking outcome rather than a review comment.
- Two copies of the same rules now exist: this document, and the reasoning in HQ that earned
them. The enforced copy wins by default, so the reasoned copy quietly stops being true —
which is why [`00-META/how-we-build.md`](../00-META/how-we-build.md) is now the source
and the governed page is derived from it, via playbook
[`05-constitution-sync.md`](../00-META/process/05-constitution-sync.md).
- Injection costs context on every eligible turn, and grows with the document. Nothing
currently bounds that.
- The amendment process requires two reviewers, which a mesh with one human operator satisfies
only by counting agents. That tension is real and unresolved.
## References
- The governed page was authored 2026-07-10 and carries its own amendment process.
- `feat(noxflow): HAL architectural conformance gate for reviewer + architect` (#296),
2026-06-10 — the check phase, predating the document it checks against.
- Knowledge base: `platform/constitution`.
@@ -0,0 +1,64 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 10. Applications live in their own repository; the monorepo is for the mesh
> Reconstructed after the fact from the evidence cited below.
## Context
The module system makes adding anything to the monorepo trivial — a directory and a manifest.
That ease is the problem. Standalone applications, sites and side-projects accumulated beside
the mesh's own components, and once there they inherited the monorepo's review cadence, its
pipeline detection, and its history.
The mesh's own code and an application that merely runs on the mesh have nothing in common
except the manifest format. They change for different reasons, are reviewed by different
criteria, and have no reason to share a branch.
## Considered options
1. **Everything in the monorepo.** Rejected — it is what existed. The monorepo becomes an
inventory of one installation's applications, and every application change queues behind
mesh review.
2. **A second monorepo for applications.** Rejected: the same coupling with an extra name.
Applications have no more in common with each other than with the mesh.
3. **One repository per application, registered with the mesh as a build source.** Chosen.
## Decision
Every standalone application, site or side-project lives in its own repository, with a manifest
at the root. It registers with the mesh as a build source and is then built, provisioned,
deployed and verified by exactly the same pipeline as anything in the monorepo.
The monorepo holds the mesh: the runtime, the core modules, the delivery machinery, and the
shared infrastructure the mesh itself provisions against.
Creating an application directory in the monorepo is a convention violation, and reviewers
reject it.
## Consequences
- An application's cadence is its own. It is not reviewed as mesh code and does not queue
behind mesh work.
- The separation is safe **only because** the pipeline and provisioning are identical either
side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without
standalone packages this decision would fork the build.
- The monorepo stops being an inventory of the installation, which is a precondition for
publishing anything about it.
- Discovery gets harder: there is no single listing of everything the mesh runs, and the
registry of build sources becomes the closest thing to one.
- A module source that is not registered is silently skipped by the pipeline. The cost of
being outside the monorepo is that being forgotten is possible.
## References
- The rule is stated in the governed constitution page authored 2026-07-10, §3, as a
convention violation reviewers must reject.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
applications to the modules already in the monorepo.
- Knowledge base: `troubleshooting/unregistered-module-source`.
@@ -0,0 +1,59 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 11. The installer owns linking; nothing else creates a symlink
> Reconstructed after the fact from the evidence cited below. The incident that earned the rule
> predates the record, and its date is not established here.
## Context
A service's definition lives in the module catalogue; its runtime directory and persistent data
live outside it. The mesh connects the two by linking the definition into the runtime location
— deliberately, so that runtime state and source stay separate while the running service reads
a current definition.
A link is also the easiest thing in the world to create by hand while fixing something, and a
container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere
it should not have, and **production data was lost**.
## Considered options
1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss
for a service running a definition nobody can find.
2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of
the mistake — the link looks right and the resolution happens inside the container engine.
3. **One component owns linking; everyone else is forbidden.** Chosen.
## Decision
The installer creates and repairs every link the mesh needs. It reconciles them: a missing
link is created, a stale one is repointed, and a real file found where a link belongs is
adopted into the node's override location and replaced.
**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person
debugging. The prohibition is absolute because the judgement required to make a safe exception
is exactly the judgement that was not available at the moment it mattered.
## Consequences
- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who
has not seen the incident. That is why it is recorded here rather than only asserted.
- Links become reconcilable state rather than incidental filesystem facts.
- The rule is stated for humans and agents and is enforced by convention, not mechanism. A
check does not exist.
- The rule as written governs the mechanism rather than removing it. A link made by the
installer resolves the same way as one made by hand, so the hazard is narrowed and not
closed. [ADR 0018](0018-the-mesh-creates-no-symlinks.md) proposes widening this to "nothing
links, the installer included"; until that is accepted, this record governs.
## References
- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or
service directories have caused production data loss via Docker volume path resolution. The
installer handles all linking. Never create symlinks manually."*
- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files.
@@ -0,0 +1,74 @@
---
status: accepted
date: 2026-07-12
deciders: jochen
reconstructed: true
---
# 12. An agent is a persistent employee, not an instance of a pool
> Reconstructed after the fact from the evidence cited below.
## Context
Agents were originally a **pool**: a named kind of worker, scaled to some number of
interchangeable instances. Work went to whichever instance was free.
That model has no place to put the things that turn out to matter. An agent that accumulates
knowledge of a domain cannot keep it, because the next task lands on a different instance. An
agent cannot own a workspace, because there are several of it. It cannot be held to a policy —
warned for a violation, then dismissed — because there is no continuing subject to warn.
Scaling was also solving a problem the mesh does not have. Instances were being multiplied to
get concurrency, when concurrency is a property of how much work one agent may hold at once.
## Considered options
1. **Keep the pool, attach memory to the pool.** Rejected: shared memory across
interchangeable workers is a knowledge base, not an agent's experience, and the mesh
already has one.
2. **Keep the pool, make instances sticky.** Rejected as a pool pretending to be identities —
identity by scheduling accident, lost on any restart.
3. **One agent is one persistent identity, with concurrency as a property of it.** Chosen.
## Decision
An agent is a **singular, named, persistent identity**: a home node, a workspace on that node,
accumulating memory, and a lifecycle — hired, active, draining, retired. Not a pool member.
Concurrency is a property of the agent, not a count of copies: an agent has a cap on how many
sessions it may hold at once.
Lifecycle is explicit and has verbs. An agent is hired onto a node; it may be reassigned while
idle; it is retired by draining first, and forced only deliberately. Retired agents are not
deleted.
Surge capacity is expressed within the model rather than against it: a template agent is a
blueprint, cloned into a real agent with a lifetime when a queue grows, drained and retired
when it expires. A temporary employee is still an employee.
Some agents are **human**. What differs is modality — how the agent acts — not category. A node
itself is an agent of a kind exempt from the hiring lifecycle.
## Consequences
- Memory, workspace and reputation have a subject to belong to. Policy becomes possible: an
agent that violates a rule can be warned, and warned agents can be dismissed.
- The mesh gained a hiring model, and with it the question of who may hire.
- Scaling by adding instances is gone. If one agent is saturated, either its session cap rises
or another agent is hired — both deliberate acts.
- The transition was not free. Lifecycle columns had to reach every query that selects an
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises:
one kind of participant, differing only in modality.
## References
- `docs(adr): agents as persistent employees + MINERVA librarian` (#495), 2026-07-12 — the
original record, in the code repository.
- `feat(B4): one persistent employee, N sessions — rename max_instances → max_sessions` (#547)
and `feat(noxflow): B3 — workspace provisioner for agent employee model` (#549), 2026-07-20.
- `feat(noxflow): warn-then-fire agents who merge to main without review` (#209), 2026-06-01 —
policy that presumes a continuing subject, predating the model that provides one.
- Knowledge base: `agents/employee-lifecycle`, `agents/temp-surge`, `agents/workspace-layout`.
- The migration cost: `troubleshooting`/`noxflow-agent-enriched-select-missing-lifecycle-columns` (#546).
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 13. An artifact is build output, never a source tree
> Reconstructed after the fact from the evidence cited below.
## Context
A module is built once and deployed to every node assigned to it. What travels between those
two events is the artifact.
For a long time the artifact was a filtered copy of the module's source directory. Deploying it
therefore meant resolving and installing its dependencies **on the target node** — which
requires the target to reach a package registry, at deploy time, for every node, every deploy.
A node with no route to the registry could not deploy code that had already been built
successfully.
## Considered options
1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy
becomes a network operation with a failure mode per node, and the code that runs is
assembled independently on each one.
2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the
dependency resolution's platform assumptions along with it.
3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen.
## Decision
The artifact is the module's **build output directory** — compiled and bundled, with its
dependency graph inlined. Deploy is extract-and-run and touches no network.
A build that cannot produce a self-contained output **fails**. It does not fall back to
shipping a dependency tree, because a fallback that works is a fallback that is never fixed —
an application of [ADR 0008](0008-a-failed-step-fails-the-job.md).
## Consequences
- A node can deploy without reaching a registry. What was built is what runs, identically, on
every node.
- Deploys are faster and their failure modes are local.
- **Everything not in the build output does not ship.** This is the decision's whole cost, and
it was paid several times before it was understood: migrations that read the source layout,
provisioning scripts that read the source layout, selection files never packaged at all. Each
worked in development, where the source is present, and silently did nothing after deploy.
- Any file a module needs at runtime must be deliberately placed into the build output. The
rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code,
and that generalisation was the expensive part.
- Bundling has its own failure modes that a compiler will not catch — a bundler can exit
successfully and produce output that cannot load.
## References
- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04.
- The consequences, in order: `Provision migrations and seeds read the source layout, not the
artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07.
- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`,
`troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`,
`troubleshooting/esbuild-silent-tla-breakage`.
@@ -0,0 +1,78 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 14. Build, publish and deploy are three silos with different cardinality
> Reconstructed after the fact from the evidence cited below.
## Context
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
run a different number of times.
- Compiling happens **once per module feature**, on the build node.
- Packaging and uploading happens **once per module feature**, on the build node.
- Installing, configuring, starting and verifying happens **once per module feature per node**.
Conflating them is what made earlier versions slow and hard to reason about. Work that should
happen once was being repeated per node, and the fan-out point was implicit rather than a
boundary anything could observe.
The split had been declared before it was real. Packaging still happened inside the build,
which meant the boundary existed in the documentation and not in the code.
## Considered options
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
Cardinality is then a property of each stage's implementation, and nothing can reason about
the pipeline as a whole.
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
so build must know every module, every feature, and how each composes its artifact —
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
a stale package instead of re-packaging.
3. **Three silos, with an explicit handover between each.** Chosen.
## Decision
Delivery is three silos, and the boundaries are real:
| Silo | Runs | Where |
|---|---|---|
| **build** | once per module feature | the build node |
| **publish** | once per module feature | the build node |
| **deploy** | once per module feature **per node** | every assigned node |
Commands and events are addressed **per feature**, not per module.
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
publishing, so a module whose artifact is a package publishes in the publish silo, not the
build one.
Modules are resolved into dependency **levels**, and a level completes before the next begins,
so a module always builds against its dependencies' freshly published versions.
## Consequences
- Work that should happen once happens once. The fan-out point is explicit and observable.
- A failed upload retries by re-packaging, because packaging belongs to the stage that
uploads.
- The handover is a staged tree in a known location rather than the build's working directory,
which is reference-counted and cannot be assumed to still exist when a later stage runs.
- The build node is now the only node that has already passed through two silos when the
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
stages; code that knew only about the first parked the build node forever while every other
node deployed cleanly.
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
reports success over a stall it cannot see.
## References
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
documents claimed otherwise and were stale until 2026-08-06.
- The build-node stage-tracking failure was observed on pipeline #5557.
@@ -0,0 +1,188 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 15. The mesh brokers capabilities; nodes host; agents think
## Context
HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split
because splitting is the only granularity the platform offers, and domains are merged
because a shared database is the only integration it offers. Both pressures push in the
same direction: boundaries end up drawn by deployment accident rather than by domain.
Three observations establish the state.
**The word "agent" means two different things.** The original design treated a node as an
agent with thinking abilities. Later, noxflow implemented agents as employees with
skills, workspaces and tasks. Both survive. The collision is visible in the data — there
are **two agent rows per node**:
| agent | skills |
|---|---|
| one named after the node | `{deploy,verify,operate}` |
| one named `hal-<node>` | `{}` |
One carries the work; the other carries only identity, existing to hold a licence for
HAL's own sessions. The same split appears in the schema: `nodes.hal_claude_account` and
`agents.claude_account` are one fact in two tables in two databases, and what was
documented as "three licence touchpoints" is one concept modelled three times.
**Domains integrate by sharing a schema.** `noxflow` is 45 tables spanning five domains —
tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose
is 7 of 45. Because agent identity lives in noxflow's database, work that belongs
elsewhere must be implemented there: per-agent Claude credentials had to be written by
the noxflow runtime, even though node identity — the same kind of fact — lives in the
mesh registry.
**The core domain has no context to live in.** The mesh's distinguishing feature is that
a module declares `requires: postgres/database` and never learns where the database
lives, who owns the credential, or how it rotates. That is capability brokering, and it
is what makes this a mesh rather than four machines with a configuration manager. Yet the
logic implementing it sits in `modules/postgres/tools/index.ts` — a provider module,
where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its
invariants were found violated simultaneously (see Consequences).
## Considered Options
1. **Keep the current layout; fix bugs as they surface.** Rejected. The faults are not
independent. Every incident on 2026-08-22 — credentials written by the wrong module, a
dependency question answered wrongly twice, an invariant enforced nowhere — traced to
a boundary that was never stated. Fixing them individually leaves the generator intact.
2. **Merge aggressively into few large modules.** Fewer names, same problem: a shared
schema across domains is what produced noxflow, and doing it deliberately would
produce it again at larger scale.
3. **Decompose by bounded context, with the mesh as a broker.** Name contexts after their
aggregates, integrate through a published record rather than a shared schema, and let
deployment granularity be a feature-level concern rather than a reason to create a
module. **Adopted.**
## Decision
### The domain, in one sentence
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
resident agent, no node-owned identity, no ownership in either direction. Agents named
after a node remain, as **ordinary agents** that happen to hold infra skills.
Consequently `nodes.node_license` and `nodes.hal_claude_account` cease to exist: a node
does not authenticate to a model provider, agents do. The two agent rows per node merge.
### There is one kind of participant, and some are human
Human and non-human participants are both **agents**. Both hold identity and credentials;
both act, remember and coordinate. What differs is **modality** — how an agent acts:
| modality | credential is delivered to |
|---|---|
| spawned session | that agent's own config directory |
| shell or desktop | that agent's home on the node it acts from |
A node holds no licence. **An agent holds credentials, and delivery follows that agent's
node bindings and modality.** A human agent's grant arrives in the home directory of the
user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's
grant in its config directory, with a different target.
This requires one fact the mesh does not record today: which user, on which node, a given
human agent acts as. Adding it is what removes the node licence — the node is currently
standing in for an identity the mesh cannot name. It is also what makes a second human
agent require no new mechanism.
### Contexts
| context | aggregate | subdomain |
|---|---|---|
| `hal/mesh` | **Provision**, Node, Module — brokering and its bookkeeping | core |
| `hal/agents` | **Agent** — identity, licence, runs, memory, thoughts | core |
| `hal/work` | **Task** — workflows, bindings | core |
| `hal/stream` | **Thread** — mentions, messages, meetings, notifications | core |
| `hal/delivery` | **Pipeline** — jobs, artifacts, features | supporting |
| `hal/knowledge` | **Document** — spaces, revisions, review | supporting |
| `hal/ai` | **Licence** — provider grants and rotation | supporting |
| `hal/observability` | **Check** | supporting |
| `hal/config` | **Setting** — env, secrets, PKI | generic |
A *brain* — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is
not a module. Anatomy makes attractive names and poor boundaries; today `hal/brain` names
infrastructure and `hal/cortex` describes itself as messaging while running nowhere.
### Provisioning is the core domain, not plumbing
`hal/mesh` is a broker; the registry is its bookkeeping. Its invariants are explicit and
owned:
- one rotation source per resource
- a credential change fans out to every consumer
- a consumer never holds a credential the provider does not know about
### Contexts integrate through the record, never a shared schema
`hal/stream` is the published language. A context publishes; it does not join across a
boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a
notification is a mention not yet read.
### Third-party software leaves the repository
`plex`, `sonarr`, `postgres`, `verdaccio`, `docker-registry` and 88 others run **on** the
mesh; they are not **of** it. The pipeline and provisioning are deliberately
module-agnostic, so HAL's own modules dogfood exactly what external modules use — which
is what makes the separation safe rather than merely tidy.
## Consequences
**The invariants now have an owner, and were measurably unowned before.** On 2026-08-22,
`provision_ensure` — documented as "NEVER rotates an existing secret" — was found to mint
a new password on every adoption and update only the provider's row. `hal_notifications`
consumers on three nodes held dead credentials for two days; `hal_transcripts` had two
rows written 216 ms apart, so at most one could match the live role.
**noxflow dissolves.** `hal/work` inherits tasks and workflows — the concept it was built
for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes
away.
**`hal/sdk` shrinks.** 155 files, 34,636 lines, containing code from every context —
including `workflow-engine.ts` and `task-commands.ts`, work-domain logic in the kernel
every module imports. Each landed there to avoid a cycle between modules that both needed
it; a domain module can only own its shared code once the domain has a module. Extraction
is therefore downstream of this decision, not independent of it.
**Two mechanisms are prerequisites, not follow-ups.**
- *Named features with per-node opt-in.* Without it, every independently deployable unit
inside a context becomes a module again and the count returns. `hal/claude-licences`
exists solely because one daemon must run on one node.
- *A local mesh in containers.* `dev_up` starts providers "via systemctl (same as
production)" — it borrows the host, and no mesh can be stood up locally. Everything that
manifests **between** nodes is therefore discoverable only in production, which is where
every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable.
**Credential delivery becomes uniform.** Per-agent credential directories, built for
spawned sessions, extend to human agents unchanged. A file previously scoped to a node
becomes scoped to an agent — which would have prevented the class of failure where a
rotation reached one node of four while the mesh reported success.
**Migration is incremental and long.** Contexts can be extracted one at a time behind the
existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
## References
- [`01-RESEARCH/001-module-domain-decomposition`](../01-RESEARCH/001-module-domain-decomposition/analysis.md)
— current-state evidence, table counts, open questions
- [`00-META/how-we-build.md`](../00-META/how-we-build.md) — naming and integration rules
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
array that makes a feature a singleton per module
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the
constraints that make a unit independently shippable, applicable unchanged to features
- impire.io / soulstream — *the record* as integration substrate, personas over services,
and "cheap awareness and expensive thinking"
@@ -0,0 +1,138 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
Two things were measured while establishing what exists
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
- **There is no local mesh.** The dev tooling starts providers through the host's own init
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
It borrows the machine because there is nowhere else to put a mesh.
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
physical machine. Every subsequent question — how faithful is faithful enough, what may be
mocked, which failures remain reachable — follows from that one answer.
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
hardware virtualisation present.
## Considered Options
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
a node as one inverts the thing being modelled: module service stacks then require nested
containers through a privileged daemon, or a shared socket that makes isolation between
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
and the least like a node.
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
is the only option that makes a twenty-node run affordable. It was rejected because **the
scale requirement that justified it was invented rather than required**: the stated goal is
to run the real mesh, which is four nodes, on one computer. And a system container still
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
then has to be answered again for every capability under test.
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
no storage pools, no snapshot management, no network management, no virtual machines.
Snapshots are what make the loop fast, so the saving is not worth what it costs.
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
turns into a node.
## Decision
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
runs the real install, and becomes a node. It is not a model of a node, so no question arises
about how good the model is.
The environment is called **the lab**.
Three things follow directly and are decided here:
### The lab is driven by `incus`
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
bridges between them, through one interface. It also manages system containers, so if a run
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
does.
### The simulated public segment uses TEST-NET-3
`203.0.113.0/24`, reserved by RFC 5737, never routable.
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
peer's underlay address against an RFC1918 regex
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
rather than an addressing mistake.
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
### The lab issues its own certificates
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
single-authority lab would hide any fault living in that split.
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure rather than a mystery.
## Consequences
**The fidelity question disappears, and with it a class of argument.** There is no "how real
is this node" to litigate per capability, because the node is real. What remains not-real is a
short, enumerable list: the model provider, the public internet, and the public certificate
authority.
**The install becomes the thing under test.** A container-shaped lab would have had to skip
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
remains possible at lower fidelity if one is ever genuinely needed.
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
time is noise.
**One change is required before the lab can issue certificates.** The reverse proxy sets no
`caServer`, so it defaults to the public authority's *production* endpoint
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
than staging means every certificate experiment on a real node consumes issuance quota.
## References
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
- RFC 5737 — reserved documentation address blocks
@@ -0,0 +1,93 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0015-mesh-brokers-nodes-host-agents-think.md
---
# 17. Modules outside the platform core are grouped by domain, not by single function
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules
into bounded contexts named after their aggregates, and sends the rest out of the monorepo on
the grounds that they run *on* the mesh rather than being *of* it.
That leaves the larger half unaddressed. Around three quarters of the catalogue are modules
that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an
SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of
media-library services. Today each is its own module, because one module is the unit of *one
piece of software*, and no other grouping exists.
The result is that the catalogue's shape records what was installed, not what anything is for.
Four modules that together constitute "how a node is reachable" have no relationship the mesh
can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a
change to how the mesh handles connectivity has to be made four times.
This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident
rather than by domain* — appearing outside it.
## Considered options
1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else
keeps accident boundaries, so the catalogue becomes harder to read after the refactor than
before it.
2. **One module per piece of software, with a tag or category field.** Rejected. A label is not
a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it
drifts from the thing it labels.
3. **Group them into domain modules, each owning the software that serves one purpose.**
Proposed here.
4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for
the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it
into that model repeats the metaphor-naming mistake ADR 0015 exists to correct.
## Decision
*Proposed — the principle is settled; the domain list is not. See "Open" below.*
Modules that are not part of the platform core are grouped into **domain modules**. A domain
is named for the concern it serves, and owns the software that serves it. The unit stops being
one piece of software and becomes one purpose.
This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own
domain stand unchanged. This decision covers what ADR 0015 leaves outside them.
Naming follows the same rule as the core: **name the domain for what it does, not for what it
is made of**. Connectivity, not a VPN implementation.
## Consequences
- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes
reach each other is a change to one module.
- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from
its module list.
- Swapping an implementation stops being a module replacement, with the data-volume and
provisioning consequences that carries, and becomes a change inside a domain.
- The count drops sharply, which is a symptom of the improvement rather than the point of it.
- **Grouping conceals.** A domain module hides which implementation is in use, and every
operational question — which port, which unit, which credential — gains an indirection.
- The migration is not free and has no obvious increments: a domain is only useful once
everything belonging to it has moved.
- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its
own sake would be the same error in the other direction.
## Open
**The domain list is not settled and this record does not invent one.** What is decided is the
principle; what is not decided is the set. Candidate groupings are visible in the catalogue —
connectivity and reachability, node presentation and desktop, media libraries, observation and
metrics, storage and data services — but naming them here would be reconstructing a decision
that has not been taken.
Settling the list is a research effort, not an act of this record. Until it concludes, this
ADR stays `proposed`.
## References
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this
extends, and its rule about naming a context after its aggregate.
- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are
already out of scope here; they are not domains and do not group.
- [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md)
— the catalogue's current shape, which is the evidence for the problem.
@@ -0,0 +1,101 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0011-the-installer-owns-linking.md
---
# 18. The mesh creates no symlinks — a derived file is a copy
## Context
[ADR 0011](0011-the-installer-owns-linking.md) responded to production data loss — a hand-made
link, resolved through a container engine's volume handling, pointing a mount somewhere it
should not have — by centralising linking in the installer and forbidding it everywhere else.
That narrowed the incident class. It did not close it. The hazard is not *who* made the link;
it is that a path can resolve somewhere other than where it appears to. A link made by the
installer resolves exactly the same way as a link made by hand. The rule made the mechanism
rarer and better-governed while leaving the mechanism in place.
Two things have changed since, and together they remove the argument that kept it.
**The original case for linking was staleness.** A copy of a service definition goes stale
silently while the catalogue moves on, so a link was the cheap way to guarantee the running
node reads a current definition. That argument assumes the node's copy is unmanaged.
**It is not.** [ADR 0004](0004-managed-files-are-generated-never-edited.md) established that
everything on a node's disk is derived from the mesh and regenerated when its inputs change,
and the installer already **reconciles** links rather than assuming them — repointing stale
ones, adopting real files it finds where a link belongs. Reconciling content is the same
operation as reconciling a pointer, plus a comparison.
So the mesh already has the machinery that makes a copy safe, and is using a link to solve a
problem that machinery solves better. Worse, a link is conceptually the wrong shape: it makes
the node's runtime state a *pointer into source*, which is the one thing
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) and ADR 0004 exist to prevent.
State is derived onto nodes; it does not reach back.
## Considered options
1. **Keep ADR 0011 as the final position** — centralised linking, forbidden elsewhere.
Rejected as the status quo. It governs the mechanism rather than removing it, and the
failure it was written for remains reachable by any code path the installer trusts.
2. **Keep links but harden them** — canonicalise before mounting, refuse a link that escapes
an expected root. Rejected: it is a check bolted onto a hazard, and it has to be correct in
every consumer, including container engines the mesh does not control.
3. **Copy, reconciled by the installer, with staleness detected rather than assumed away.**
Proposed here.
## Decision
*Proposed — the position is settled; the migration is not designed. See "Open" below.*
**The mesh creates no symlinks.** A file a node needs is placed on that node as a real file,
derived from the mesh and reconciled by the installer like every other managed file
([ADR 0004](0004-managed-files-are-generated-never-edited.md)).
The prohibition in ADR 0011 stands and widens: it ceases to be "only the installer may link"
and becomes "nothing links, the installer included".
When this is accepted, ADR 0011 becomes superseded rather than edited — its reasoning is why
the rule exists at all, and the incident behind it is the reason anyone believes either record.
## Consequences
- The path-resolution hazard is removed rather than governed. There is no link for a container
engine to resolve, so the class of failure that cost production data is closed by
construction.
- A node's runtime state stops pointing into source. What a node holds is derived output, which
is what the mesh's model already says it is everywhere else.
- **Staleness becomes a real problem that must be answered, not assumed away.** This is the
cost, and it is the whole cost: today a link cannot be stale, and a copy can. The answer has
to be detection — the installer comparing what is on disk against what the mesh says should
be — and it must be loud, because a silently stale definition is exactly the failure shape
this mesh keeps producing ([ADR 0008](0008-a-failed-step-fails-the-job.md)).
- Reconciliation gets more expensive: comparing content rather than checking a pointer's
target, on every module, on every node.
- Disk usage rises, trivially, and is not a consideration.
- Existing links must be converted. A node mid-migration holds both forms, so reconciliation
has to handle finding a link where a file now belongs — the mirror image of the adoption it
already does.
## Open
- **How staleness is detected.** Content hash, version marker, or regeneration on every
reconcile. This is the decision that makes or breaks the change and it is not taken here.
- **Whether anything must keep a link** for reasons outside the mesh's control. If something
does, that is a finding worth recording rather than an exception worth granting quietly.
- **Migration order.** Converting a node's links is a change to how its services resolve their
own definitions, which is not a change to make everywhere at once.
Until those are answered this record stays `proposed`, and ADR 0011 remains the governing rule.
## References
- [ADR 0011](0011-the-installer-owns-linking.md) — the incident, and the rule this widens.
- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — the machinery that makes a
copy safe.
- [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
— what the installer does today, including reconciliation and adoption.
+60
View File
@@ -0,0 +1,60 @@
# 02-DECISIONS
Architecture decision records — the "why" trail behind the rules in
[`00-META`](../00-META/) and the specifications in [`03-DESIGN`](../03-DESIGN/).
**Numbered `02` because a decision precedes the design it authorises.** Research concludes,
the decision is recorded here, and only then is the design written. Following the folder
numbers walks the process in the order it happens.
One file per decision, numbered, never deleted. A superseded record has its `status:` changed
and gains a pointer to what replaced it — **its text is never edited**. The reasoning that was
rejected is the expensive half to rediscover.
The records are a **ledger**: they run in the order the decisions were taken, oldest first.
## Frontmatter
```yaml
---
status: proposed | accepted | superseded
date: YYYY-MM-DD # when the decision was taken, not when it was written down
deciders: name
reconstructed: true|false # true when the record was written after the fact from evidence
superseded-by: # 02-DECISIONS/NNNN-....md, when status is superseded
extends: # 02-DECISIONS/NNNN-....md, when this record widens an earlier one
---
```
## Body
```
# N. Title in plain language
## Context what was true, with evidence
## Considered Options numbered, each with why it was rejected
## Decision what was decided
## Consequences what follows, including what got harder
## References commits, pull requests, knowledge-base entries, prior art
```
State evidence, not assertion. *"Zero of 124 modules declare `brain` as a dependency"*
outranks *"the dependency rule is not followed"*.
## Reconstructed records
Records 0001–0014 were written on 2026-08-23, after the decisions they describe. Records 0015 onward were taken as records. Those
decisions were taken in implementation rather than in a document; the records state what was
decided and the evidence it was decided from, and each carries `reconstructed: true` and says
so in its first lines.
A reconstructed record is not a transcript. Where the deliberation is not recoverable, the
options section states what the alternatives were and why the chosen one won on the evidence
available — not a discussion that did not happen. Where a date is not establishable it says so
rather than guessing.
## Index
The index is **generated, not maintained** — run the `hal-status` skill, which reads the
frontmatter of every record. A hand-written index drifts from the folder it describes, and
this one had already done so after a single addition.