Base layer: the mesh as it is, under the mesh as it should be

HQ held only the to-be. Every reader had to already know the system the
decisions were about, and an as-is claim had nowhere to live except inside
an intention.

Adds 02-DESIGN/00-as-is — eleven documents written from the implementation
and the operational record, not from intent, including the parts nobody
would choose again. The two existing designs move under 01-to-be. Layers
are declared in frontmatter and never mix: a design that ships does not
move, its as-is counterpart is written, and both stand.

Back-fills adr/0001-0014 for decisions taken in implementation and never
recorded — the broker, the module abstraction, the mesh database, managed
files, provisioning, migrations, the workspace removal, failing loudly,
the constitution, application placement, linking, the employee model, the
artifact, the three silos. Each marked reconstructed, dated from the
history, and citing the evidence it was recovered from. The two existing
records renumber to 0015 and 0016 so the ledger runs oldest first;
0017 extends 0015 to modules outside the core, principle only — the
domain list is deliberately not invented here.

how-we-build.md becomes the source of the mesh constitution, with a sync
playbook, so the enforced copy stops being the only one that is true.

Process becomes explicit: five playbooks, eight thin skills that defer to
them, a repository map, and AGENTS.md with CLAUDE.md as its include.

The five Observations become 04-ISSUES 001-005 where they can be owned and
closed. 006 is new and uncomfortable: HQ is not indexed into the knowledge
base. That claim is what decision 27 rests on, it was never checked, and
the README now says so instead of repeating it.

Also corrects the ADR index into something generated, the "02-DESIGN is
empty" claim, the VISION.md pointer that did not survive the repo split,
and a note asserting the symlink rule was contradicted — it was a
misreading; the rule forbids hand-made links, the installer links by design.
This commit is contained in:
2026-08-23 03:08:26 +02:00
parent cf9357e8e9
commit 702efca6bb
74 changed files with 3676 additions and 138 deletions
@@ -0,0 +1,68 @@
---
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 1. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
> deliberation that happened.
## Context
The mesh is a set of machines that must call each other's capabilities. On the day the
repository was founded there was no inter-node transport at all — each node was configured
independently and shared nothing at runtime.
Three properties were required and are visible in everything built since:
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
should wait, not error.
- Adding a node must not require editing anything on the nodes that already exist.
## Considered options
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
participant. It also makes node liveness a caller's problem: a request to a sleeping node
is an error rather than a wait.
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
independent of demand, and request/reply has to be built on top of it by hand.
3. **A central message broker with per-node exchanges.** Chosen.
## Decision
All inter-node communication goes through a message broker. Every node owns a topic exchange
named for itself and a request queue; a shared mesh exchange carries commands and events that
are not addressed to one node.
Three message shapes, and only three:
- **RPC** — request/reply, for calling a capability that lives on another node.
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
- **Events** — statements that something happened, addressed to nobody.
Every node dials the broker outbound. Nothing dials a node.
## Consequences
- NAT stops being an architectural concern. A node's reachability is a property of the broker
connection, not of its network position.
- A call to a node that is down waits in that node's queue instead of failing. This is usually
right and occasionally the wrong thing entirely — a queued command for a node that never
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
- The broker is a single point of failure and a single point of trust. Its credential is
mesh-wide, so rotating it is a mesh-wide operation.
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
where the capability is.
## References
- The broker was stood up on 2026-02-25, the second day of the repository.
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.
+73
View File
@@ -0,0 +1,73 @@
---
status: accepted
date: 2026-03-14
deciders: jochen
reconstructed: true
---
# 2. Everything is a module, and one manifest describes all of them
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh carries several kinds of thing: containerised services with data and ports, pure
capability providers with no service at all, and bare markers whose only content is that a
node has them. Before this decision these were separate concepts with separate handling —
the earlier vocabulary was *capabilities*, and services were installed by a different path
than tools.
Every distinct kind of thing needs its own install path, its own change detection, its own
place in the delivery pipeline, and its own documentation. Three kinds means three of each,
and every new feature has to be built three times or, more commonly, once — leaving two kinds
quietly unsupported.
## Considered options
1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag
list. Rejected: it is what existed, and the cost was paid in every cross-cutting change.
2. **One manifest, kind inferred from directory contents.** Chosen.
3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service
still declares itself — but the general rule became inference, because a declared list and
the directory it describes drift, and the directory is the one that is true.
## Decision
Everything the mesh installs is a **module**: a directory with a manifest. The manifest
declares identity, environment variables, what the module provides, what it requires, and how
it is exposed. What kind of module it is follows from what the directory contains:
| Contains | Is |
|---|---|
| a compose definition | a service |
| a tools directory | a capability provider |
| a daemon or unit directory | a long-running process |
| a configs directory | a source of managed files |
| nothing but a manifest | a flag — presence is the whole content |
A module may be several of these at once. Each is a **feature**, and the delivery pipeline
addresses features, not modules.
The mesh's own components are modules on exactly these terms. They get no privileged install
path, no separate registry, and no exemption from the pipeline.
## Consequences
- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches
everything the mesh carries.
- Dogfooding stops being a discipline and becomes structural: if the mesh's own components
need an exception, the machinery is unfinished, and that is visible immediately.
- Feature detection from directory contents means a directory rename silently changes what a
module *is*. This has bitten repeatedly — a hook named for a feature the module does not
have is skipped without complaint.
- The manifest becomes load-bearing and grows. It is now the largest single point of
coupling in the mesh.
## References
- `Rename capabilities → modules across the entire codebase`, 2026-03-14.
- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive
by conversion rather than by creation.
- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`.
- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`,
`troubleshooting/health-check-tools-index-false-positive`.
@@ -0,0 +1,65 @@
---
status: accepted
date: 2026-04-02
deciders: jochen
reconstructed: true
---
# 3. The mesh database is the source of truth; the repository is node-agnostic
> Reconstructed after the fact from the evidence cited below.
## Context
Two things must be known to run the mesh: **what exists** — which modules there are, what each
declares, how each is built — and **what runs where** — which node hosts which module, with
which settings, at which version.
The repository is the natural home of the first. It was initially also the home of the second:
per-node directories held that node's configuration, and adopting a machine meant committing
its files. That has three costs. A node cannot be changed without a commit, so runtime state
and source share a review cadence they do not share a rhythm with. Two nodes cannot be
reconciled, because nothing holds both. And the repository becomes an inventory of the
installation, which is exactly the content that cannot be made public.
## Considered options
1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding
change is a commit and a deploy, and the repository accumulates an inventory of one
particular mesh.
2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no
authority: two nodes disagreeing have no arbiter, and drift is invisible until something
breaks.
3. **A mesh database as the single authority, cached locally for resilience.** Chosen.
## Decision
A single database holds every binding: which node hosts which module, at which selection, with
which environment overrides, plus mesh-level settings that all nodes read. The runtime loads
its configuration from that database at startup and falls back to a local cache when the
database is unreachable.
**The repository defines what exists. The database defines what runs where.** No node-to-module
mapping is ever committed.
A node is therefore not described anywhere in source. Bringing one into the mesh is a database
operation.
## Consequences
- The repository becomes node-agnostic, and can be published without disclosing an
installation. This repository's public stance rests on that property.
- A binding changes without a commit, a build, or a deploy.
- The local cache means a node survives losing the database, but a node running from cache is
running from a snapshot with no indication of its age. Divergence is silent by construction.
- The database is the hardest dependency in the mesh. It is also a module, provisioned like
any other, which makes its bootstrap circular — resolved by the first-node initialisation
script, and the reason such a script exists.
- Nothing on a node is authoritative. That is what makes the next decision necessary.
## References
- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration
tables that landed with it.
- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments."
- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`.
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-04-03
deciders: jochen
reconstructed: true
---
# 4. Managed files are generated onto nodes and never edited there
> Reconstructed after the fact from the evidence cited below.
## Context
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh
database. But the things that consume those bindings — environment files, service
definitions, daemon configuration, firewall rules — are files on a node's disk, because that
is what the software reading them requires.
So the same value exists twice: authoritatively in the database, and materialised in a file.
Any edit to the file is a change to a copy. Before this decision, environment values could be
pushed from a node back into the database, which made the direction ambiguous in both
directions at once.
## Considered options
1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed.
Two writers and no arbiter: whichever synced last wins, and neither is authority.
2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts
ADR 0003 and returns to state that cannot be reconciled across nodes.
3. **Strictly one-directional: the database is written, files are generated.** Chosen.
## Decision
Every managed file is **derived**. A synchroniser regenerates it from the mesh database
whenever the underlying values change. The write path is the mesh tool that owns the value;
the file is an output.
This applies to generated environment files, service definitions, managed configuration, and
anything else a synchroniser lists as its own.
An edit to a managed file survives until the next synchronisation and is then overwritten,
without a warning, taking whatever it was fixing with it.
## Consequences
- **A file edited on a node is a bug with a delay on it.** This is now one of the mesh's core
values, and it is a consequence of this decision rather than a stance taken independently.
- To change a value you must know which tool owns it. That is a real cost, paid every time,
and the reason the mesh provides a way to ask whether a given file is managed.
- Debugging by editing a file no longer works, and fails in the most confusing way available:
it works, and then stops working later for no locally visible reason.
- Recovery is cheap. A node's entire managed surface can be regenerated from the database.
- Values resolve by precedence — database override, then existing file value, then generated,
then manifest default — which means an unset override does not clobber a generated
password. The subtlety is real and has caused its own confusion.
## References
- `Extract hal/env-sync module, remove syncEnvToDb`, 2026-04-03 — the commit that removed the
node-to-database direction.
- Knowledge base: `conventions/no-direct-file-mutation`, `provisioning` (resolution priority).
- Regeneration gaps: `troubleshooting/changed-manifest-default-not-rerendered`,
`troubleshooting/config-removed-from-manifest-not-pruned`.
@@ -0,0 +1,67 @@
---
status: accepted
date: 2026-04-06
deciders: jochen
reconstructed: true
---
# 5. Capabilities are provisioned on declaration, not configured by hand
> Reconstructed after the fact from the evidence cited below.
## Context
Most modules need something another module holds — a database, a cache, a bucket, a message
vhost, an identity client. Wiring that by hand means creating the resource, creating a user,
generating a credential, putting it in the consumer's configuration, and repeating all of it
on every node the consumer runs on.
Every step is a place to make a mistake that surfaces much later, and the credential ends up
written somewhere it can be read.
## Considered options
1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a
description of the mistakes people will make.
2. **A shared credential per resource type**, distributed to all consumers. Rejected: no
isolation, and rotation becomes a mesh-wide outage.
3. **Declared requirements, satisfied by the provider module.** Chosen.
## Decision
A module declares what it **provides** and what it **requires**. A requirement names the
provider, the resource type, optionally a name and a target node, and a mapping from the
resource's connection fields to the consumer's environment variables.
The mesh satisfies it: a provisioner belonging to the provider creates the resource and its
credential, records the grant, and writes the mapped values as database overrides. The
synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then
materialises them. Neither the credential nor the topology is ever written by hand.
A requirement may name a provider on another node. The grant records consumer and provider
nodes separately, so cross-node wiring is the same declaration.
## Consequences
- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a
capability; where it lives is the mesh's problem. This is the property
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model
around.
- Credentials are never authored, so they are never authored badly, and they are never in the
repository.
- Each consumer gets its own credential, so revocation is per-consumer.
- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the
peers holding the old one, and this has taken the mesh down. The declaration model makes
granting easy and says nothing about fan-out.
- A module with no requirements skips the stage entirely, which is correct and also means the
absence of provisioning is indistinguishable from provisioning that did not run.
## References
- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the
provisioner daemon becomes its own component.
- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then-
environment-then-start sequence becomes the coordinator's.
- Knowledge base: `provisioning`, `provisioning/requires`.
- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`,
`troubleshooting/provision-adoption-rotates-live-credential`.
@@ -0,0 +1,68 @@
---
status: accepted
date: 2026-05-14
deciders: jochen
reconstructed: true
---
# 6. Schema and state changes are numbered migrations, in the same language as the code
> Reconstructed after the fact from the evidence cited below.
## Context
Modules own persistent state. That state has to change as they change, across nodes that are
at different versions, some of which have data that predates the change.
Two things were being done that do not survive contact with a second node. Schema was created
at startup, so what a table looked like depended on which version last started. And migrations
were shell scripts, so they could not use the types, connection handling or helpers the module
already had, and were not compiled or checked with it.
## Considered options
1. **Startup SQL / create-if-missing.** Rejected. It converges only for a node that started
with the newest version. A node that never restarts never migrates; a node that restarts on
an old version can undo a change.
2. **Shell migrations.** Rejected. Unchecked, untyped, and a separate dialect from the module
they belong to. Also, as later discovered, packaged differently — and therefore
occasionally not packaged at all.
3. **Numbered migrations in the module's own language, compiled with it.** Chosen.
## Decision
Every schema or state change is a numbered migration file, written in the same language as the
module and compiled with it. There is no startup schema creation and no ad-hoc statement.
Rules that come with it:
- The initial migration is **frozen** once it has run anywhere. It is never modified; a change
is a new number.
- Every statement is **idempotent** — guarded so that re-running is safe.
- A change needs **both** a baseline for a fresh installation and an incremental migration for
installations that already exist. Code referencing a column requires that the migration
creating it exists.
- Migration numbers are unique. A duplicate prefix is a defect, not a style issue.
## Consequences
- A node at any version converges to the current schema by running the migrations it has not
run.
- Migrations are checked by the same compiler as the code, and a migration that does not
compile fails the build rather than the deployment.
- The rules are enforced unevenly. Duplicate prefixes have shipped repeatedly and been fixed by
renumbering afterwards; the mesh now validates for them, which is the check this rule needed
in order to be real.
- A migration directory is a feature like any other, which means it is packaged like any other
— and when packaging is wrong, migrations silently do not ship. This has happened.
- Freezing the initial migration means a fresh installation replays the entire history. That
cost grows and nothing currently bounds it.
## References
- `fix(agents): convert workflow migrations to TypeScript` (#50), 2026-05-14.
- `feat(dev_validate): guard against duplicate migration numeric prefixes` (#379), 2026-06-26 —
the rule acquiring a check.
- Knowledge base: `migrations/schema-drift`, `noxflow/troubleshooting/migration-number-collision`.
- Packaging failures: `troubleshooting/shell-migrations-never-packaged`,
`troubleshooting/provision-migration-never-applied`.
+65
View File
@@ -0,0 +1,65 @@
---
status: accepted
date: 2026-06-04
deciders: jochen
reconstructed: true
---
# 7. No workspace — each module is a standalone package consuming published dependencies
> Reconstructed after the fact from the evidence cited below.
## Context
Modules depend on each other, above all on the shared library every module builds against.
A workspace was the obvious way to express that: sibling packages, resolved locally, one
install at the root.
It produced a divergence that is worth stating precisely, because it is not obvious. In
development, a workspace member importing a sibling resolves to that sibling's **local source**.
In the pipeline, each module is built alone, from a clone, without its siblings present — so
the same import resolves to the **published version**. The two environments were therefore
building different code from identical source, and the failure appeared only in the pipeline,
in a module that had not been touched.
## Considered options
1. **Keep the workspace and make the pipeline replicate it** — clone every module, build the
graph. Rejected: it makes every build a whole-repository build, which is the cost the
per-module pipeline exists to avoid, and it does not extend to modules in their own
repositories.
2. **Keep the workspace and pin siblings to published versions.** Rejected as the worst of
both: the workspace's local resolution silently overrides the pin, so the divergence
remains while looking solved.
3. **No workspace. Every module is standalone and consumes published dependencies.** Chosen.
## Decision
There is no workspace. Each module is an independent package that declares its dependencies
and consumes them from the private registry, including the mesh's own shared library.
A cross-package change is therefore two steps: publish the producer, then consume it. The
pipeline does the first on push and resolves the levels so that a module always builds against
its dependencies' freshly published versions.
## Consequences
- Development and the pipeline resolve imports identically. The divergence is gone by
construction rather than by discipline.
- A module in its own repository is not a special case. It builds exactly as a module in the
monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md)
cheap.
- A cross-package change costs a publish-and-consume round trip. This is the real price, paid
on every shared-library change.
- There is no repository-wide install and no repository-wide build. Anything that assumed one
broke, and one thing that assumed one has stayed broken: the end-to-end pipeline harness has
not built since this decision landed. See
[`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).
## References
- `fix(noxflow): kill npm workspace, restore encryption inside PgAdminRepo` (#240),
2026-06-04. The reason is recorded in the root package manifest, which still carries the
note explaining why no workspace exists.
- The divergence it fixed is named there: workspace members importing each other resolved to
local unbuilt source in the pipeline.
+71
View File
@@ -0,0 +1,71 @@
---
status: accepted
date: 2026-06-05
deciders: jochen
reconstructed: true
---
# 8. A step that fails must fail the job
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh's expensive faults are not crashes. They are the operations that reported success and
did nothing: an artifact that partially downloaded and was extracted anyway, a package that
404ed from every mirror while the job went green, a hook that never ran because it was named
for a feature the module does not declare, a deploy that reported the transport succeeded
rather than that the effect happened.
Each of these was found long after it happened, by someone investigating an unrelated symptom.
The cost is not the failure; it is the interval between the failure and anyone learning of it,
during which decisions are made on the assumption that the thing worked.
## Considered options
1. **Continue on error and report at the end.** Rejected — it is largely what existed. A
summary nobody reads is not a report, and later steps run against the state the failed step
should have produced.
2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a
precise, located failure into a vague one discovered elsewhere, and requires a health check
for every possible partial state.
3. **Fail the step, fail the job, say which step.** Chosen.
## Decision
A step that fails stops the sequence it is part of, and the failure is surfaced where the work
was requested — not only in a log.
Concretely, and these are the forms it takes:
- A scripted sequence gates each step on the previous one. A directory change that fails must
stop the commands that assumed it.
- An artifact that does not fully download is not extracted.
- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must
mean the thing is running, not that a command returned.
- A template that cannot resolve a variable is not written half-rendered.
**Prefer failing to lying.** A green result that is not true costs more than a red one.
## Consequences
- Failures are noisier and land earlier, on the person who caused them.
- Some jobs that used to complete now stop. In every case examined so far, that job was
producing a partial result that something downstream trusted.
- This is a rule the mesh has adopted repeatedly rather than once, because each instance is
written in a different place — a shell hook, a download path, a deploy stage. It is not
enforced by a mechanism, and cannot currently be checked in general. New instances are still
being found; the package-install case remains open as
[`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md).
## References
- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05.
- `A flavor template with an unresolved variable is written to disk instead of failing`
(#710), 2026-08-08.
- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`,
`troubleshooting/service-started-is-not-ready`,
`troubleshooting/green-pipeline-means-transport-not-effect`,
`troubleshooting/silent-failures-and-stale-state`.
- The core value it became: [`00-GENESIS/mission.md`](../00-GENESIS/mission.md), "Failure must
be loud."
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 9. The mesh is governed by a constitution, injected where work is decided
> Reconstructed after the fact from the evidence cited below.
## Context
By mid-2026 the mesh was doing a large share of its own design and implementation work through
agents. The rules those agents were expected to follow existed — in operating instructions, in
convention documents, in the knowledge base — but they were **retrieved**: an agent had to know
a rule existed in order to look it up.
Rules that must be looked up are followed by whoever already knows them, which is precisely the
population that does not need them. The rules being violated were the ones nobody thought to
search for.
## Considered options
1. **Documentation plus review.** Rejected — it is what existed. Review catches a violation
after the work is done, and only if the reviewer knows the rule.
2. **Lint and automated checks only.** Rejected as insufficient, not wrong. A check catches
what can be expressed mechanically; most of these rules are about judgement — what belongs
in a repository, when a criterion counts as verified.
3. **A canonical rule set, injected into context wherever work is decided, with a check phase
before output is accepted.** Chosen.
## Decision
A single canonical document states the mesh's non-negotiable rules. It is **injected
proactively** into every eligible design and analysis session — agents do not fetch it, it
arrives — and a check phase verifies the session's output against it before the work proceeds.
It is a governed document, not a page. Changing it requires a proposal, sign-off by reviewers
who are not the proposer, and a recorded decision. Drive-by edits are reverted.
Scoped override pages may **tighten** it for a team or product. They may never relax it.
## Consequences
- A rule reaches the work whether or not anyone remembered it existed.
- The check phase makes a violation a blocking outcome rather than a review comment.
- Two copies of the same rules now exist: this document, and the reasoning in HQ that earned
them. The enforced copy wins by default, so the reasoned copy quietly stops being true —
which is why [`00-GENESIS/how-we-build.md`](../00-GENESIS/how-we-build.md) is now the source
and the governed page is derived from it, via playbook
[`05-constitution-sync.md`](../00-GENESIS/process/05-constitution-sync.md).
- Injection costs context on every eligible turn, and grows with the document. Nothing
currently bounds that.
- The amendment process requires two reviewers, which a mesh with one human operator satisfies
only by counting agents. That tension is real and unresolved.
## References
- The governed page was authored 2026-07-10 and carries its own amendment process.
- `feat(noxflow): HAL architectural conformance gate for reviewer + architect` (#296),
2026-06-10 — the check phase, predating the document it checks against.
- Knowledge base: `platform/constitution`.
@@ -0,0 +1,64 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 10. Applications live in their own repository; the monorepo is for the mesh
> Reconstructed after the fact from the evidence cited below.
## Context
The module system makes adding anything to the monorepo trivial — a directory and a manifest.
That ease is the problem. Standalone applications, sites and side-projects accumulated beside
the mesh's own components, and once there they inherited the monorepo's review cadence, its
pipeline detection, and its history.
The mesh's own code and an application that merely runs on the mesh have nothing in common
except the manifest format. They change for different reasons, are reviewed by different
criteria, and have no reason to share a branch.
## Considered options
1. **Everything in the monorepo.** Rejected — it is what existed. The monorepo becomes an
inventory of one installation's applications, and every application change queues behind
mesh review.
2. **A second monorepo for applications.** Rejected: the same coupling with an extra name.
Applications have no more in common with each other than with the mesh.
3. **One repository per application, registered with the mesh as a build source.** Chosen.
## Decision
Every standalone application, site or side-project lives in its own repository, with a manifest
at the root. It registers with the mesh as a build source and is then built, provisioned,
deployed and verified by exactly the same pipeline as anything in the monorepo.
The monorepo holds the mesh: the runtime, the core modules, the delivery machinery, and the
shared infrastructure the mesh itself provisions against.
Creating an application directory in the monorepo is a convention violation, and reviewers
reject it.
## Consequences
- An application's cadence is its own. It is not reviewed as mesh code and does not queue
behind mesh work.
- The separation is safe **only because** the pipeline and provisioning are identical either
side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without
standalone packages this decision would fork the build.
- The monorepo stops being an inventory of the installation, which is a precondition for
publishing anything about it.
- Discovery gets harder: there is no single listing of everything the mesh runs, and the
registry of build sources becomes the closest thing to one.
- A module source that is not registered is silently skipped by the pipeline. The cost of
being outside the monorepo is that being forgotten is possible.
## References
- The rule is stated in the governed constitution page authored 2026-07-10, §3, as a
convention violation reviewers must reject.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
applications to the modules already in the monorepo.
- Knowledge base: `troubleshooting/unregistered-module-source`.
+58
View File
@@ -0,0 +1,58 @@
---
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 11. The installer owns linking; nothing else creates a symlink
> Reconstructed after the fact from the evidence cited below. The incident that earned the rule
> predates the record, and its date is not established here.
## Context
A service's definition lives in the module catalogue; its runtime directory and persistent data
live outside it. The mesh connects the two by linking the definition into the runtime location
— deliberately, so that runtime state and source stay separate while the running service reads
a current definition.
A link is also the easiest thing in the world to create by hand while fixing something, and a
container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere
it should not have, and **production data was lost**.
## Considered options
1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss
for a service running a definition nobody can find.
2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of
the mistake — the link looks right and the resolution happens inside the container engine.
3. **One component owns linking; everyone else is forbidden.** Chosen.
## Decision
The installer creates and repairs every link the mesh needs. It reconciles them: a missing
link is created, a stale one is repointed, and a real file found where a link belongs is
adopted into the node's override location and replaced.
**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person
debugging. The prohibition is absolute because the judgement required to make a safe exception
is exactly the judgement that was not available at the moment it mattered.
## Consequences
- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who
has not seen the incident. That is why it is recorded here rather than only asserted.
- Links become reconcilable state rather than incidental filesystem facts.
- The rule is stated for humans and agents and is enforced by convention, not mechanism. A
check does not exist.
- The rule is regularly misread as "the mesh does not use symlinks", which is false and makes
the design documentation look self-contradictory. It uses them; it centralises who may make
them.
## References
- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or
service directories have caused production data loss via Docker volume path resolution. The
installer handles all linking. Never create symlinks manually."*
- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files.
@@ -0,0 +1,74 @@
---
status: accepted
date: 2026-07-12
deciders: jochen
reconstructed: true
---
# 12. An agent is a persistent employee, not an instance of a pool
> Reconstructed after the fact from the evidence cited below.
## Context
Agents were originally a **pool**: a named kind of worker, scaled to some number of
interchangeable instances. Work went to whichever instance was free.
That model has no place to put the things that turn out to matter. An agent that accumulates
knowledge of a domain cannot keep it, because the next task lands on a different instance. An
agent cannot own a workspace, because there are several of it. It cannot be held to a policy —
warned for a violation, then dismissed — because there is no continuing subject to warn.
Scaling was also solving a problem the mesh does not have. Instances were being multiplied to
get concurrency, when concurrency is a property of how much work one agent may hold at once.
## Considered options
1. **Keep the pool, attach memory to the pool.** Rejected: shared memory across
interchangeable workers is a knowledge base, not an agent's experience, and the mesh
already has one.
2. **Keep the pool, make instances sticky.** Rejected as a pool pretending to be identities —
identity by scheduling accident, lost on any restart.
3. **One agent is one persistent identity, with concurrency as a property of it.** Chosen.
## Decision
An agent is a **singular, named, persistent identity**: a home node, a workspace on that node,
accumulating memory, and a lifecycle — hired, active, draining, retired. Not a pool member.
Concurrency is a property of the agent, not a count of copies: an agent has a cap on how many
sessions it may hold at once.
Lifecycle is explicit and has verbs. An agent is hired onto a node; it may be reassigned while
idle; it is retired by draining first, and forced only deliberately. Retired agents are not
deleted.
Surge capacity is expressed within the model rather than against it: a template agent is a
blueprint, cloned into a real agent with a lifetime when a queue grows, drained and retired
when it expires. A temporary employee is still an employee.
Some agents are **human**. What differs is modality — how the agent acts — not category. A node
itself is an agent of a kind exempt from the hiring lifecycle.
## Consequences
- Memory, workspace and reputation have a subject to belong to. Policy becomes possible: an
agent that violates a rule can be warned, and warned agents can be dismissed.
- The mesh gained a hiring model, and with it the question of who may hire.
- Scaling by adding instances is gone. If one agent is saturated, either its session cap rises
or another agent is hired — both deliberate acts.
- The transition was not free. Lifecycle columns had to reach every query that selects an
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises:
one kind of participant, differing only in modality.
## References
- `docs(adr): agents as persistent employees + MINERVA librarian` (#495), 2026-07-12 — the
original record, in the code repository.
- `feat(B4): one persistent employee, N sessions — rename max_instances → max_sessions` (#547)
and `feat(noxflow): B3 — workspace provisioner for agent employee model` (#549), 2026-07-20.
- `feat(noxflow): warn-then-fire agents who merge to main without review` (#209), 2026-06-01 —
policy that presumes a continuing subject, predating the model that provides one.
- Knowledge base: `agents/employee-lifecycle`, `agents/temp-surge`, `agents/workspace-layout`.
- The migration cost: `troubleshooting`/`noxflow-agent-enriched-select-missing-lifecycle-columns` (#546).
+63
View File
@@ -0,0 +1,63 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 13. An artifact is build output, never a source tree
> Reconstructed after the fact from the evidence cited below.
## Context
A module is built once and deployed to every node assigned to it. What travels between those
two events is the artifact.
For a long time the artifact was a filtered copy of the module's source directory. Deploying it
therefore meant resolving and installing its dependencies **on the target node** — which
requires the target to reach a package registry, at deploy time, for every node, every deploy.
A node with no route to the registry could not deploy code that had already been built
successfully.
## Considered options
1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy
becomes a network operation with a failure mode per node, and the code that runs is
assembled independently on each one.
2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the
dependency resolution's platform assumptions along with it.
3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen.
## Decision
The artifact is the module's **build output directory** — compiled and bundled, with its
dependency graph inlined. Deploy is extract-and-run and touches no network.
A build that cannot produce a self-contained output **fails**. It does not fall back to
shipping a dependency tree, because a fallback that works is a fallback that is never fixed —
an application of [ADR 0008](0008-a-failed-step-fails-the-job.md).
## Consequences
- A node can deploy without reaching a registry. What was built is what runs, identically, on
every node.
- Deploys are faster and their failure modes are local.
- **Everything not in the build output does not ship.** This is the decision's whole cost, and
it was paid several times before it was understood: migrations that read the source layout,
provisioning scripts that read the source layout, selection files never packaged at all. Each
worked in development, where the source is present, and silently did nothing after deploy.
- Any file a module needs at runtime must be deliberately placed into the build output. The
rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code,
and that generalisation was the expensive part.
- Bundling has its own failure modes that a compiler will not catch — a bundler can exit
successfully and produce output that cannot load.
## References
- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04.
- The consequences, in order: `Provision migrations and seeds read the source layout, not the
artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07.
- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`,
`troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`,
`troubleshooting/esbuild-silent-tla-breakage`.
@@ -0,0 +1,78 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 14. Build, publish and deploy are three silos with different cardinality
> Reconstructed after the fact from the evidence cited below.
## Context
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
run a different number of times.
- Compiling happens **once per module feature**, on the build node.
- Packaging and uploading happens **once per module feature**, on the build node.
- Installing, configuring, starting and verifying happens **once per module feature per node**.
Conflating them is what made earlier versions slow and hard to reason about. Work that should
happen once was being repeated per node, and the fan-out point was implicit rather than a
boundary anything could observe.
The split had been declared before it was real. Packaging still happened inside the build,
which meant the boundary existed in the documentation and not in the code.
## Considered options
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
Cardinality is then a property of each stage's implementation, and nothing can reason about
the pipeline as a whole.
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
so build must know every module, every feature, and how each composes its artifact —
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
a stale package instead of re-packaging.
3. **Three silos, with an explicit handover between each.** Chosen.
## Decision
Delivery is three silos, and the boundaries are real:
| Silo | Runs | Where |
|---|---|---|
| **build** | once per module feature | the build node |
| **publish** | once per module feature | the build node |
| **deploy** | once per module feature **per node** | every assigned node |
Commands and events are addressed **per feature**, not per module.
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
publishing, so a module whose artifact is a package publishes in the publish silo, not the
build one.
Modules are resolved into dependency **levels**, and a level completes before the next begins,
so a module always builds against its dependencies' freshly published versions.
## Consequences
- Work that should happen once happens once. The fan-out point is explicit and observable.
- A failed upload retries by re-packaging, because packaging belongs to the stage that
uploads.
- The handover is a staged tree in a known location rather than the build's working directory,
which is reference-counted and cannot be assumed to still exist when a later stage runs.
- The build node is now the only node that has already passed through two silos when the
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
stages; code that knew only about the first parked the build node forever while every other
node deployed cleanly.
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
reports success over a stall it cannot see.
## References
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
documents claimed otherwise and were stale until 2026-08-06.
- The build-node stage-tracking failure was observed on pipeline #5557.
@@ -1,8 +1,11 @@
# 1. The mesh brokers capabilities; nodes host; agents think
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
- **Status:** Accepted
- **Date:** 2026-08-22
- **Deciders:** jochen
# 15. The mesh brokers capabilities; nodes host; agents think
## Context
@@ -1,12 +1,15 @@
# 2. A lab node is a virtual machine running the real install
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
- **Status:** Accepted
- **Date:** 2026-08-22
- **Deciders:** jochen
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
@@ -128,7 +131,7 @@ than staging means every certificate experiment on a real node consumes issuance
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`02-DESIGN/01-end-to-end-testing.md`](../02-DESIGN/01-end-to-end-testing.md) — what the lab
- [`02-DESIGN/01-end-to-end-testing.md`](../02-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
@@ -0,0 +1,93 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0015-mesh-brokers-nodes-host-agents-think.md
---
# 17. Modules outside the platform core are grouped by domain, not by single function
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules
into bounded contexts named after their aggregates, and sends the rest out of the monorepo on
the grounds that they run *on* the mesh rather than being *of* it.
That leaves the larger half unaddressed. Around three quarters of the catalogue are modules
that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an
SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of
media-library services. Today each is its own module, because one module is the unit of *one
piece of software*, and no other grouping exists.
The result is that the catalogue's shape records what was installed, not what anything is for.
Four modules that together constitute "how a node is reachable" have no relationship the mesh
can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a
change to how the mesh handles connectivity has to be made four times.
This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident
rather than by domain* — appearing outside it.
## Considered options
1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else
keeps accident boundaries, so the catalogue becomes harder to read after the refactor than
before it.
2. **One module per piece of software, with a tag or category field.** Rejected. A label is not
a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it
drifts from the thing it labels.
3. **Group them into domain modules, each owning the software that serves one purpose.**
Proposed here.
4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for
the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it
into that model repeats the metaphor-naming mistake ADR 0015 exists to correct.
## Decision
*Proposed — the principle is settled; the domain list is not. See "Open" below.*
Modules that are not part of the platform core are grouped into **domain modules**. A domain
is named for the concern it serves, and owns the software that serves it. The unit stops being
one piece of software and becomes one purpose.
This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own
domain stand unchanged. This decision covers what ADR 0015 leaves outside them.
Naming follows the same rule as the core: **name the domain for what it does, not for what it
is made of**. Connectivity, not a VPN implementation.
## Consequences
- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes
reach each other is a change to one module.
- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from
its module list.
- Swapping an implementation stops being a module replacement, with the data-volume and
provisioning consequences that carries, and becomes a change inside a domain.
- The count drops sharply, which is a symptom of the improvement rather than the point of it.
- **Grouping conceals.** A domain module hides which implementation is in use, and every
operational question — which port, which unit, which credential — gains an indirection.
- The migration is not free and has no obvious increments: a domain is only useful once
everything belonging to it has moved.
- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its
own sake would be the same error in the other direction.
## Open
**The domain list is not settled and this record does not invent one.** What is decided is the
principle; what is not decided is the set. Candidate groupings are visible in the catalogue —
connectivity and reachability, node presentation and desktop, media libraries, observation and
metrics, storage and data services — but naming them here would be reconstructing a decision
that has not been taken.
Settling the list is a research effort, not an act of this record. Until it concludes, this
ADR stays `proposed`.
## References
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this
extends, and its rule about naming a context after its aggregate.
- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are
already out of scope here; they are not domains and do not group.
- [`02-DESIGN/00-as-is/10-module-catalogue.md`](../02-DESIGN/00-as-is/10-module-catalogue.md)
— the catalogue's current shape, which is the evidence for the problem.
+40 -17
View File
@@ -1,30 +1,53 @@
# Architecture Decision Records
One file per decision, numbered, never deleted. A superseded ADR gets its status changed
and a pointer to what replaced it — the reasoning that was rejected is the expensive half
to rediscover.
One file per decision, numbered, never deleted. A superseded record has its `status:` changed
and gains a pointer to what replaced it — **its text is never edited**. The reasoning that was
rejected is the expensive half to rediscover.
## Format
The records are a **ledger**: they run in the order the decisions were taken, oldest first.
## Frontmatter
```yaml
---
status: proposed | accepted | superseded
date: YYYY-MM-DD # when the decision was taken, not when it was written down
deciders: name
reconstructed: true|false # true when the record was written after the fact from evidence
superseded-by: # adr/NNNN-....md, when status is superseded
extends: # adr/NNNN-....md, when this record widens an earlier one
---
```
## Body
```
# N. Title in plain language
- **Status:** Proposed | Accepted | Superseded by ADR-XXXX
- **Date:** YYYY-MM-DD
- **Deciders:**
## Context what is true today, with evidence
## Context what was true, with evidence
## Considered Options numbered, each with why it was rejected
## Decision what we are doing
## Consequences what follows, including what gets harder
## References code, data, prior art
## Decision what was decided
## Consequences what follows, including what got harder
## References commits, pull requests, knowledge-base entries, prior art
```
State evidence, not assertion. "Zero of 124 modules declare `brain` as a dependency"
outranks "the dependency rule is not followed".
State evidence, not assertion. *"Zero of 124 modules declare `brain` as a dependency"*
outranks *"the dependency rule is not followed"*.
## Reconstructed records
Records 0001–0014 were written on 2026-08-23, after the decisions they describe. Those
decisions were taken in implementation rather than in a document; the records state what was
decided and the evidence it was decided from, and each carries `reconstructed: true` and says
so in its first lines.
A reconstructed record is not a transcript. Where the deliberation is not recoverable, the
options section states what the alternatives were and why the chosen one won on the evidence
available — not a discussion that did not happen. Where a date is not establishable it says so
rather than guessing.
## Index
| ADR | Title | Status |
|-----|-------|--------|
| [0001](0001-mesh-brokers-nodes-host-agents-think.md) | The mesh brokers capabilities; nodes host; agents think | Accepted |
The index is **generated, not maintained** — run the `hal-status` skill, which reads the
frontmatter of every record. A hand-written index drifts from the folder it describes, and
this one had already done so after a single addition.