The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a scar, not a choice: 02-DESIGN existed from its initial commit, and when adr/ was finally promoted on 2026-07-13 it took the next free number rather than its place in the sequence. By then design was too settled to renumber. hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and 02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks the process in the order it happens: research produces a decision, the decision authorises a design. 00-GENESIS becomes 00-META, matching papa's rename from the same restructure. Every path reference rewritten across documents, frontmatter, playbooks and skills. All links resolve; all 58 frontmatter blocks parse and their path fields still point at files that exist.
This commit is contained in:
@@ -0,0 +1,68 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-02-25
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 1. Nodes communicate over a message broker, not over HTTP
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below. The decision was taken in
|
||||
> implementation, not in a record; this document states what was decided and why, not a
|
||||
> deliberation that happened.
|
||||
|
||||
## Context
|
||||
|
||||
The mesh is a set of machines that must call each other's capabilities. On the day the
|
||||
repository was founded there was no inter-node transport at all — each node was configured
|
||||
independently and shared nothing at runtime.
|
||||
|
||||
Three properties were required and are visible in everything built since:
|
||||
|
||||
- A node behind a household NAT must participate fully. It can dial out; nothing can dial in.
|
||||
- A node that is asleep, rebooting or upgrading must not cause a caller to fail — the request
|
||||
should wait, not error.
|
||||
- Adding a node must not require editing anything on the nodes that already exist.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **HTTP APIs between nodes.** Rejected. Every node becomes a server that every other node
|
||||
must be able to reach, which the NAT case makes impossible without inbound tunnels to each
|
||||
participant. It also makes node liveness a caller's problem: a request to a sleeping node
|
||||
is an error rather than a wait.
|
||||
2. **Polling a shared database.** Rejected. Latency is the poll interval, load is constant and
|
||||
independent of demand, and request/reply has to be built on top of it by hand.
|
||||
3. **A central message broker with per-node exchanges.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
All inter-node communication goes through a message broker. Every node owns a topic exchange
|
||||
named for itself and a request queue; a shared mesh exchange carries commands and events that
|
||||
are not addressed to one node.
|
||||
|
||||
Three message shapes, and only three:
|
||||
|
||||
- **RPC** — request/reply, for calling a capability that lives on another node.
|
||||
- **Commands** — instructions to do a stage of work, addressed by what is to be done.
|
||||
- **Events** — statements that something happened, addressed to nobody.
|
||||
|
||||
Every node dials the broker outbound. Nothing dials a node.
|
||||
|
||||
## Consequences
|
||||
|
||||
- NAT stops being an architectural concern. A node's reachability is a property of the broker
|
||||
connection, not of its network position.
|
||||
- A call to a node that is down waits in that node's queue instead of failing. This is usually
|
||||
right and occasionally the wrong thing entirely — a queued command for a node that never
|
||||
returns is a stall with no error, which is the failure shape this mesh keeps rediscovering.
|
||||
- The broker is a single point of failure and a single point of trust. Its credential is
|
||||
mesh-wide, so rotating it is a mesh-wide operation.
|
||||
- Tools never leave the host: a remote call proxies over the broker and the credentials stay
|
||||
where the capability is.
|
||||
|
||||
## References
|
||||
|
||||
- The broker was stood up on 2026-02-25, the second day of the repository.
|
||||
- Knowledge base: `mesh` (transport, exchanges, queue naming), `troubleshooting/amqp-credential-rotation`.
|
||||
- The stall shape is recorded in `troubleshooting/empty-pipeline-blocks-the-queue` and
|
||||
`troubleshooting/daemon-and-tool-server-share-a-request-queue`.
|
||||
@@ -0,0 +1,73 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-03-14
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 2. Everything is a module, and one manifest describes all of them
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
The mesh carries several kinds of thing: containerised services with data and ports, pure
|
||||
capability providers with no service at all, and bare markers whose only content is that a
|
||||
node has them. Before this decision these were separate concepts with separate handling —
|
||||
the earlier vocabulary was *capabilities*, and services were installed by a different path
|
||||
than tools.
|
||||
|
||||
Every distinct kind of thing needs its own install path, its own change detection, its own
|
||||
place in the delivery pipeline, and its own documentation. Three kinds means three of each,
|
||||
and every new feature has to be built three times or, more commonly, once — leaving two kinds
|
||||
quietly unsupported.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag
|
||||
list. Rejected: it is what existed, and the cost was paid in every cross-cutting change.
|
||||
2. **One manifest, kind inferred from directory contents.** Chosen.
|
||||
3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service
|
||||
still declares itself — but the general rule became inference, because a declared list and
|
||||
the directory it describes drift, and the directory is the one that is true.
|
||||
|
||||
## Decision
|
||||
|
||||
Everything the mesh installs is a **module**: a directory with a manifest. The manifest
|
||||
declares identity, environment variables, what the module provides, what it requires, and how
|
||||
it is exposed. What kind of module it is follows from what the directory contains:
|
||||
|
||||
| Contains | Is |
|
||||
|---|---|
|
||||
| a compose definition | a service |
|
||||
| a tools directory | a capability provider |
|
||||
| a daemon or unit directory | a long-running process |
|
||||
| a configs directory | a source of managed files |
|
||||
| nothing but a manifest | a flag — presence is the whole content |
|
||||
|
||||
A module may be several of these at once. Each is a **feature**, and the delivery pipeline
|
||||
addresses features, not modules.
|
||||
|
||||
The mesh's own components are modules on exactly these terms. They get no privileged install
|
||||
path, no separate registry, and no exemption from the pipeline.
|
||||
|
||||
## Consequences
|
||||
|
||||
- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches
|
||||
everything the mesh carries.
|
||||
- Dogfooding stops being a discipline and becomes structural: if the mesh's own components
|
||||
need an exception, the machinery is unfinished, and that is visible immediately.
|
||||
- Feature detection from directory contents means a directory rename silently changes what a
|
||||
module *is*. This has bitten repeatedly — a hook named for a feature the module does not
|
||||
have is skipped without complaint.
|
||||
- The manifest becomes load-bearing and grows. It is now the largest single point of
|
||||
coupling in the mesh.
|
||||
|
||||
## References
|
||||
|
||||
- `Rename capabilities → modules across the entire codebase`, 2026-03-14.
|
||||
- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive
|
||||
by conversion rather than by creation.
|
||||
- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`.
|
||||
- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`,
|
||||
`troubleshooting/health-check-tools-index-false-positive`.
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-04-02
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 3. The mesh database is the source of truth; the repository is node-agnostic
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
Two things must be known to run the mesh: **what exists** — which modules there are, what each
|
||||
declares, how each is built — and **what runs where** — which node hosts which module, with
|
||||
which settings, at which version.
|
||||
|
||||
The repository is the natural home of the first. It was initially also the home of the second:
|
||||
per-node directories held that node's configuration, and adopting a machine meant committing
|
||||
its files. That has three costs. A node cannot be changed without a commit, so runtime state
|
||||
and source share a review cadence they do not share a rhythm with. Two nodes cannot be
|
||||
reconciled, because nothing holds both. And the repository becomes an inventory of the
|
||||
installation, which is exactly the content that cannot be made public.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding
|
||||
change is a commit and a deploy, and the repository accumulates an inventory of one
|
||||
particular mesh.
|
||||
2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no
|
||||
authority: two nodes disagreeing have no arbiter, and drift is invisible until something
|
||||
breaks.
|
||||
3. **A mesh database as the single authority, cached locally for resilience.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
A single database holds every binding: which node hosts which module, at which selection, with
|
||||
which environment overrides, plus mesh-level settings that all nodes read. The runtime loads
|
||||
its configuration from that database at startup and falls back to a local cache when the
|
||||
database is unreachable.
|
||||
|
||||
**The repository defines what exists. The database defines what runs where.** No node-to-module
|
||||
mapping is ever committed.
|
||||
|
||||
A node is therefore not described anywhere in source. Bringing one into the mesh is a database
|
||||
operation.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The repository becomes node-agnostic, and can be published without disclosing an
|
||||
installation. This repository's public stance rests on that property.
|
||||
- A binding changes without a commit, a build, or a deploy.
|
||||
- The local cache means a node survives losing the database, but a node running from cache is
|
||||
running from a snapshot with no indication of its age. Divergence is silent by construction.
|
||||
- The database is the hardest dependency in the mesh. It is also a module, provisioned like
|
||||
any other, which makes its bootstrap circular — resolved by the first-node initialisation
|
||||
script, and the reason such a script exists.
|
||||
- Nothing on a node is authoritative. That is what makes the next decision necessary.
|
||||
|
||||
## References
|
||||
|
||||
- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration
|
||||
tables that landed with it.
|
||||
- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments."
|
||||
- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`.
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-04-03
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 4. Managed files are generated onto nodes and never edited there
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh
|
||||
database. But the things that consume those bindings — environment files, service
|
||||
definitions, daemon configuration, firewall rules — are files on a node's disk, because that
|
||||
is what the software reading them requires.
|
||||
|
||||
So the same value exists twice: authoritatively in the database, and materialised in a file.
|
||||
Any edit to the file is a change to a copy. Before this decision, environment values could be
|
||||
pushed from a node back into the database, which made the direction ambiguous in both
|
||||
directions at once.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed.
|
||||
Two writers and no arbiter: whichever synced last wins, and neither is authority.
|
||||
2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts
|
||||
ADR 0003 and returns to state that cannot be reconciled across nodes.
|
||||
3. **Strictly one-directional: the database is written, files are generated.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
Every managed file is **derived**. A synchroniser regenerates it from the mesh database
|
||||
whenever the underlying values change. The write path is the mesh tool that owns the value;
|
||||
the file is an output.
|
||||
|
||||
This applies to generated environment files, service definitions, managed configuration, and
|
||||
anything else a synchroniser lists as its own.
|
||||
|
||||
An edit to a managed file survives until the next synchronisation and is then overwritten,
|
||||
without a warning, taking whatever it was fixing with it.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **A file edited on a node is a bug with a delay on it.** This is now one of the mesh's core
|
||||
values, and it is a consequence of this decision rather than a stance taken independently.
|
||||
- To change a value you must know which tool owns it. That is a real cost, paid every time,
|
||||
and the reason the mesh provides a way to ask whether a given file is managed.
|
||||
- Debugging by editing a file no longer works, and fails in the most confusing way available:
|
||||
it works, and then stops working later for no locally visible reason.
|
||||
- Recovery is cheap. A node's entire managed surface can be regenerated from the database.
|
||||
- Values resolve by precedence — database override, then existing file value, then generated,
|
||||
then manifest default — which means an unset override does not clobber a generated
|
||||
password. The subtlety is real and has caused its own confusion.
|
||||
|
||||
## References
|
||||
|
||||
- `Extract hal/env-sync module, remove syncEnvToDb`, 2026-04-03 — the commit that removed the
|
||||
node-to-database direction.
|
||||
- Knowledge base: `conventions/no-direct-file-mutation`, `provisioning` (resolution priority).
|
||||
- Regeneration gaps: `troubleshooting/changed-manifest-default-not-rerendered`,
|
||||
`troubleshooting/config-removed-from-manifest-not-pruned`.
|
||||
@@ -0,0 +1,67 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-04-06
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 5. Capabilities are provisioned on declaration, not configured by hand
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
Most modules need something another module holds — a database, a cache, a bucket, a message
|
||||
vhost, an identity client. Wiring that by hand means creating the resource, creating a user,
|
||||
generating a credential, putting it in the consumer's configuration, and repeating all of it
|
||||
on every node the consumer runs on.
|
||||
|
||||
Every step is a place to make a mistake that surfaces much later, and the credential ends up
|
||||
written somewhere it can be read.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a
|
||||
description of the mistakes people will make.
|
||||
2. **A shared credential per resource type**, distributed to all consumers. Rejected: no
|
||||
isolation, and rotation becomes a mesh-wide outage.
|
||||
3. **Declared requirements, satisfied by the provider module.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
A module declares what it **provides** and what it **requires**. A requirement names the
|
||||
provider, the resource type, optionally a name and a target node, and a mapping from the
|
||||
resource's connection fields to the consumer's environment variables.
|
||||
|
||||
The mesh satisfies it: a provisioner belonging to the provider creates the resource and its
|
||||
credential, records the grant, and writes the mapped values as database overrides. The
|
||||
synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then
|
||||
materialises them. Neither the credential nor the topology is ever written by hand.
|
||||
|
||||
A requirement may name a provider on another node. The grant records consumer and provider
|
||||
nodes separately, so cross-node wiring is the same declaration.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a
|
||||
capability; where it lives is the mesh's problem. This is the property
|
||||
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model
|
||||
around.
|
||||
- Credentials are never authored, so they are never authored badly, and they are never in the
|
||||
repository.
|
||||
- Each consumer gets its own credential, so revocation is per-consumer.
|
||||
- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the
|
||||
peers holding the old one, and this has taken the mesh down. The declaration model makes
|
||||
granting easy and says nothing about fan-out.
|
||||
- A module with no requirements skips the stage entirely, which is correct and also means the
|
||||
absence of provisioning is indistinguishable from provisioning that did not run.
|
||||
|
||||
## References
|
||||
|
||||
- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the
|
||||
provisioner daemon becomes its own component.
|
||||
- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then-
|
||||
environment-then-start sequence becomes the coordinator's.
|
||||
- Knowledge base: `provisioning`, `provisioning/requires`.
|
||||
- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`,
|
||||
`troubleshooting/provision-adoption-rotates-live-credential`.
|
||||
@@ -0,0 +1,68 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-05-14
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 6. Schema and state changes are numbered migrations, in the same language as the code
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
Modules own persistent state. That state has to change as they change, across nodes that are
|
||||
at different versions, some of which have data that predates the change.
|
||||
|
||||
Two things were being done that do not survive contact with a second node. Schema was created
|
||||
at startup, so what a table looked like depended on which version last started. And migrations
|
||||
were shell scripts, so they could not use the types, connection handling or helpers the module
|
||||
already had, and were not compiled or checked with it.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Startup SQL / create-if-missing.** Rejected. It converges only for a node that started
|
||||
with the newest version. A node that never restarts never migrates; a node that restarts on
|
||||
an old version can undo a change.
|
||||
2. **Shell migrations.** Rejected. Unchecked, untyped, and a separate dialect from the module
|
||||
they belong to. Also, as later discovered, packaged differently — and therefore
|
||||
occasionally not packaged at all.
|
||||
3. **Numbered migrations in the module's own language, compiled with it.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
Every schema or state change is a numbered migration file, written in the same language as the
|
||||
module and compiled with it. There is no startup schema creation and no ad-hoc statement.
|
||||
|
||||
Rules that come with it:
|
||||
|
||||
- The initial migration is **frozen** once it has run anywhere. It is never modified; a change
|
||||
is a new number.
|
||||
- Every statement is **idempotent** — guarded so that re-running is safe.
|
||||
- A change needs **both** a baseline for a fresh installation and an incremental migration for
|
||||
installations that already exist. Code referencing a column requires that the migration
|
||||
creating it exists.
|
||||
- Migration numbers are unique. A duplicate prefix is a defect, not a style issue.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A node at any version converges to the current schema by running the migrations it has not
|
||||
run.
|
||||
- Migrations are checked by the same compiler as the code, and a migration that does not
|
||||
compile fails the build rather than the deployment.
|
||||
- The rules are enforced unevenly. Duplicate prefixes have shipped repeatedly and been fixed by
|
||||
renumbering afterwards; the mesh now validates for them, which is the check this rule needed
|
||||
in order to be real.
|
||||
- A migration directory is a feature like any other, which means it is packaged like any other
|
||||
— and when packaging is wrong, migrations silently do not ship. This has happened.
|
||||
- Freezing the initial migration means a fresh installation replays the entire history. That
|
||||
cost grows and nothing currently bounds it.
|
||||
|
||||
## References
|
||||
|
||||
- `fix(agents): convert workflow migrations to TypeScript` (#50), 2026-05-14.
|
||||
- `feat(dev_validate): guard against duplicate migration numeric prefixes` (#379), 2026-06-26 —
|
||||
the rule acquiring a check.
|
||||
- Knowledge base: `migrations/schema-drift`, `noxflow/troubleshooting/migration-number-collision`.
|
||||
- Packaging failures: `troubleshooting/shell-migrations-never-packaged`,
|
||||
`troubleshooting/provision-migration-never-applied`.
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-06-04
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 7. No workspace — each module is a standalone package consuming published dependencies
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
Modules depend on each other, above all on the shared library every module builds against.
|
||||
A workspace was the obvious way to express that: sibling packages, resolved locally, one
|
||||
install at the root.
|
||||
|
||||
It produced a divergence that is worth stating precisely, because it is not obvious. In
|
||||
development, a workspace member importing a sibling resolves to that sibling's **local source**.
|
||||
In the pipeline, each module is built alone, from a clone, without its siblings present — so
|
||||
the same import resolves to the **published version**. The two environments were therefore
|
||||
building different code from identical source, and the failure appeared only in the pipeline,
|
||||
in a module that had not been touched.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Keep the workspace and make the pipeline replicate it** — clone every module, build the
|
||||
graph. Rejected: it makes every build a whole-repository build, which is the cost the
|
||||
per-module pipeline exists to avoid, and it does not extend to modules in their own
|
||||
repositories.
|
||||
2. **Keep the workspace and pin siblings to published versions.** Rejected as the worst of
|
||||
both: the workspace's local resolution silently overrides the pin, so the divergence
|
||||
remains while looking solved.
|
||||
3. **No workspace. Every module is standalone and consumes published dependencies.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
There is no workspace. Each module is an independent package that declares its dependencies
|
||||
and consumes them from the private registry, including the mesh's own shared library.
|
||||
|
||||
A cross-package change is therefore two steps: publish the producer, then consume it. The
|
||||
pipeline does the first on push and resolves the levels so that a module always builds against
|
||||
its dependencies' freshly published versions.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Development and the pipeline resolve imports identically. The divergence is gone by
|
||||
construction rather than by discipline.
|
||||
- A module in its own repository is not a special case. It builds exactly as a module in the
|
||||
monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md)
|
||||
cheap.
|
||||
- A cross-package change costs a publish-and-consume round trip. This is the real price, paid
|
||||
on every shared-library change.
|
||||
- There is no repository-wide install and no repository-wide build. Anything that assumed one
|
||||
broke, and one thing that assumed one has stayed broken: the end-to-end pipeline harness has
|
||||
not built since this decision landed. See
|
||||
[`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md).
|
||||
|
||||
## References
|
||||
|
||||
- `fix(noxflow): kill npm workspace, restore encryption inside PgAdminRepo` (#240),
|
||||
2026-06-04. The reason is recorded in the root package manifest, which still carries the
|
||||
note explaining why no workspace exists.
|
||||
- The divergence it fixed is named there: workspace members importing each other resolved to
|
||||
local unbuilt source in the pipeline.
|
||||
@@ -0,0 +1,71 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-06-05
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 8. A step that fails must fail the job
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
The mesh's expensive faults are not crashes. They are the operations that reported success and
|
||||
did nothing: an artifact that partially downloaded and was extracted anyway, a package that
|
||||
404ed from every mirror while the job went green, a hook that never ran because it was named
|
||||
for a feature the module does not declare, a deploy that reported the transport succeeded
|
||||
rather than that the effect happened.
|
||||
|
||||
Each of these was found long after it happened, by someone investigating an unrelated symptom.
|
||||
The cost is not the failure; it is the interval between the failure and anyone learning of it,
|
||||
during which decisions are made on the assumption that the thing worked.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Continue on error and report at the end.** Rejected — it is largely what existed. A
|
||||
summary nobody reads is not a report, and later steps run against the state the failed step
|
||||
should have produced.
|
||||
2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a
|
||||
precise, located failure into a vague one discovered elsewhere, and requires a health check
|
||||
for every possible partial state.
|
||||
3. **Fail the step, fail the job, say which step.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
A step that fails stops the sequence it is part of, and the failure is surfaced where the work
|
||||
was requested — not only in a log.
|
||||
|
||||
Concretely, and these are the forms it takes:
|
||||
|
||||
- A scripted sequence gates each step on the previous one. A directory change that fails must
|
||||
stop the commands that assumed it.
|
||||
- An artifact that does not fully download is not extracted.
|
||||
- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must
|
||||
mean the thing is running, not that a command returned.
|
||||
- A template that cannot resolve a variable is not written half-rendered.
|
||||
|
||||
**Prefer failing to lying.** A green result that is not true costs more than a red one.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Failures are noisier and land earlier, on the person who caused them.
|
||||
- Some jobs that used to complete now stop. In every case examined so far, that job was
|
||||
producing a partial result that something downstream trusted.
|
||||
- This is a rule the mesh has adopted repeatedly rather than once, because each instance is
|
||||
written in a different place — a shell hook, a download path, a deploy stage. It is not
|
||||
enforced by a mechanism, and cannot currently be checked in general. New instances are still
|
||||
being found; the package-install case remains open as
|
||||
[`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md).
|
||||
|
||||
## References
|
||||
|
||||
- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05.
|
||||
- `A flavor template with an unresolved variable is written to disk instead of failing`
|
||||
(#710), 2026-08-08.
|
||||
- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`,
|
||||
`troubleshooting/service-started-is-not-ready`,
|
||||
`troubleshooting/green-pipeline-means-transport-not-effect`,
|
||||
`troubleshooting/silent-failures-and-stale-state`.
|
||||
- The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must
|
||||
be loud."
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-07-10
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 9. The mesh is governed by a constitution, injected where work is decided
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
By mid-2026 the mesh was doing a large share of its own design and implementation work through
|
||||
agents. The rules those agents were expected to follow existed — in operating instructions, in
|
||||
convention documents, in the knowledge base — but they were **retrieved**: an agent had to know
|
||||
a rule existed in order to look it up.
|
||||
|
||||
Rules that must be looked up are followed by whoever already knows them, which is precisely the
|
||||
population that does not need them. The rules being violated were the ones nobody thought to
|
||||
search for.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Documentation plus review.** Rejected — it is what existed. Review catches a violation
|
||||
after the work is done, and only if the reviewer knows the rule.
|
||||
2. **Lint and automated checks only.** Rejected as insufficient, not wrong. A check catches
|
||||
what can be expressed mechanically; most of these rules are about judgement — what belongs
|
||||
in a repository, when a criterion counts as verified.
|
||||
3. **A canonical rule set, injected into context wherever work is decided, with a check phase
|
||||
before output is accepted.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
A single canonical document states the mesh's non-negotiable rules. It is **injected
|
||||
proactively** into every eligible design and analysis session — agents do not fetch it, it
|
||||
arrives — and a check phase verifies the session's output against it before the work proceeds.
|
||||
|
||||
It is a governed document, not a page. Changing it requires a proposal, sign-off by reviewers
|
||||
who are not the proposer, and a recorded decision. Drive-by edits are reverted.
|
||||
|
||||
Scoped override pages may **tighten** it for a team or product. They may never relax it.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A rule reaches the work whether or not anyone remembered it existed.
|
||||
- The check phase makes a violation a blocking outcome rather than a review comment.
|
||||
- Two copies of the same rules now exist: this document, and the reasoning in HQ that earned
|
||||
them. The enforced copy wins by default, so the reasoned copy quietly stops being true —
|
||||
which is why [`00-META/how-we-build.md`](../00-META/how-we-build.md) is now the source
|
||||
and the governed page is derived from it, via playbook
|
||||
[`05-constitution-sync.md`](../00-META/process/05-constitution-sync.md).
|
||||
- Injection costs context on every eligible turn, and grows with the document. Nothing
|
||||
currently bounds that.
|
||||
- The amendment process requires two reviewers, which a mesh with one human operator satisfies
|
||||
only by counting agents. That tension is real and unresolved.
|
||||
|
||||
## References
|
||||
|
||||
- The governed page was authored 2026-07-10 and carries its own amendment process.
|
||||
- `feat(noxflow): HAL architectural conformance gate for reviewer + architect` (#296),
|
||||
2026-06-10 — the check phase, predating the document it checks against.
|
||||
- Knowledge base: `platform/constitution`.
|
||||
@@ -0,0 +1,64 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-07-10
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 10. Applications live in their own repository; the monorepo is for the mesh
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
The module system makes adding anything to the monorepo trivial — a directory and a manifest.
|
||||
That ease is the problem. Standalone applications, sites and side-projects accumulated beside
|
||||
the mesh's own components, and once there they inherited the monorepo's review cadence, its
|
||||
pipeline detection, and its history.
|
||||
|
||||
The mesh's own code and an application that merely runs on the mesh have nothing in common
|
||||
except the manifest format. They change for different reasons, are reviewed by different
|
||||
criteria, and have no reason to share a branch.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Everything in the monorepo.** Rejected — it is what existed. The monorepo becomes an
|
||||
inventory of one installation's applications, and every application change queues behind
|
||||
mesh review.
|
||||
2. **A second monorepo for applications.** Rejected: the same coupling with an extra name.
|
||||
Applications have no more in common with each other than with the mesh.
|
||||
3. **One repository per application, registered with the mesh as a build source.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
Every standalone application, site or side-project lives in its own repository, with a manifest
|
||||
at the root. It registers with the mesh as a build source and is then built, provisioned,
|
||||
deployed and verified by exactly the same pipeline as anything in the monorepo.
|
||||
|
||||
The monorepo holds the mesh: the runtime, the core modules, the delivery machinery, and the
|
||||
shared infrastructure the mesh itself provisions against.
|
||||
|
||||
Creating an application directory in the monorepo is a convention violation, and reviewers
|
||||
reject it.
|
||||
|
||||
## Consequences
|
||||
|
||||
- An application's cadence is its own. It is not reviewed as mesh code and does not queue
|
||||
behind mesh work.
|
||||
- The separation is safe **only because** the pipeline and provisioning are identical either
|
||||
side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without
|
||||
standalone packages this decision would fork the build.
|
||||
- The monorepo stops being an inventory of the installation, which is a precondition for
|
||||
publishing anything about it.
|
||||
- Discovery gets harder: there is no single listing of everything the mesh runs, and the
|
||||
registry of build sources becomes the closest thing to one.
|
||||
- A module source that is not registered is silently skipped by the pipeline. The cost of
|
||||
being outside the monorepo is that being forgotten is possible.
|
||||
|
||||
## References
|
||||
|
||||
- The rule is stated in the governed constitution page authored 2026-07-10, §3, as a
|
||||
convention violation reviewers must reject.
|
||||
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
|
||||
applications to the modules already in the monorepo.
|
||||
- Knowledge base: `troubleshooting/unregistered-module-source`.
|
||||
@@ -0,0 +1,59 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-07-10
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 11. The installer owns linking; nothing else creates a symlink
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below. The incident that earned the rule
|
||||
> predates the record, and its date is not established here.
|
||||
|
||||
## Context
|
||||
|
||||
A service's definition lives in the module catalogue; its runtime directory and persistent data
|
||||
live outside it. The mesh connects the two by linking the definition into the runtime location
|
||||
— deliberately, so that runtime state and source stay separate while the running service reads
|
||||
a current definition.
|
||||
|
||||
A link is also the easiest thing in the world to create by hand while fixing something, and a
|
||||
container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere
|
||||
it should not have, and **production data was lost**.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss
|
||||
for a service running a definition nobody can find.
|
||||
2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of
|
||||
the mistake — the link looks right and the resolution happens inside the container engine.
|
||||
3. **One component owns linking; everyone else is forbidden.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
The installer creates and repairs every link the mesh needs. It reconciles them: a missing
|
||||
link is created, a stale one is repointed, and a real file found where a link belongs is
|
||||
adopted into the node's override location and replaced.
|
||||
|
||||
**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person
|
||||
debugging. The prohibition is absolute because the judgement required to make a safe exception
|
||||
is exactly the judgement that was not available at the moment it mattered.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who
|
||||
has not seen the incident. That is why it is recorded here rather than only asserted.
|
||||
- Links become reconcilable state rather than incidental filesystem facts.
|
||||
- The rule is stated for humans and agents and is enforced by convention, not mechanism. A
|
||||
check does not exist.
|
||||
- The rule as written governs the mechanism rather than removing it. A link made by the
|
||||
installer resolves the same way as one made by hand, so the hazard is narrowed and not
|
||||
closed. [ADR 0018](0018-the-mesh-creates-no-symlinks.md) proposes widening this to "nothing
|
||||
links, the installer included"; until that is accepted, this record governs.
|
||||
|
||||
## References
|
||||
|
||||
- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or
|
||||
service directories have caused production data loss via Docker volume path resolution. The
|
||||
installer handles all linking. Never create symlinks manually."*
|
||||
- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files.
|
||||
@@ -0,0 +1,74 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-07-12
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 12. An agent is a persistent employee, not an instance of a pool
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
Agents were originally a **pool**: a named kind of worker, scaled to some number of
|
||||
interchangeable instances. Work went to whichever instance was free.
|
||||
|
||||
That model has no place to put the things that turn out to matter. An agent that accumulates
|
||||
knowledge of a domain cannot keep it, because the next task lands on a different instance. An
|
||||
agent cannot own a workspace, because there are several of it. It cannot be held to a policy —
|
||||
warned for a violation, then dismissed — because there is no continuing subject to warn.
|
||||
|
||||
Scaling was also solving a problem the mesh does not have. Instances were being multiplied to
|
||||
get concurrency, when concurrency is a property of how much work one agent may hold at once.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Keep the pool, attach memory to the pool.** Rejected: shared memory across
|
||||
interchangeable workers is a knowledge base, not an agent's experience, and the mesh
|
||||
already has one.
|
||||
2. **Keep the pool, make instances sticky.** Rejected as a pool pretending to be identities —
|
||||
identity by scheduling accident, lost on any restart.
|
||||
3. **One agent is one persistent identity, with concurrency as a property of it.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
An agent is a **singular, named, persistent identity**: a home node, a workspace on that node,
|
||||
accumulating memory, and a lifecycle — hired, active, draining, retired. Not a pool member.
|
||||
|
||||
Concurrency is a property of the agent, not a count of copies: an agent has a cap on how many
|
||||
sessions it may hold at once.
|
||||
|
||||
Lifecycle is explicit and has verbs. An agent is hired onto a node; it may be reassigned while
|
||||
idle; it is retired by draining first, and forced only deliberately. Retired agents are not
|
||||
deleted.
|
||||
|
||||
Surge capacity is expressed within the model rather than against it: a template agent is a
|
||||
blueprint, cloned into a real agent with a lifetime when a queue grows, drained and retired
|
||||
when it expires. A temporary employee is still an employee.
|
||||
|
||||
Some agents are **human**. What differs is modality — how the agent acts — not category. A node
|
||||
itself is an agent of a kind exempt from the hiring lifecycle.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Memory, workspace and reputation have a subject to belong to. Policy becomes possible: an
|
||||
agent that violates a rule can be warned, and warned agents can be dismissed.
|
||||
- The mesh gained a hiring model, and with it the question of who may hire.
|
||||
- Scaling by adding instances is gone. If one agent is saturated, either its session cap rises
|
||||
or another agent is hired — both deliberate acts.
|
||||
- The transition was not free. Lifecycle columns had to reach every query that selects an
|
||||
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
|
||||
- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises:
|
||||
one kind of participant, differing only in modality.
|
||||
|
||||
## References
|
||||
|
||||
- `docs(adr): agents as persistent employees + MINERVA librarian` (#495), 2026-07-12 — the
|
||||
original record, in the code repository.
|
||||
- `feat(B4): one persistent employee, N sessions — rename max_instances → max_sessions` (#547)
|
||||
and `feat(noxflow): B3 — workspace provisioner for agent employee model` (#549), 2026-07-20.
|
||||
- `feat(noxflow): warn-then-fire agents who merge to main without review` (#209), 2026-06-01 —
|
||||
policy that presumes a continuing subject, predating the model that provides one.
|
||||
- Knowledge base: `agents/employee-lifecycle`, `agents/temp-surge`, `agents/workspace-layout`.
|
||||
- The migration cost: `troubleshooting`/`noxflow-agent-enriched-select-missing-lifecycle-columns` (#546).
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-04
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 13. An artifact is build output, never a source tree
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
A module is built once and deployed to every node assigned to it. What travels between those
|
||||
two events is the artifact.
|
||||
|
||||
For a long time the artifact was a filtered copy of the module's source directory. Deploying it
|
||||
therefore meant resolving and installing its dependencies **on the target node** — which
|
||||
requires the target to reach a package registry, at deploy time, for every node, every deploy.
|
||||
A node with no route to the registry could not deploy code that had already been built
|
||||
successfully.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy
|
||||
becomes a network operation with a failure mode per node, and the code that runs is
|
||||
assembled independently on each one.
|
||||
2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the
|
||||
dependency resolution's platform assumptions along with it.
|
||||
3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
The artifact is the module's **build output directory** — compiled and bundled, with its
|
||||
dependency graph inlined. Deploy is extract-and-run and touches no network.
|
||||
|
||||
A build that cannot produce a self-contained output **fails**. It does not fall back to
|
||||
shipping a dependency tree, because a fallback that works is a fallback that is never fixed —
|
||||
an application of [ADR 0008](0008-a-failed-step-fails-the-job.md).
|
||||
|
||||
## Consequences
|
||||
|
||||
- A node can deploy without reaching a registry. What was built is what runs, identically, on
|
||||
every node.
|
||||
- Deploys are faster and their failure modes are local.
|
||||
- **Everything not in the build output does not ship.** This is the decision's whole cost, and
|
||||
it was paid several times before it was understood: migrations that read the source layout,
|
||||
provisioning scripts that read the source layout, selection files never packaged at all. Each
|
||||
worked in development, where the source is present, and silently did nothing after deploy.
|
||||
- Any file a module needs at runtime must be deliberately placed into the build output. The
|
||||
rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code,
|
||||
and that generalisation was the expensive part.
|
||||
- Bundling has its own failure modes that a compiler will not catch — a bundler can exit
|
||||
successfully and produce output that cannot load.
|
||||
|
||||
## References
|
||||
|
||||
- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04.
|
||||
- The consequences, in order: `Provision migrations and seeds read the source layout, not the
|
||||
artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07.
|
||||
- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`,
|
||||
`troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`,
|
||||
`troubleshooting/esbuild-silent-tla-breakage`.
|
||||
@@ -0,0 +1,78 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-04
|
||||
deciders: jochen
|
||||
reconstructed: true
|
||||
---
|
||||
|
||||
# 14. Build, publish and deploy are three silos with different cardinality
|
||||
|
||||
> Reconstructed after the fact from the evidence cited below.
|
||||
|
||||
## Context
|
||||
|
||||
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
|
||||
run a different number of times.
|
||||
|
||||
- Compiling happens **once per module feature**, on the build node.
|
||||
- Packaging and uploading happens **once per module feature**, on the build node.
|
||||
- Installing, configuring, starting and verifying happens **once per module feature per node**.
|
||||
|
||||
Conflating them is what made earlier versions slow and hard to reason about. Work that should
|
||||
happen once was being repeated per node, and the fan-out point was implicit rather than a
|
||||
boundary anything could observe.
|
||||
|
||||
The split had been declared before it was real. Packaging still happened inside the build,
|
||||
which meant the boundary existed in the documentation and not in the code.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
|
||||
Cardinality is then a property of each stage's implementation, and nothing can reason about
|
||||
the pipeline as a whole.
|
||||
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
|
||||
so build must know every module, every feature, and how each composes its artifact —
|
||||
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
|
||||
a stale package instead of re-packaging.
|
||||
3. **Three silos, with an explicit handover between each.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
Delivery is three silos, and the boundaries are real:
|
||||
|
||||
| Silo | Runs | Where |
|
||||
|---|---|---|
|
||||
| **build** | once per module feature | the build node |
|
||||
| **publish** | once per module feature | the build node |
|
||||
| **deploy** | once per module feature **per node** | every assigned node |
|
||||
|
||||
Commands and events are addressed **per feature**, not per module.
|
||||
|
||||
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
|
||||
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
|
||||
publishing, so a module whose artifact is a package publishes in the publish silo, not the
|
||||
build one.
|
||||
|
||||
Modules are resolved into dependency **levels**, and a level completes before the next begins,
|
||||
so a module always builds against its dependencies' freshly published versions.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Work that should happen once happens once. The fan-out point is explicit and observable.
|
||||
- A failed upload retries by re-packaging, because packaging belongs to the stage that
|
||||
uploads.
|
||||
- The handover is a staged tree in a known location rather than the build's working directory,
|
||||
which is reference-counted and cannot be assumed to still exist when a later stage runs.
|
||||
- The build node is now the only node that has already passed through two silos when the
|
||||
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
|
||||
stages; code that knew only about the first parked the build node forever while every other
|
||||
node deployed cleanly.
|
||||
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
|
||||
reports success over a stall it cannot see.
|
||||
|
||||
## References
|
||||
|
||||
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
|
||||
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
|
||||
documents claimed otherwise and were stale until 2026-08-06.
|
||||
- The build-node stage-tracking failure was observed on pipeline #5557.
|
||||
@@ -0,0 +1,188 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-22
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
---
|
||||
|
||||
# 15. The mesh brokers capabilities; nodes host; agents think
|
||||
|
||||
## Context
|
||||
|
||||
HAL has 124 modules. The count is not the problem — it is the symptom. Modules are split
|
||||
because splitting is the only granularity the platform offers, and domains are merged
|
||||
because a shared database is the only integration it offers. Both pressures push in the
|
||||
same direction: boundaries end up drawn by deployment accident rather than by domain.
|
||||
|
||||
Three observations establish the state.
|
||||
|
||||
**The word "agent" means two different things.** The original design treated a node as an
|
||||
agent with thinking abilities. Later, noxflow implemented agents as employees with
|
||||
skills, workspaces and tasks. Both survive. The collision is visible in the data — there
|
||||
are **two agent rows per node**:
|
||||
|
||||
| agent | skills |
|
||||
|---|---|
|
||||
| one named after the node | `{deploy,verify,operate}` |
|
||||
| one named `hal-<node>` | `{}` |
|
||||
|
||||
One carries the work; the other carries only identity, existing to hold a licence for
|
||||
HAL's own sessions. The same split appears in the schema: `nodes.hal_claude_account` and
|
||||
`agents.claude_account` are one fact in two tables in two databases, and what was
|
||||
documented as "three licence touchpoints" is one concept modelled three times.
|
||||
|
||||
**Domains integrate by sharing a schema.** `noxflow` is 45 tables spanning five domains —
|
||||
tasks (7), agents (13), knowledge (8), meetings (7), scheduling (2). Its original purpose
|
||||
is 7 of 45. Because agent identity lives in noxflow's database, work that belongs
|
||||
elsewhere must be implemented there: per-agent Claude credentials had to be written by
|
||||
the noxflow runtime, even though node identity — the same kind of fact — lives in the
|
||||
mesh registry.
|
||||
|
||||
**The core domain has no context to live in.** The mesh's distinguishing feature is that
|
||||
a module declares `requires: postgres/database` and never learns where the database
|
||||
lives, who owns the credential, or how it rotates. That is capability brokering, and it
|
||||
is what makes this a mesh rather than four machines with a configuration manager. Yet the
|
||||
logic implementing it sits in `modules/postgres/tools/index.ts` — a provider module,
|
||||
where "the mesh brokers credentials" cannot be expressed. On 2026-08-22 three of its
|
||||
invariants were found violated simultaneously (see Consequences).
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Keep the current layout; fix bugs as they surface.** Rejected. The faults are not
|
||||
independent. Every incident on 2026-08-22 — credentials written by the wrong module, a
|
||||
dependency question answered wrongly twice, an invariant enforced nowhere — traced to
|
||||
a boundary that was never stated. Fixing them individually leaves the generator intact.
|
||||
|
||||
2. **Merge aggressively into few large modules.** Fewer names, same problem: a shared
|
||||
schema across domains is what produced noxflow, and doing it deliberately would
|
||||
produce it again at larger scale.
|
||||
|
||||
3. **Decompose by bounded context, with the mesh as a broker.** Name contexts after their
|
||||
aggregates, integrate through a published record rather than a shared schema, and let
|
||||
deployment granularity be a feature-level concern rather than a reason to create a
|
||||
module. **Adopted.**
|
||||
|
||||
## Decision
|
||||
|
||||
### The domain, in one sentence
|
||||
|
||||
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
|
||||
that think and act.** Everything else supports one of those three.
|
||||
|
||||
### Nodes and agents are decoupled
|
||||
|
||||
A node is a place where an agent can run — that is the entire relationship. There is no
|
||||
resident agent, no node-owned identity, no ownership in either direction. Agents named
|
||||
after a node remain, as **ordinary agents** that happen to hold infra skills.
|
||||
|
||||
Consequently `nodes.node_license` and `nodes.hal_claude_account` cease to exist: a node
|
||||
does not authenticate to a model provider, agents do. The two agent rows per node merge.
|
||||
|
||||
### There is one kind of participant, and some are human
|
||||
|
||||
Human and non-human participants are both **agents**. Both hold identity and credentials;
|
||||
both act, remember and coordinate. What differs is **modality** — how an agent acts:
|
||||
|
||||
| modality | credential is delivered to |
|
||||
|---|---|
|
||||
| spawned session | that agent's own config directory |
|
||||
| shell or desktop | that agent's home on the node it acts from |
|
||||
|
||||
A node holds no licence. **An agent holds credentials, and delivery follows that agent's
|
||||
node bindings and modality.** A human agent's grant arrives in the home directory of the
|
||||
user it acts as, on the nodes it is bound to — the same rule that puts a spawned agent's
|
||||
grant in its config directory, with a different target.
|
||||
|
||||
This requires one fact the mesh does not record today: which user, on which node, a given
|
||||
human agent acts as. Adding it is what removes the node licence — the node is currently
|
||||
standing in for an identity the mesh cannot name. It is also what makes a second human
|
||||
agent require no new mechanism.
|
||||
|
||||
### Contexts
|
||||
|
||||
| context | aggregate | subdomain |
|
||||
|---|---|---|
|
||||
| `hal/mesh` | **Provision**, Node, Module — brokering and its bookkeeping | core |
|
||||
| `hal/agents` | **Agent** — identity, licence, runs, memory, thoughts | core |
|
||||
| `hal/work` | **Task** — workflows, bindings | core |
|
||||
| `hal/stream` | **Thread** — mentions, messages, meetings, notifications | core |
|
||||
| `hal/delivery` | **Pipeline** — jobs, artifacts, features | supporting |
|
||||
| `hal/knowledge` | **Document** — spaces, revisions, review | supporting |
|
||||
| `hal/ai` | **Licence** — provider grants and rotation | supporting |
|
||||
| `hal/observability` | **Check** | supporting |
|
||||
| `hal/config` | **Setting** — env, secrets, PKI | generic |
|
||||
|
||||
A *brain* — memory, thoughts, cognition — is a concept the Agent aggregate owns. It is
|
||||
not a module. Anatomy makes attractive names and poor boundaries; today `hal/brain` names
|
||||
infrastructure and `hal/cortex` describes itself as messaging while running nowhere.
|
||||
|
||||
### Provisioning is the core domain, not plumbing
|
||||
|
||||
`hal/mesh` is a broker; the registry is its bookkeeping. Its invariants are explicit and
|
||||
owned:
|
||||
|
||||
- one rotation source per resource
|
||||
- a credential change fans out to every consumer
|
||||
- a consumer never holds a credential the provider does not know about
|
||||
|
||||
### Contexts integrate through the record, never a shared schema
|
||||
|
||||
`hal/stream` is the published language. A context publishes; it does not join across a
|
||||
boundary. This is what dissolves "meetings" as a domain — a meeting is a thread, and a
|
||||
notification is a mention not yet read.
|
||||
|
||||
### Third-party software leaves the repository
|
||||
|
||||
`plex`, `sonarr`, `postgres`, `verdaccio`, `docker-registry` and 88 others run **on** the
|
||||
mesh; they are not **of** it. The pipeline and provisioning are deliberately
|
||||
module-agnostic, so HAL's own modules dogfood exactly what external modules use — which
|
||||
is what makes the separation safe rather than merely tidy.
|
||||
|
||||
## Consequences
|
||||
|
||||
**The invariants now have an owner, and were measurably unowned before.** On 2026-08-22,
|
||||
`provision_ensure` — documented as "NEVER rotates an existing secret" — was found to mint
|
||||
a new password on every adoption and update only the provider's row. `hal_notifications`
|
||||
consumers on three nodes held dead credentials for two days; `hal_transcripts` had two
|
||||
rows written 216 ms apart, so at most one could match the live role.
|
||||
|
||||
**noxflow dissolves.** `hal/work` inherits tasks and workflows — the concept it was built
|
||||
for. Agents, knowledge, meetings and scheduling return to their contexts. The name goes
|
||||
away.
|
||||
|
||||
**`hal/sdk` shrinks.** 155 files, 34,636 lines, containing code from every context —
|
||||
including `workflow-engine.ts` and `task-commands.ts`, work-domain logic in the kernel
|
||||
every module imports. Each landed there to avoid a cycle between modules that both needed
|
||||
it; a domain module can only own its shared code once the domain has a module. Extraction
|
||||
is therefore downstream of this decision, not independent of it.
|
||||
|
||||
**Two mechanisms are prerequisites, not follow-ups.**
|
||||
|
||||
- *Named features with per-node opt-in.* Without it, every independently deployable unit
|
||||
inside a context becomes a module again and the count returns. `hal/claude-licences`
|
||||
exists solely because one daemon must run on one node.
|
||||
- *A local mesh in containers.* `dev_up` starts providers "via systemctl (same as
|
||||
production)" — it borrows the host, and no mesh can be stood up locally. Everything that
|
||||
manifests **between** nodes is therefore discoverable only in production, which is where
|
||||
every fault of 2026-08-22 was found. A refactor of this size is otherwise unverifiable.
|
||||
|
||||
**Credential delivery becomes uniform.** Per-agent credential directories, built for
|
||||
spawned sessions, extend to human agents unchanged. A file previously scoped to a node
|
||||
becomes scoped to an agent — which would have prevented the class of failure where a
|
||||
rotation reached one node of four while the mesh reported success.
|
||||
|
||||
**Migration is incremental and long.** Contexts can be extracted one at a time behind the
|
||||
existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
|
||||
|
||||
## References
|
||||
|
||||
- [`01-RESEARCH/001-module-domain-decomposition`](../01-RESEARCH/001-module-domain-decomposition/analysis.md)
|
||||
— current-state evidence, table counts, open questions
|
||||
- [`00-META/how-we-build.md`](../00-META/how-we-build.md) — naming and integration rules
|
||||
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
|
||||
array that makes a feature a singleton per module
|
||||
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
|
||||
- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the
|
||||
constraints that make a unit independently shippable, applicable unchanged to features
|
||||
- impire.io / soulstream — *the record* as integration substrate, personas over services,
|
||||
and "cheap awareness and expensive thinking"
|
||||
@@ -0,0 +1,138 @@
|
||||
---
|
||||
status: accepted
|
||||
date: 2026-08-22
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
---
|
||||
|
||||
# 16. A lab node is a virtual machine running the real install
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
|
||||
rather than a convenience: *"everything that manifests between nodes is discoverable only in
|
||||
production, which is where every fault of 2026-08-22 was found."*
|
||||
|
||||
Two things were measured while establishing what exists
|
||||
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
|
||||
|
||||
- **There is no local mesh.** The dev tooling starts providers through the host's own init
|
||||
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
|
||||
It borrows the machine because there is nowhere else to put a mesh.
|
||||
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
|
||||
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
|
||||
|
||||
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
|
||||
physical machine. Every subsequent question — how faithful is faithful enough, what may be
|
||||
mocked, which failures remain reachable — follows from that one answer.
|
||||
|
||||
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
|
||||
hardware virtualisation present.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
|
||||
a node as one inverts the thing being modelled: module service stacks then require nested
|
||||
containers through a privileged daemon, or a shared socket that makes isolation between
|
||||
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
|
||||
and the least like a node.
|
||||
|
||||
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
|
||||
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
|
||||
is the only option that makes a twenty-node run affordable. It was rejected because **the
|
||||
scale requirement that justified it was invented rather than required**: the stated goal is
|
||||
to run the real mesh, which is four nodes, on one computer. And a system container still
|
||||
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
|
||||
then has to be answered again for every capability under test.
|
||||
|
||||
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
|
||||
no storage pools, no snapshot management, no network management, no virtual machines.
|
||||
Snapshots are what make the loop fast, so the saving is not worth what it costs.
|
||||
|
||||
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
|
||||
turns into a node.
|
||||
|
||||
## Decision
|
||||
|
||||
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
|
||||
runs the real install, and becomes a node. It is not a model of a node, so no question arises
|
||||
about how good the model is.
|
||||
|
||||
The environment is called **the lab**.
|
||||
|
||||
Three things follow directly and are decided here:
|
||||
|
||||
### The lab is driven by `incus`
|
||||
|
||||
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
|
||||
bridges between them, through one interface. It also manages system containers, so if a run
|
||||
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
|
||||
|
||||
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
|
||||
does.
|
||||
|
||||
### The simulated public segment uses TEST-NET-3
|
||||
|
||||
`203.0.113.0/24`, reserved by RFC 5737, never routable.
|
||||
|
||||
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
|
||||
peer's underlay address against an RFC1918 regex
|
||||
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
|
||||
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
|
||||
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
|
||||
rather than an addressing mistake.
|
||||
|
||||
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
|
||||
|
||||
### The lab issues its own certificates
|
||||
|
||||
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
|
||||
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
|
||||
single-authority lab would hide any fault living in that split.
|
||||
|
||||
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
|
||||
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
|
||||
failure rather than a mystery.
|
||||
|
||||
## Consequences
|
||||
|
||||
**The fidelity question disappears, and with it a class of argument.** There is no "how real
|
||||
is this node" to litigate per capability, because the node is real. What remains not-real is a
|
||||
short, enumerable list: the model provider, the public internet, and the public certificate
|
||||
authority.
|
||||
|
||||
**The install becomes the thing under test.** A container-shaped lab would have had to skip
|
||||
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
|
||||
|
||||
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
|
||||
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
|
||||
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
|
||||
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
|
||||
|
||||
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
|
||||
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
|
||||
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
|
||||
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
|
||||
remains possible at lower fidelity if one is ever genuinely needed.
|
||||
|
||||
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
|
||||
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
|
||||
time is noise.
|
||||
|
||||
**One change is required before the lab can issue certificates.** The reverse proxy sets no
|
||||
`caServer`, so it defaults to the public authority's *production* endpoint
|
||||
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
|
||||
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
|
||||
than staging means every certificate experiment on a real node consumes issuance quota.
|
||||
|
||||
## References
|
||||
|
||||
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
|
||||
the four host couplings that only obstruct a container-shaped node
|
||||
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
|
||||
being reproduced and the endpoint constraint
|
||||
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
|
||||
is for
|
||||
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
|
||||
recording what it cost to get right
|
||||
- RFC 5737 — reserved documentation address blocks
|
||||
@@ -0,0 +1,93 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-23
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0015-mesh-brokers-nodes-host-agents-think.md
|
||||
---
|
||||
|
||||
# 17. Modules outside the platform core are grouped by domain, not by single function
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules
|
||||
into bounded contexts named after their aggregates, and sends the rest out of the monorepo on
|
||||
the grounds that they run *on* the mesh rather than being *of* it.
|
||||
|
||||
That leaves the larger half unaddressed. Around three quarters of the catalogue are modules
|
||||
that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an
|
||||
SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of
|
||||
media-library services. Today each is its own module, because one module is the unit of *one
|
||||
piece of software*, and no other grouping exists.
|
||||
|
||||
The result is that the catalogue's shape records what was installed, not what anything is for.
|
||||
Four modules that together constitute "how a node is reachable" have no relationship the mesh
|
||||
can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a
|
||||
change to how the mesh handles connectivity has to be made four times.
|
||||
|
||||
This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident
|
||||
rather than by domain* — appearing outside it.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else
|
||||
keeps accident boundaries, so the catalogue becomes harder to read after the refactor than
|
||||
before it.
|
||||
2. **One module per piece of software, with a tag or category field.** Rejected. A label is not
|
||||
a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it
|
||||
drifts from the thing it labels.
|
||||
3. **Group them into domain modules, each owning the software that serves one purpose.**
|
||||
Proposed here.
|
||||
4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for
|
||||
the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it
|
||||
into that model repeats the metaphor-naming mistake ADR 0015 exists to correct.
|
||||
|
||||
## Decision
|
||||
|
||||
*Proposed — the principle is settled; the domain list is not. See "Open" below.*
|
||||
|
||||
Modules that are not part of the platform core are grouped into **domain modules**. A domain
|
||||
is named for the concern it serves, and owns the software that serves it. The unit stops being
|
||||
one piece of software and becomes one purpose.
|
||||
|
||||
This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own
|
||||
domain stand unchanged. This decision covers what ADR 0015 leaves outside them.
|
||||
|
||||
Naming follows the same rule as the core: **name the domain for what it does, not for what it
|
||||
is made of**. Connectivity, not a VPN implementation.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes
|
||||
reach each other is a change to one module.
|
||||
- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from
|
||||
its module list.
|
||||
- Swapping an implementation stops being a module replacement, with the data-volume and
|
||||
provisioning consequences that carries, and becomes a change inside a domain.
|
||||
- The count drops sharply, which is a symptom of the improvement rather than the point of it.
|
||||
- **Grouping conceals.** A domain module hides which implementation is in use, and every
|
||||
operational question — which port, which unit, which credential — gains an indirection.
|
||||
- The migration is not free and has no obvious increments: a domain is only useful once
|
||||
everything belonging to it has moved.
|
||||
- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its
|
||||
own sake would be the same error in the other direction.
|
||||
|
||||
## Open
|
||||
|
||||
**The domain list is not settled and this record does not invent one.** What is decided is the
|
||||
principle; what is not decided is the set. Candidate groupings are visible in the catalogue —
|
||||
connectivity and reachability, node presentation and desktop, media libraries, observation and
|
||||
metrics, storage and data services — but naming them here would be reconstructing a decision
|
||||
that has not been taken.
|
||||
|
||||
Settling the list is a research effort, not an act of this record. Until it concludes, this
|
||||
ADR stays `proposed`.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this
|
||||
extends, and its rule about naming a context after its aggregate.
|
||||
- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are
|
||||
already out of scope here; they are not domains and do not group.
|
||||
- [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md)
|
||||
— the catalogue's current shape, which is the evidence for the problem.
|
||||
@@ -0,0 +1,101 @@
|
||||
---
|
||||
status: proposed
|
||||
date: 2026-08-23
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0011-the-installer-owns-linking.md
|
||||
---
|
||||
|
||||
# 18. The mesh creates no symlinks — a derived file is a copy
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0011](0011-the-installer-owns-linking.md) responded to production data loss — a hand-made
|
||||
link, resolved through a container engine's volume handling, pointing a mount somewhere it
|
||||
should not have — by centralising linking in the installer and forbidding it everywhere else.
|
||||
|
||||
That narrowed the incident class. It did not close it. The hazard is not *who* made the link;
|
||||
it is that a path can resolve somewhere other than where it appears to. A link made by the
|
||||
installer resolves exactly the same way as a link made by hand. The rule made the mechanism
|
||||
rarer and better-governed while leaving the mechanism in place.
|
||||
|
||||
Two things have changed since, and together they remove the argument that kept it.
|
||||
|
||||
**The original case for linking was staleness.** A copy of a service definition goes stale
|
||||
silently while the catalogue moves on, so a link was the cheap way to guarantee the running
|
||||
node reads a current definition. That argument assumes the node's copy is unmanaged.
|
||||
|
||||
**It is not.** [ADR 0004](0004-managed-files-are-generated-never-edited.md) established that
|
||||
everything on a node's disk is derived from the mesh and regenerated when its inputs change,
|
||||
and the installer already **reconciles** links rather than assuming them — repointing stale
|
||||
ones, adopting real files it finds where a link belongs. Reconciling content is the same
|
||||
operation as reconciling a pointer, plus a comparison.
|
||||
|
||||
So the mesh already has the machinery that makes a copy safe, and is using a link to solve a
|
||||
problem that machinery solves better. Worse, a link is conceptually the wrong shape: it makes
|
||||
the node's runtime state a *pointer into source*, which is the one thing
|
||||
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) and ADR 0004 exist to prevent.
|
||||
State is derived onto nodes; it does not reach back.
|
||||
|
||||
## Considered options
|
||||
|
||||
1. **Keep ADR 0011 as the final position** — centralised linking, forbidden elsewhere.
|
||||
Rejected as the status quo. It governs the mechanism rather than removing it, and the
|
||||
failure it was written for remains reachable by any code path the installer trusts.
|
||||
2. **Keep links but harden them** — canonicalise before mounting, refuse a link that escapes
|
||||
an expected root. Rejected: it is a check bolted onto a hazard, and it has to be correct in
|
||||
every consumer, including container engines the mesh does not control.
|
||||
3. **Copy, reconciled by the installer, with staleness detected rather than assumed away.**
|
||||
Proposed here.
|
||||
|
||||
## Decision
|
||||
|
||||
*Proposed — the position is settled; the migration is not designed. See "Open" below.*
|
||||
|
||||
**The mesh creates no symlinks.** A file a node needs is placed on that node as a real file,
|
||||
derived from the mesh and reconciled by the installer like every other managed file
|
||||
([ADR 0004](0004-managed-files-are-generated-never-edited.md)).
|
||||
|
||||
The prohibition in ADR 0011 stands and widens: it ceases to be "only the installer may link"
|
||||
and becomes "nothing links, the installer included".
|
||||
|
||||
When this is accepted, ADR 0011 becomes superseded rather than edited — its reasoning is why
|
||||
the rule exists at all, and the incident behind it is the reason anyone believes either record.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The path-resolution hazard is removed rather than governed. There is no link for a container
|
||||
engine to resolve, so the class of failure that cost production data is closed by
|
||||
construction.
|
||||
- A node's runtime state stops pointing into source. What a node holds is derived output, which
|
||||
is what the mesh's model already says it is everywhere else.
|
||||
- **Staleness becomes a real problem that must be answered, not assumed away.** This is the
|
||||
cost, and it is the whole cost: today a link cannot be stale, and a copy can. The answer has
|
||||
to be detection — the installer comparing what is on disk against what the mesh says should
|
||||
be — and it must be loud, because a silently stale definition is exactly the failure shape
|
||||
this mesh keeps producing ([ADR 0008](0008-a-failed-step-fails-the-job.md)).
|
||||
- Reconciliation gets more expensive: comparing content rather than checking a pointer's
|
||||
target, on every module, on every node.
|
||||
- Disk usage rises, trivially, and is not a consideration.
|
||||
- Existing links must be converted. A node mid-migration holds both forms, so reconciliation
|
||||
has to handle finding a link where a file now belongs — the mirror image of the adoption it
|
||||
already does.
|
||||
|
||||
## Open
|
||||
|
||||
- **How staleness is detected.** Content hash, version marker, or regeneration on every
|
||||
reconcile. This is the decision that makes or breaks the change and it is not taken here.
|
||||
- **Whether anything must keep a link** for reasons outside the mesh's control. If something
|
||||
does, that is a finding worth recording rather than an exception worth granting quietly.
|
||||
- **Migration order.** Converting a node's links is a change to how its services resolve their
|
||||
own definitions, which is not a change to make everywhere at once.
|
||||
|
||||
Until those are answered this record stays `proposed`, and ADR 0011 remains the governing rule.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0011](0011-the-installer-owns-linking.md) — the incident, and the rule this widens.
|
||||
- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — the machinery that makes a
|
||||
copy safe.
|
||||
- [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
|
||||
— what the installer does today, including reconciliation and adoption.
|
||||
@@ -0,0 +1,60 @@
|
||||
# 02-DECISIONS
|
||||
|
||||
Architecture decision records — the "why" trail behind the rules in
|
||||
[`00-META`](../00-META/) and the specifications in [`03-DESIGN`](../03-DESIGN/).
|
||||
|
||||
**Numbered `02` because a decision precedes the design it authorises.** Research concludes,
|
||||
the decision is recorded here, and only then is the design written. Following the folder
|
||||
numbers walks the process in the order it happens.
|
||||
|
||||
One file per decision, numbered, never deleted. A superseded record has its `status:` changed
|
||||
and gains a pointer to what replaced it — **its text is never edited**. The reasoning that was
|
||||
rejected is the expensive half to rediscover.
|
||||
|
||||
The records are a **ledger**: they run in the order the decisions were taken, oldest first.
|
||||
|
||||
## Frontmatter
|
||||
|
||||
```yaml
|
||||
---
|
||||
status: proposed | accepted | superseded
|
||||
date: YYYY-MM-DD # when the decision was taken, not when it was written down
|
||||
deciders: name
|
||||
reconstructed: true|false # true when the record was written after the fact from evidence
|
||||
superseded-by: # 02-DECISIONS/NNNN-....md, when status is superseded
|
||||
extends: # 02-DECISIONS/NNNN-....md, when this record widens an earlier one
|
||||
---
|
||||
```
|
||||
|
||||
## Body
|
||||
|
||||
```
|
||||
# N. Title in plain language
|
||||
|
||||
## Context what was true, with evidence
|
||||
## Considered Options numbered, each with why it was rejected
|
||||
## Decision what was decided
|
||||
## Consequences what follows, including what got harder
|
||||
## References commits, pull requests, knowledge-base entries, prior art
|
||||
```
|
||||
|
||||
State evidence, not assertion. *"Zero of 124 modules declare `brain` as a dependency"*
|
||||
outranks *"the dependency rule is not followed"*.
|
||||
|
||||
## Reconstructed records
|
||||
|
||||
Records 0001–0014 were written on 2026-08-23, after the decisions they describe. Records 0015 onward were taken as records. Those
|
||||
decisions were taken in implementation rather than in a document; the records state what was
|
||||
decided and the evidence it was decided from, and each carries `reconstructed: true` and says
|
||||
so in its first lines.
|
||||
|
||||
A reconstructed record is not a transcript. Where the deliberation is not recoverable, the
|
||||
options section states what the alternatives were and why the chosen one won on the evidence
|
||||
available — not a discussion that did not happen. Where a date is not establishable it says so
|
||||
rather than guessing.
|
||||
|
||||
## Index
|
||||
|
||||
The index is **generated, not maintained** — run the `hal-status` skill, which reads the
|
||||
frontmatter of every record. A hand-written index drifts from the folder it describes, and
|
||||
this one had already done so after a single addition.
|
||||
Reference in New Issue
Block a user