Merge branch 'issue/021-provider-port-published-on-loopback' into design/bootstrap-is-a-pivot

# Conflicts:
#	03-DESIGN/01-to-be/04-lab-installation.md
This commit is contained in:
2026-09-11 00:45:59 +02:00
188 changed files with 14552 additions and 2968 deletions
+3 -3
View File
@@ -34,7 +34,7 @@ comparing it against the code rather than by anyone noticing:
- It described the pipeline as having a separate builder process and a build stage that
packages. Neither was true after 2026-08-04; the documents stayed stale until 2026-08-06
([ADR 0014](../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)).
([ADR 0010](../02-DECISIONS/0010-delivery.md)).
- It listed the mesh as spanning a fixed number of named machines, which is exactly the
content this repository cannot carry.
@@ -44,10 +44,10 @@ symlinks at all — the rule is not merely "only the installer may link", and a
elevating linking to a principle points the opposite way from where this is going.
What exists today is that the installer owns and reconciles every link
([ADR 0011](../02-DECISIONS/0011-the-installer-owns-linking.md)) — an as-is fact, recorded in
([ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md)) — an as-is fact, recorded in
[`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md).
Centralising who may link narrowed the incident class; it did not close it. The intent is to
remove the mechanism, recorded as [ADR 0018](../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md).
remove the mechanism, recorded as [ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md).
A founding document contradicting the direction of travel is precisely the failure this folder
exists to prevent.
+50
View File
@@ -0,0 +1,50 @@
# Checks
```
python3 00-META/checks/records.py structure: links, citations, supersession, topics
python3 00-META/checks/index.py the reading order in 02-DECISIONS/README.md is current
python3 00-META/checks/index.py --write regenerate it
```
Non-zero exit on any problem, so it can be a gate rather than a report.
**Why this exists.** Until now nothing in this repository was verified by anything but reading,
which is how a superseded decision stayed live in the constitution for days and in
`01-to-be/README.md` alongside it. Both were found by a person looking. `how-we-build` §5 says
*an unenforced rule is indistinguishable from a wrong one, and costs more, because people
believe it* — this repository was carrying several.
**Every check here failed on something real before it passed.** A check that has never failed is
indistinguishable from one that cannot.
| Check | Asserts | Found |
|---|---|---|
| `links` | every relative link resolves | — (run ad hoc during authoring; now permanent) |
| `rests-on` | `decisions:` and `extends:` name records that exist and are **accepted** | the class behind both incidents |
| `live-citation` | a governing document citing a **superseded** record names its replacement in the same paragraph | `01-to-be/README.md` citing ADR 0022 as live guidance |
| `supersession` | if A says it was superseded by B, B says it supersedes A | ADR 0012 never declared that it superseded 0011 |
| `numbering` | the number in the filename is the number in the heading | — |
| `topics` | every record names a topic the index knows | — |
| *(index.py)* | the written reading order matches what the records say | — |
| `status-vs-code` | a to-be document naming specific code is not still `designed` | **ten documents**, several with a *What was built* section, describing lab-proven code |
## What is deliberately not checked
- **`02-DECISIONS/` and `01-RESEARCH/` may cite superseded records freely.** A decision record
discusses history; research records what was observed. Flagging those would produce noise on
correct documents, and a check that cries wolf gets suppressed — which costs more than not
having it.
- **`03-DESIGN/00-as-is/` may rest on a superseded record.** It describes what runs, and what
runs was built under whatever was decided at the time
([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md):
*as-is describing a superseded decision is exactly what as-is is for*).
- **Whether a citation's prose is still true.** Only whether the record it points at is live.
A document can cite an accepted record and describe it wrongly, and nothing here notices.
So "governing" means `00-META/` and `03-DESIGN/01-to-be/` — the documents that tell somebody
what to do.
## Adding a check
State what incident it would have caught, and make it fail before you make it pass. A check
whose failure has never been observed is a guess about its own correctness.
+108
View File
@@ -0,0 +1,108 @@
#!/usr/bin/env python3
"""Generate the decision index, and check the written one still matches.
A number identifies a record and never changes, so the folder listing is creation order rather
than reading order. The index is what carries the path — and it is written rather than only
generated on demand, because a reader on a forge sees the folder and not a command.
The objection to a written index is that it drifts. That objection is answered by checking it
rather than by refusing to write one, which is `how-we-build` §5: a rule states how it is
checked.
python3 00-META/checks/index.py --write regenerate it
python3 00-META/checks/index.py fail if it is stale
"""
import glob
import io
import os
import re
import sys
README = "02-DECISIONS/README.md"
START = "<!-- index:start -->"
END = "<!-- index:end -->"
# The reading order. Topics a record may belong to, in the order somebody would learn the system.
TOPICS = [
("the mesh", "What the mesh is"),
("the tiers", "Its tiers, from the bottom up"),
("what runs on it", "What runs on them, and how it gets there"),
("building it", "How it is built"),
("checking it", "How it is checked"),
("how we work", "How we work"),
]
def field(text, name):
m = re.search(r'^%s:\s*(.+)$' % name, text, re.M)
return m.group(1).strip() if m else None
def records():
out = []
for path in sorted(glob.glob('02-DECISIONS/0*.md')):
text = io.open(path, encoding='utf-8').read()
heading = re.search(r'^# \d+\.\s*(.+)$', text, re.M)
out.append({
"file": os.path.basename(path),
"number": os.path.basename(path)[:4],
"title": heading.group(1).strip() if heading else "(no heading)",
"topic": field(text, "topic"),
"status": field(text, "status"),
})
return out
def render(rs):
known = {t for t, _ in TOPICS}
lines = [START, ""]
for topic, label in TOPICS:
rows = [r for r in rs if r["topic"] == topic]
if not rows:
continue
lines.append("### %s" % label)
lines.append("")
for r in rows:
mark = "" if r["status"] == "accepted" else " *(%s)*" % r["status"]
lines.append("- **%s** — [%s](%s)%s" % (r["number"], r["title"], r["file"], mark))
lines.append("")
stray = [r for r in rs if r["topic"] not in known]
if stray:
lines.append("### Unfiled")
lines.append("")
for r in stray:
lines.append("- **%s** — [%s](%s) — `topic:` is %r, which is not one of %s" % (
r["number"], r["title"], r["file"], r["topic"], ", ".join(sorted(known))))
lines.append("")
lines.append(END)
return "\n".join(lines)
def main():
text = io.open(README, encoding='utf-8').read()
wanted = render(records())
if START not in text or END not in text:
print("index: %s has no index markers (%s / %s)" % (README, START, END))
return 1
current = text[text.index(START):text.index(END) + len(END)]
if "--write" in sys.argv:
if current == wanted:
print("index: already current")
return 0
io.open(README, 'w', encoding='utf-8').write(text.replace(current, wanted, 1))
print("index: written")
return 0
if current != wanted:
print("index: %s is stale. Regenerate it:\n"
" python3 00-META/checks/index.py --write" % README)
return 1
print("index: current")
return 0
if __name__ == "__main__":
sys.exit(main())
+333
View File
@@ -0,0 +1,333 @@
#!/usr/bin/env python3
"""Structural checks over HQ's own records.
Every check here exists because the thing it checks for actually happened. See README.md
for which incident is behind which check. Run from the repository root:
python3 00-META/checks/records.py
Exits non-zero if anything fails, so it can be a gate rather than a report.
"""
import os
import re
import sys
ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
# Where a citation is *guidance* rather than history. A document here tells somebody what to
# do, so a link to a superseded record is an instruction to follow a withdrawn decision.
# 02-DECISIONS and 01-RESEARCH are deliberately absent: they record what was decided and what
# was observed, and both legitimately discuss superseded records at length.
GOVERNING = ("00-META/", "03-DESIGN/01-to-be/")
LINK = re.compile(r"\[[^\]]*\]\((?!https?:|mailto:)([^)]+)\)")
ADR_FILE = re.compile(r"^(\d{4})-")
def markdown_files():
for base, dirs, files in os.walk(ROOT):
dirs[:] = [d for d in dirs if d not in (".git", ".claude")]
for name in sorted(files):
if name.endswith(".md"):
yield os.path.join(base, name)
def rel(path):
return os.path.relpath(path, ROOT)
def read(path):
with open(path, encoding="utf-8") as handle:
return handle.read()
def frontmatter(text):
"""Minimal frontmatter reader — enough for the fields these checks use.
Not a YAML parser on purpose: a dependency in a repository that has none, to read four
scalar fields and one list, would cost more than it returns.
"""
if not text.startswith("---\n"):
return {}
end = text.find("\n---", 4)
if end == -1:
return {}
fields, key = {}, None
for line in text[4:end].split("\n"):
item = re.match(r"^\s+-\s+(.*)$", line)
if item and key:
fields.setdefault(key, []).append(item.group(1).strip())
continue
pair = re.match(r"^([A-Za-z_-]+):\s*(.*)$", line)
if pair:
key = pair.group(1)
value = pair.group(2).strip()
if value.startswith("[") and value.endswith("]"):
# inline list: `decisions: [a, b]`, `code: []`
inner = value[1:-1].strip()
fields[key] = [v.strip() for v in inner.split(",") if v.strip()]
elif value:
fields[key] = value
else:
fields[key] = []
return fields
def load_records():
"""Every decision record, by its four-digit number."""
records = {}
folder = os.path.join(ROOT, "02-DECISIONS")
for name in sorted(os.listdir(folder)):
match = ADR_FILE.match(name)
if not name.endswith(".md") or not match:
continue
path = os.path.join(folder, name)
text = read(path)
records[match.group(1)] = {
"number": match.group(1),
"name": name,
"path": path,
"text": text,
"front": frontmatter(text),
}
return records
class Failures:
def __init__(self):
self.items = []
def add(self, check, location, message):
self.items.append((check, location, message))
def report(self):
if not self.items:
print("records: all checks passed")
return 0
by_check = {}
for check, location, message in self.items:
by_check.setdefault(check, []).append((location, message))
for check in sorted(by_check):
print(f"\n{check} — {len(by_check[check])} problem(s)")
for location, message in by_check[check]:
print(f" {location}\n {message}")
print(f"\nrecords: {len(self.items)} problem(s)")
return 1
def check_links(failures):
"""Every relative link resolves to something that exists."""
for path in markdown_files():
folder = os.path.dirname(path)
for number, line in enumerate(read(path).split("\n"), 1):
for target in LINK.findall(line):
target = target.split("#")[0].strip()
if not target:
continue
if not os.path.exists(os.path.normpath(os.path.join(folder, target))):
failures.add("links", f"{rel(path)}:{number}", f"link does not resolve: {target}")
def check_rests_on(failures, records):
"""A document's `decisions:` and a record's `extends:` must name a live record.
These are the load-bearing citations: the document declares that it rests on that
decision. Resting on a withdrawn one is the defect this whole check set exists for.
"""
for path in markdown_files():
front = frontmatter(read(path))
cited = list(front.get("decisions", []) or [])
extends = front.get("extends")
if isinstance(extends, str) and extends:
cited.append(extends)
for entry in cited:
match = ADR_FILE.match(os.path.basename(entry))
if not match:
failures.add("rests-on", rel(path), f"not a decision record: {entry}")
continue
number = match.group(1)
if number not in records:
failures.add("rests-on", rel(path), f"no such record: {entry}")
continue
status = records[number]["front"].get("status")
if status != "accepted":
# A proposed record may extend another proposed one. Decisions are drafted in
# chains -- 0059 extends 0057 while both await review -- and refusing that would
# mean either drafting out of order or marking records accepted to satisfy a
# check, which is the failure this repository already made once.
if frontmatter(read(path)).get("status") == "proposed":
continue
# An as-is document describes what runs, and what runs was built under
# whatever was decided at the time. ADR 0056: "as-is describing a superseded
# decision is exactly what as-is is for."
if rel(path).startswith("03-DESIGN/00-as-is/"):
continue
# An extension that supersedes legitimately names what it replaced.
this = ADR_FILE.match(os.path.basename(path))
supersedes = records[number]["front"].get("superseded-by", "")
if this and supersedes and os.path.basename(path) in str(supersedes):
continue
failures.add(
"rests-on",
rel(path),
f"rests on ADR {number}, which is '{status}' — a document may not rest on a "
f"record that is not accepted",
)
def check_live_citations(failures, records):
"""In a governing document, a link to a superseded record must name its replacement.
The reader of a rule needs to know the rule was withdrawn, and needs somewhere to go.
Naming the superseder in the same paragraph is both, and it is what a person would
write anyway.
"""
superseded = {
number: record["front"].get("superseded-by", "")
for number, record in records.items()
if record["front"].get("status") == "superseded"
}
for path in markdown_files():
if not any(rel(path).startswith(prefix) for prefix in GOVERNING):
continue
text = read(path)
offset = 0
for paragraph in text.split("\n\n"):
line_no = text[:offset].count("\n") + 1
offset += len(paragraph) + 2
targets = LINK.findall(paragraph)
for target in targets:
match = ADR_FILE.match(os.path.basename(target.split("#")[0]))
if not match or match.group(1) not in superseded:
continue
number = match.group(1)
replacement = os.path.basename(str(superseded[number]))
if not replacement:
failures.add(
"live-citation",
f"{rel(path)}:{line_no}",
f"cites superseded ADR {number}, which names no superseder",
)
continue
if not any(replacement in t for t in targets):
failures.add(
"live-citation",
f"{rel(path)}:{line_no}",
f"cites superseded ADR {number} without naming its replacement "
f"({replacement}) in the same paragraph",
)
def check_supersession_symmetry(failures, records):
"""If A says it was superseded by B, B must say it supersedes A."""
for number, record in records.items():
front = record["front"]
status = front.get("status")
by = os.path.basename(str(front.get("superseded-by", "")))
if status == "superseded" and not by:
failures.add("supersession", rel(record["path"]), "marked superseded but names no superseder")
continue
if by and status != "superseded":
failures.add("supersession", rel(record["path"]), f"names a superseder but status is '{status}'")
if not by:
continue
match = ADR_FILE.match(by)
if not match or match.group(1) not in records:
failures.add("supersession", rel(record["path"]), f"superseder does not exist: {by}")
continue
other = records[match.group(1)]
claims = os.path.basename(str(other["front"].get("supersedes", "")))
if claims != record["name"]:
failures.add(
"supersession",
rel(other["path"]),
f"ADR {number} says this supersedes it; this record does not say so "
f"(supersedes: {claims or 'absent'})",
)
def check_topics(failures, records):
"""Every record names a topic the index knows.
The topic is what puts a record in the reading order, so a record without one — or with one
nobody defined — disappears from the index rather than appearing in the wrong place. That is
the quiet failure, so it is the one checked.
"""
known = {"the mesh", "the tiers", "what runs on it", "building it", "checking it",
"how we work"}
for number, record in sorted(records.items()):
topic = record["front"].get("topic")
if not topic:
failures.add("topics", rel(record["path"]),
"no topic, so it has no place in the reading order")
elif topic not in known:
failures.add("topics", rel(record["path"]),
"topic %r is not one of: %s" % (topic, ", ".join(sorted(known))))
def check_numbering(failures, records):
"""The number in the filename is the number in the heading."""
for number, record in records.items():
heading = re.search(r"^# (\d+)\.", record["text"], re.M)
if not heading:
failures.add("numbering", rel(record["path"]), "no '# N. Title' heading")
elif heading.group(1) != number.lstrip("0"):
failures.add(
"numbering",
rel(record["path"]),
f"filename says {number}, heading says {heading.group(1)}",
)
def check_status_against_code(failures):
"""A design document naming specific code may not still call itself `designed`.
**Naming a file is a claim that the file implements this**, so the two fields have to agree.
They drifted: ten to-be documents named working, lab-proven code — several with a *What was
built* or *Raised, and observed* section — while still saying nothing had been built.
Deliberately weak, and that is the point of it being mechanical. It cannot tell whether the
prose is true, only that a document has stopped claiming to be unbuilt once it points at
something. `code: [mesh-control]` — a repository with no path — is a plan and stays
`designed`.
"""
for path in markdown_files():
if not rel(path).startswith("03-DESIGN/01-to-be/") or path.endswith("README.md"):
continue
front = frontmatter(read(path))
if front.get("status") != "designed":
continue
for entry in front.get("code") or []:
named = re.sub(r"\s*\(.*\)$", "", entry).strip().split(None, 1)
if len(named) > 1:
failures.add(
"status-vs-code",
rel(path),
f"`designed`, but names {named[1]!r} in {named[0]}. Naming a file claims "
f"it implements this — use `in-progress`, or `implemented` once it is "
f"defensible from that repository's main branch.",
)
break
def main():
failures = Failures()
records = load_records()
check_links(failures)
check_rests_on(failures, records)
check_live_citations(failures, records)
check_supersession_symmetry(failures, records)
check_numbering(failures, records)
check_topics(failures, records)
check_status_against_code(failures)
print(f"records: {len(records)} decision records checked")
return failures.report()
if __name__ == "__main__":
sys.exit(main())
+18 -16
View File
@@ -3,7 +3,7 @@ status: canonical
updated: 2026-08-23
derives: knowledge-base constitution page
decisions:
- 02-DECISIONS/0009-the-mesh-is-governed-by-a-constitution.md
- 02-DECISIONS/0020-the-mesh-is-governed-by-a-constitution.md
---
# How we build
@@ -37,13 +37,13 @@ incident behind it is not written down, and the fix is to write it down, not to
| Rule | What it means |
|---|---|
| **Never write to a production database directly** | No insert, update, delete or schema statement executed against production by hand. Schema changes go through numbered migrations; data changes go through application code or the module's own capabilities. Raw statements skip every side effect the proper path has — events, audit, cache invalidation, fan-out. |
| **Every schema change is a migration** | Numbered, in the module's own language, compiled with it. Both a baseline for a fresh installation *and* an incremental migration for installations that already exist. If code references a column, the migration creating it must exist. [ADR 0006](../02-DECISIONS/0006-schema-changes-are-numbered-migrations.md) |
| **Every schema change is a migration** | Numbered, in the module's own language, compiled with it. Both a baseline for a fresh installation *and* an incremental migration for installations that already exist. If code references a column, the migration creating it must exist. [ADR 0013](../02-DECISIONS/0013-schema-changes-are-numbered-migrations.md) |
| **Never bypass the pipeline** | No manual database edit, no manual restart as a workaround. Fix the cause and deploy. A workaround that works is a workaround that is never removed, and the next person cannot tell the node from its declaration. |
| **Never create a symlink** | A hand-made link caused production data loss through container volume resolution, and the judgement needed to make a safe exception is exactly the judgement unavailable at the moment it matters. **The mesh creates none at all** ([ADR 0018](../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), which supersedes [ADR 0011](../02-DECISIONS/0011-the-installer-owns-linking.md)). The links the installer still reconciles are a migration, not a permission. |
| **Never create a symlink** | A hand-made link caused production data loss through container volume resolution, and the judgement needed to make a safe exception is exactly the judgement unavailable at the moment it matters. **The mesh creates none at all** ([ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md), which supersedes [ADR 0012](../02-DECISIONS/0012-the-mesh-creates-no-symlinks.md)). The links the installer still reconciles are a migration, not a permission. |
| **Never push directly to the main branch** | Branch, push, review, merge. Every merge is a human checkpoint, without exception — **including in this repository**. A documentation repository is not a lower tier of care; a decision record lands the same way a service does. |
| **One change per pull request, and never merge unapproved work** | Unrelated improvements bundled together cannot be reviewed or reverted separately. And the checkpoint is **a person deciding, not a person clicking** — work may be merged by whoever wrote it once a human has explicitly approved *that merge*, and never on a standing permission, an instruction to do the work, silence, or the author's own judgement that it is ready. [ADR 0042](../02-DECISIONS/0042-approval-is-the-checkpoint.md) |
| **One change per pull request, and never merge unapproved work** | Unrelated improvements bundled together cannot be reviewed or reverted separately. And the checkpoint is **a person deciding, not a person clicking** — work may be merged by whoever wrote it once a human has explicitly approved *that merge*, and never on a standing permission, an instruction to do the work, silence, or the author's own judgement that it is ready. [ADR 0023](../02-DECISIONS/0023-approval-is-the-checkpoint.md) |
| **Never open a pull request unprompted** | A permissions list saying it is allowed is not a request. |
| **A failed step fails the job** | A sequence that continues past a failure does the next thing in the wrong place. Gate each step on the last. [ADR 0008](../02-DECISIONS/0008-a-failed-step-fails-the-job.md), and §5. |
| **A failed step fails the job** | A sequence that continues past a failure does the next thing in the wrong place. Gate each step on the last. [ADR 0010](../02-DECISIONS/0010-delivery.md), and §5. |
### A failed step must stop the steps after it — how it was earned
@@ -69,7 +69,7 @@ reported failure, nothing stopped, and the damage happened somewhere nobody was
- **Every runtime variable is declared.** A variable the module reads and the manifest does not
declare is invisible to the mesh: it will not be generated, injected, or audited.
- **Provisioned credentials arrive through declared requirements**, never hardcoded in code,
compose files or scripts. [ADR 0005](../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)
compose files or scripts. [ADR 0009](../02-DECISIONS/0009-modules-and-the-graph.md)
- **Never install a package by hand.** A package is declared in the manifest and arrives the
way every other package does. A hand-installed package is invisible to the mesh: it is not
declared, not reproduced on the next node, and not present after a rebuild — and the node
@@ -82,7 +82,7 @@ reported failure, nothing stopped, and the damage happened somewhere nobody was
- **Every standalone application gets its own repository**, with a manifest at its root,
registered as a build source. Creating an application directory in the monorepo is a
convention violation and reviewers reject it.
[ADR 0010](../02-DECISIONS/0010-applications-live-in-their-own-repository.md)
[ADR 0015](../02-DECISIONS/0015-applications-live-in-their-own-repository.md)
### Migrations
@@ -96,7 +96,7 @@ reported failure, nothing stopped, and the damage happened somewhere nobody was
surface is regenerated from the mesh database; a local edit survives one synchronisation and is
then silently overwritten, bringing back whatever it fixed. Use the mesh operation that owns
the value. If unsure whether a file is managed, ask the tooling — the answer is not visible
from the file. [ADR 0004](../02-DECISIONS/0004-managed-files-are-generated-never-edited.md)
from the file. [ADR 0011](../02-DECISIONS/0011-managed-files-are-generated-never-edited.md)
---
@@ -114,12 +114,14 @@ and it runs on no node at all.
Anatomy makes attractive names and poor boundaries. Name the thing the domain calls it.
### Group by domain, not by single function
### Things that change together share an authority, not a package
A module is a purpose, not a piece of software. Four modules that together constitute "how a
node is reachable" and cannot be assigned, versioned or replaced as one thing are four
accidents, not four boundaries.
[ADR 0017](../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)
When several modules always change together under one intent, name the **context** that decides
for them. Do not merge them into one module: they are delivered to different nodes, and a module
that must be assigned where half of it is unwanted is not a boundary either.
Coherence is a context. Delivery is a module. Relationships are edges, not folders.
[ADR 0009](../02-DECISIONS/0009-modules-and-the-graph.md)
### Contexts integrate through the record, never through a shared schema
@@ -224,7 +226,7 @@ No drive-by edits. Every change traces to a recorded decision.
a design meeting with at least two node operators — which has never been met and cannot be, as
there is one operator. A rule that cannot be satisfied is not a high standard; it is a rule
everything silently violates. Recorded here as resolved in favour of what is achievable, and
what has in fact been practised ([ADR 0040](../02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md)).
what has in fact been practised ([ADR 0022](../02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md)).
---
@@ -241,7 +243,7 @@ process. Absence of an override means these rules apply unmodified.
## 8. Code quality
*Absorbed 2026-08-26 from the enforced page, which carried these rules while this document did
not — [ADR 0040](../02-DECISIONS/0040-the-constitution-absorbs-what-is-enforced.md).*
not — [ADR 0022](../02-DECISIONS/0022-the-constitution-absorbs-what-is-enforced.md).*
**These rules are recorded because they are enforced, not because this repository earned them.**
Every other rule here states the incident or measurement behind it. These state nothing,
@@ -270,7 +272,7 @@ Data access, business logic and the interface layer are separate.
*Scope: the mesh's services and surfaces. Tier 0 is a statically linked binary that must depend
on nothing installed first, and is written in Go —
[ADR 0041](../02-DECISIONS/0041-the-host-depends-on-nothing.md).*
[ADR 0005](../02-DECISIONS/0005-the-node-host.md).*
- TypeScript throughout; no new untyped JavaScript.
- Strict, with no implicit `any` and no unchecked index access.
+2
View File
@@ -49,6 +49,8 @@ the expensive half.
| [03](03-issues.md) | Issues | Something is wrong — often with the owner unknown |
| [04](04-build-handoff.md) | Build handoff | A design is ready to be built |
| [05](05-constitution-sync.md) | Constitution sync | `how-we-build.md` changed a rule the mesh enforces |
| [06](06-writing-a-module.md) | Writing a module | Something that runs today must run on the mesh |
| [07](07-feature-branches.md) | Feature branches across repos | Work that changes code, in one repo or several at once |
## Status lives in frontmatter
+107
View File
@@ -0,0 +1,107 @@
# Playbook 06 — Writing a module
**Trigger.** Something that runs today must run on the mesh, or a new capability must be
declarable.
**Who runs it.** Whoever is porting or writing it.
*Written 2026-09-01 from doing this for the first time end to end. Every step below exists
because skipping it cost something.*
## Before anything: read what runs
**A module is written from the thing, not from memory of the thing.** For a port, that means its
current compose file, its environment, and where its data actually sits. Assumptions about any of
the three have been wrong every time they were not checked.
Three questions, answered from the machine:
| | why it decides something |
|---|---|
| **what containers, and how do they find each other?** | more than one means a `network`; names between them must match what the software is configured to dial |
| **where is its data?** | a bind mount moves with a path; a named volume does not; an anonymous volume is already losing data on every redeploy |
| **which values are secret, and which are merely settings?** | a secret goes in `own-secrets` or a grant; a setting goes in the manifest and may be overridden per node |
## The steps
1. **Name what it provides and requires**, if anything. A name is what a consumer is coupled to,
not the role it plays ([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)):
`postgres-database`, not `database`. Most modules provide nothing and require nothing — an
application is usually a leaf.
2. **Declare capabilities, not dependencies, for facts about the machine.** `container-runtime`,
`package-manager`, `seat`. A capability is detected and refused against; it is not something a
module can install.
3. **Write the resources in the order they must happen.** They are applied in the order written
and orphans are removed in reverse, so a `network` is written before the containers that join
it and removed after them.
4. **Put every secret in a file, never in `env`.** A declaration travels over the broker in plain
text: a password in `env` is a password the broker sees. **The mesh delivers parts; a module
that needs them combined combines them.**
**A sealed file holds the password and nothing else** — no key, no `=`, no newline that means
anything. So `env-file` must never point at one. It points at a file the module *declares*,
whose content leaves a hole:
```
own-secrets superuser → /var/lib/postgres/superuser.secret the password, alone
a file /var/lib/postgres/superuser.env, mode 0600,
content: POSTGRES_PASSWORD=${secret:superuser}
the container env-file: [/var/lib/postgres/superuser.env]
```
The host fills the hole on the machine, which is the only place both halves exist — the mesh
discarded the value
([credentials and their rotation](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)).
A **provisioner** is the exception: it reads a password file, so it mounts the `.secret`
directly.
Every example module in `mesh-control` had this wrong and shipped: `own-secrets` pointing at a
path *named* `.env`, mounted as `env-file`, holding a bare password. Docker reads that as a
malformed line and the container starts **with no password set at all** — not a failure to
start, a service running on the wrong credential. They parsed and they resolved. Two tests in
`examples/modules` now refuse both halves of it.
Add `restart-on` naming the env file, or the container keeps the credential it started with
through every rotation.
5. **Pin every image by digest.** A tag moves. The manifest in a repository names artifacts; the
manifest the mesh holds names digests, and they are not the same document.
6. **Decide generate or accept.** A new module's credential is generated. **An adopted one keeps
the credential it already has** — `secret accept` — because minting a new password for a
database that already exists locks the application out of its own data.
7. **Add a provisioner only if the software cannot read a file.** A proxy that watches a
directory needs nothing. PostgreSQL needs `CREATE ROLE`, an object store needs a bucket and a
policy, an identity provider needs a realm and a client — those need a small program beside
them. It reads what the mesh granted and reconciles; it does not decide anything.
8. **Prove it in the lab, against the real software.** Not that a container started — that the
thing works: the credential authenticates, a wrong one is refused, the containers reach each
other, the data survives a restart.
## What the first port actually cost
Six attempts, one real bug. Recorded because the ratio is the lesson: **the mesh was right every
time and the scaffolding was not.**
- A shape existed in the language and no host implemented it, so every declaration carrying one
was refused whole — correctly, and the host said exactly that. **Nobody was reading the host's
log.** Read it first; it is the only place that says why a machine did nothing.
- A blind find-and-replace renamed a provision in quotes and missed the same word bare.
- A command was tested only for the invocations that should fail, so it rejected every real one
and the suite stayed green.
- A test asserted on a helper rather than on the code that calls it, three separate times. **A
test that cannot fail when the behaviour is deleted is not defending the behaviour.**
## Rules
- **Read the host's log before theorising.** A declaration that was sent and not applied says so
there and nowhere else.
- **A failing test is kept, not skipped.** It is the reproduction.
- **Never rotate during an adoption.** Rotation is a separate act, afterwards, deliberately.
- **A data directory is never removed by the mesh** ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)),
and that protects against the mesh only — not against a disk or a mistaken command.
+57
View File
@@ -0,0 +1,57 @@
# Playbook 07 — Feature branches across repos
**Trigger.** Work that changes code — in one code repo or in several at once (`mesh-sdk`,
`mesh-control`, `mesh-catalog`, `mesh-host`, `mesh-lab`, and `hq` when a decision rides along).
**Who runs it.** Anyone who writes code, engineers and agents alike. Agents follow it exactly —
it is the guard against the failure it was written for.
## The failure it prevents
A feature was worked as a branch-and-MR per *unit of thought* — one per decision, one per
stacked increment — and each MR was treated as finished when it was *opened*, not when it was
*merged*. Across repos the same feature took a different branch name in each. The MRs piled up
unmerged: one session left **sixteen** stacked intermediate MRs that had to be consolidated and
closed by hand. An MR is a review checkpoint, not a scratchpad.
## The rule
One feature is **one branch name**, **one worktree per repo**, **one MR per repo**, opened
**once, at the end**.
1. **Name the feature once.** `feat/<slug>`. The *same* branch name in every repo the feature
touches — never a different name per repo, never a fresh branch per increment within the
feature.
2. **Isolate each repo.** One git worktree per touched repo under `.work/<slug>/<repo>`, branched
off `main`:
```
git worktree add .work/<slug>/<repo> -b feat/<slug> origin/main
```
Parallel features never collide, and no shared checkout is edited.
3. **Commit as you go — locally.** Increments land on the one branch. Nothing is pushed and no
MR is opened mid-feature.
4. **Finish, then publish.** When the whole feature is done — every repo, tests green — push
every branch and open **one MR per touched repo**, together.
5. **Merge promptly, once approved.** Every merge into `main` is notified and approved
([ADR 0023](../../02-DECISIONS/0023-approval-is-the-checkpoint.md)); once it is, merge —
do not leave it sitting. The branch is deleted on merge.
6. **Leave nothing behind.** After the MRs merge, no `feat/<slug>` branch and no `.work/<slug>`
worktree survive.
## What this is not
- **Not a licence to batch unbounded work.** A feature is a *bounded* unit; if it sprawls for
days, end-of-feature bloat merely replaces per-increment bloat. Split it into features, each
its own branch and MR.
- **Not a second trunk.** Every repo branches off `main`. There is no longer an `initialization`
trunk.
## How it is checked
The end state is visible, and its absence is the smell:
- After a feature merges, `git branch -r | grep feat/<slug>` and `git worktree list` return
nothing for it. A surviving branch or worktree means step 6 was skipped.
- More than one open MR in a repo that share no feature name, or a stack of MRs none of which is
merged, is the failure this playbook exists to prevent — stop and consolidate before opening
more.
+10 -10
View File
@@ -15,26 +15,26 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| Repository | Owns |
|---|---|
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0028](../02-DECISIONS/0028-hq-is-company-scoped.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
## What the mesh becomes
[ADR 0030](../02-DECISIONS/0030-the-repository-structure.md) records the repositories the
monorepo decomposes into. **`mesh-lab` and `mesh-host` exist so far** — the lab is built first
([ADR 0029](../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)); the rest are the
[ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the
monorepo decomposes into. **`mesh-lab`, `mesh-host` and `mesh-control` exist so far** — the lab is built first
([ADR 0016](../02-DECISIONS/0016-the-lab.md)); the rest are the
target, not the present.
| Repository | Tier | Holds |
|---|---|---|
| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0041](../02-DECISIONS/0041-the-host-depends-on-nothing.md)) |
| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0005](../02-DECISIONS/0005-the-node-host.md)) |
| `mesh-substrate` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | the control plane and its contexts |
| `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers |
| `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). |
| `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. |
Tier 4's shape is open, and deliberately so: see ADR 0030 and
Tier 4's shape is open, and deliberately so: see ADR 0019 and
[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md).
## What lives where inside the monorepo
@@ -53,7 +53,7 @@ Named by role, because the layout is itself part of the as-is design — see
## Why applications do not live in the monorepo
A standalone application in the monorepo is a convention violation, and reviewers reject it.
The reasoning is recorded in [`02-DECISIONS/0010`](../02-DECISIONS/0010-applications-live-in-their-own-repository.md):
The reasoning is recorded in [`02-DECISIONS/0010`](../02-DECISIONS/0015-applications-live-in-their-own-repository.md):
the mesh installs, provisions for, and ships an application through exactly the same machinery
whether or not its source sits beside the mesh's own — so co-location buys nothing and costs
the monorepo's review cadence.
@@ -64,7 +64,7 @@ Each module is a standalone package that consumes its dependencies from the priv
not from a sibling directory. The workspace was removed after it caused build-versus-development
divergence — a workspace member importing another resolved to local unbuilt source in the
pipeline and to a published version in development. Recorded in
[`02-DECISIONS/0007`](../02-DECISIONS/0007-no-npm-workspace.md).
[`02-DECISIONS/0007`](../02-DECISIONS/0014-no-npm-workspace.md).
Consequence, and it is a real one: a cross-package change is two steps — publish, then consume
— and a repository-wide `npm install` does not exist.
@@ -2,7 +2,7 @@
status: active
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/02-modules-and-manifests.md, 03-DESIGN/00-as-is/10-module-catalogue.md, 03-DESIGN/01-to-be/00-work-breakdown.md]
became: [02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md]
became: [02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md]
---
# 001 — Module domain decomposition
@@ -54,7 +54,7 @@ Tracked in [`analysis.md`](analysis.md) under "Open questions".
## Deliberately not decided
Recorded so they are not mistaken for oversights. Each is open, and each comes out of
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md); this effort stays
[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md); this effort stays
`active` until they are answered.
| Question | Status |
@@ -64,4 +64,4 @@ Recorded so they are not mistaken for oversights. Each is open, and each comes o
| Catalogue destination — one repository or many. | Open. Phase 4. |
| What the shared library keeps after extraction. | Open. Phase 3. |
| Where human agent modality is recorded — which user, on which node, a human agent acts as. | Open. Required by the model; not yet stored. |
| Which domains the modules outside the platform core group into. | Open, from [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md), which settles the principle and deliberately not the list. |
| Which domains the modules outside the platform core group into. | Open, from [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), which settles the principle and deliberately not the list. |
@@ -237,6 +237,6 @@ returns. That mechanism is the subject of a separate ADR.
pipeline resolves dependencies across the registry rather than the filesystem?
4. **SDK residue** — after extraction, does `hal/sdk` keep transport (`amqp-client`), or
does that belong to `hal/stream`? Everything imports it, which argues both ways.
5. **Human agent modality.** ADR 0015 requires a fact the mesh does not record: which
5. **Human agent modality.** ADR 0001 requires a fact the mesh does not record: which
user, on which node, a human agent acts as. Where does it live — an attribute of the
agent, or of the agent-node binding?
+2 -2
View File
@@ -2,7 +2,7 @@
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/04-delivery.md, 03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md]
became: [03-DESIGN/01-to-be/01-end-to-end-testing.md, 02-DECISIONS/0016-the-lab.md]
---
# 002 — A mesh that runs locally
@@ -24,7 +24,7 @@ This effort establishes what already runs in a container, what is welded to the
what it would take to close the gap. It does **not** choose an approach: the central
question — how a containerised node executes a module service, when a module service is
defined today as a systemd unit shelling to `docker compose` in `/services/` — is not
answered by ADR 0015 and is recorded below rather than decided.
answered by ADR 0001 and is recorded below rather than decided.
## What was established
+1 -1
View File
@@ -253,7 +253,7 @@ over either way.
## References
- [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — the decision this
- [`02-DECISIONS/0001`](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) — the decision this
phase unblocks
- [`03-DESIGN/00-work-breakdown.md`](../../03-DESIGN/01-to-be/00-work-breakdown.md) — Phase 0 tasks
and checkpoint
@@ -1,8 +1,11 @@
---
status: active
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: []
became:
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0005-the-node-host.md
- 03-DESIGN/01-to-be/05-the-node-host.md
---
# 003 — Who supervises a service
@@ -40,14 +43,38 @@ This effort answers the cost half. It does not choose.
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
production then have the same shape rather than a translation layer between them.
- **It cannot be all-or-nothing**, and ADR 0015 already says why: a human agent acts through a
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes **does not
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
Detail and costs in [`analysis.md`](analysis.md).
## Decision needed
## What it became
Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds
anything. The options and their costs are in `analysis.md` under "Options".
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
which is why the effort sat `active` for five days after being answered. Recorded here because
finding that is the point of a sweep.
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): the
host is a plain process on the machine and everything above tier 0 is a container. The substrate
bootstrap declares no service at all — it is package, container, action, container — so the
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
tree*. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts
the launcher there: it supervises the host as a child and shares no code with it, so a host that
cannot start is still recovered.
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
desktop, and those are on the host. So is the host itself — the one thing an init starts.
## What is not closed
**The automatic node rescue the documentation describes does not exist.** No unit declares
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
which never happens, and it outlives this effort — filed as
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
rather than closed with it.
@@ -128,10 +128,10 @@ What it costs, honestly:
---
## 5. Why it cannot be all-or-nothing — and ADR 0015 already says so
## 5. Why it cannot be all-or-nothing — and ADR 0001 already says so
Some of what runs under systemd today **cannot** be containerised, and the reason is
already in the domain model. ADR 0015:
already in the domain model. ADR 0001:
> a non-human agent acts through a spawned session — a human agent acts through a shell or
> desktop
@@ -223,7 +223,7 @@ fate-sharing reason in §3.
## References
- [`002-local-mesh`](../002-local-mesh/analysis.md) — the effort this came out of
- [`02-DECISIONS/0001`](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) — agent modality, which
- [`02-DECISIONS/0001`](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) — agent modality, which
decides what cannot leave the host
- `modules/hal/meshware/daemon/src/cerebellum.ts:815-828` — the self-restart workaround
- `modules/hal/meshware/systemd/hal-module@.service` — the per-module Docker lifecycle
+3 -3
View File
@@ -3,9 +3,9 @@ status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/01-mesh-and-transport.md, 03-DESIGN/01-to-be/01-end-to-end-testing.md]
became:
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0031-the-lab-provides-the-underlay.md
- 02-DECISIONS/0033-a-router-is-scenery-not-a-node.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
- 03-DESIGN/01-to-be/02-scenario-declaration.md
- 03-DESIGN/01-to-be/01-end-to-end-testing.md
---
+32 -6
View File
@@ -1,16 +1,18 @@
---
status: active
status: graduated
initiated: 2026-08-23
touches:
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 03-DESIGN/00-as-is/10-module-catalogue.md
- 03-DESIGN/00-as-is/02-modules-and-manifests.md
became: []
became:
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
---
# 005 — Which domains the catalogue groups into
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) settles
[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) settles
that modules outside the platform core are grouped by domain rather than by single function,
and deliberately does not settle the list. This effort settles the list — and, first, tests
whether the premise survives measurement.
@@ -23,7 +25,7 @@ together**, measured across the full history of the code repository.
## Why
The argument in ADR 0017 is that the catalogue's shape records what was installed rather than
The argument in ADR 0022 is that the catalogue's shape records what was installed rather than
what anything is for — that four modules constituting "how a node is reachable" have no
relationship the mesh can see, so a change to connectivity is made four times.
@@ -51,7 +53,31 @@ premise**, in a way that narrows the effort usefully:
The remaining work is the list itself, for the modules where grouping is justified, plus the
open questions below.
## Open questions
## What it became
*Closed 2026-08-28.* Three of the four questions are answered, and by records that did not cite
this effort — which is why it stayed open after being resolved.
**Whether provider modules group at all** — *no.*
[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md):
there is no `networking` thing to install, there are concrete modules named individually. Folders
assert relationships; edges record them. *Provider* stops being a category at the same time.
**Whether "group or leave" is even the right pair of options** — *it was not*, and that is the
useful finding. [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)
reframes it: things that change together share an **authority**, not a package. This effort's own
measurement is what that record rests on — reachability being the *only* place modules genuinely
co-change is why connectivity is a context and why nothing else needed one.
**What to do with the ~50 modules that co-change with nothing** — *nothing.* They are modules.
Grouping is a tag and a query over the graph, neither of which anybody keeps true by hand.
**What remains is a sequencing question, not a grouping one**, and it moves rather than closes:
*whether applications leave the monorepo before or after they group* is
[research 009](../009-migration/00-overview.md)'s, because it is about how to get from here to
there rather than about what the shape is.
## Open (superseded by the above) questions
| Question | Why it is open |
|---|---|
+4 -4
View File
@@ -10,7 +10,7 @@ updated: 2026-08-23
Every commit in the code repository's main branch that touches the module catalogue, reduced
to the set of modules it touched. Platform-namespace modules are excluded — their
decomposition is settled by
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md). Modules that no
[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md). Modules that no
longer exist are excluded, because pre-rename names dominate the raw signal and describe a
catalogue nobody works in.
@@ -95,14 +95,14 @@ remainder are genuine:
| 2026-08-06 | firewall mesh-only by default, public by declaration |
Each is one intent — *change how a node is reachable* — landing across the proxy, the
resolver, the firewall and the VPN together. That is exactly the shape ADR 0017 describes, and
resolver, the firewall and the VPN together. That is exactly the shape ADR 0022 describes, and
it is the only place in the catalogue where the measurement finds it.
The 2026-08-23 scoping commit is the sharpest case: it spans the reachability cluster **and**
two providers, because "which network is this exposed on" is a reachability question asked of
a database.
## What this means for ADR 0017
## What this means for ADR 0022
The record's principle stands, and its scope needs narrowing. Grouping by domain is:
@@ -130,7 +130,7 @@ Asked directly, and stated as an opinion because it is not yet decided.
is implementation selection, a substantially larger design with its own failure modes, and
nothing currently asks for it.
3. **It would hide which implementation serves a requirement** — the one place the mesh most
needs to be explicit, and precisely the indirection ADR 0017 warns grouping causes.
needs to be explicit, and precisely the indirection ADR 0022 warns grouping causes.
A provider module is already exactly one purpose: it provisions one resource type. That is a
boundary, not an accident of installation.
@@ -2,8 +2,8 @@
status: active
initiated: 2026-08-23
touches:
- 02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
- 02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 03-DESIGN/00-as-is/00-overview.md
- 03-DESIGN/01-to-be/00-work-breakdown.md
became: []
@@ -24,9 +24,9 @@ disk, and where today's catalogue lands.
## Why
Every structural decision so far has been a **correction**: eight contexts replacing thirty-three
modules ([ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)), domains
modules ([ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md)), domains
replacing single-function modules
([ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md)). A
([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)). A
correction inherits the frame of the thing it corrects, and two of the mesh's oldest problems
look unsolvable from inside that frame:
@@ -68,7 +68,7 @@ against taste:
circle, self-hosted. Personal cloud infrastructure.
8. **Agents make it self-improving and self-healing.**
9. It is **end-to-end testable on one machine**
([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Status
@@ -76,7 +76,7 @@ A first skeleton exists, with four design moves that the current shape does not
`active` because two of them are unproven and one contradicts a record that is already
accepted.
**Finding worth stating up front:** [ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md)
**Finding worth stating up front:** [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md)
names nine bounded contexts and **none of them owns connectivity** — no overlay, no resolution,
no firewall, no ingress. Requirement 4 has no home in the accepted decomposition, while
[research 005](../005-domain-grouping/analysis.md) found reachability to be the *only* part of
@@ -87,8 +87,8 @@ the catalogue where modules genuinely change together under one intent. The skel
| Question | Why it is open |
|---|---|
| Does the record — the event log contexts integrate through — belong to the substrate or the control plane? | It is infrastructure by shape and domain by content. Placing it wrong reintroduces a circularity. |
| One repository per tier, or per context? | Already open from ADR 0015 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. |
| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md). |
| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). |
| Four substrate services or five? | The identity provider passes the tier test only if the control plane delegates authentication rather than doing it natively. |
| One repository per tier, or per context? | Already open from ADR 0001 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. |
| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md). |
| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0005](../../02-DECISIONS/0005-the-node-host.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). |
| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. |
| Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. |
@@ -160,8 +160,8 @@ neither option covers, and it is the most common one.
| **Absorbed into the host** | It is not a module at all. It is part of what "managing a machine" means, and belongs in tier 0. | overlay membership, packet filtering, package management, service supervision, container runtime, filesystem management |
| **Substrate** | The control plane cannot exist without it. Pinned, host-applied. | relational store, bus, object store, image registry |
| **Control-plane context** | It decides something across nodes. | connectivity policy, inventory, delivery, provisioning, observability |
| **Workload module** | The mesh hosts it. Grouped per [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md). | media library, desktop session, collaboration tooling |
| **Leaves the repository** | A standalone application, per [ADR 0010](../../02-DECISIONS/0010-applications-live-in-their-own-repository.md). | the applications identified in research 005 |
| **Workload module** | The mesh hosts it. Grouped per [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md). | media library, desktop session, collaboration tooling |
| **Leaves the repository** | A standalone application, per [ADR 0015](../../02-DECISIONS/0015-applications-live-in-their-own-repository.md). | the applications identified in research 005 |
**The first fate is the finding.** Research 005 measured the reachability cluster — proxy,
resolver, firewall, overlay — as the only place in the catalogue where modules genuinely change
@@ -256,7 +256,7 @@ mesh-surfaces/ TIER 3
mesh-catalog/ TIER 4
<domain>/<module>/ layout as above
mesh-lab/ mesh-sdk/ hq/ (company-scoped — ADR 0028)
mesh-lab/ mesh-sdk/ hq/ (company-scoped — ADR 0019)
```
## What this does not settle
+12 -12
View File
@@ -80,12 +80,12 @@ mesh-surfaces/ TIER 3 — thin; no logic lives here
cli/ the shell-facing interface
mesh-catalog/ TIER 4 — what the mesh hosts
<domain>/ grouped per ADR 0017, list per research 005
<domain>/ grouped per ADR 0022, list per research 005
mesh-lab/ the whole mesh, disposable, on one machine
mesh-sdk/ contracts shared across tiers — types, not behaviour
hq/ company-scoped, not a mesh repository — ADR 0028
hq/ company-scoped, not a mesh repository — ADR 0019
```
## The dependency rule
@@ -97,7 +97,7 @@ a second surface would have to reimplement.
This is the whole of the bootstrap answer, and per this repository's own rule it must say how
it is checked: a dependency-direction lint in the build, failing on an upward import. A tier
rule enforced by intention is the same as no tier rule — that is
[ADR 0008](../../02-DECISIONS/0008-a-failed-step-fails-the-job.md) applied to architecture.
[ADR 0010](../../02-DECISIONS/0010-delivery.md) applied to architecture.
## Move 1 — the substrate is applied, not delivered
@@ -144,7 +144,7 @@ assumption that every node is equivalent — already false, and today handled by
## Move 3 — connectivity becomes a context
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) names nine contexts
[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) names nine contexts
and none of them owns the overlay, the resolver, the firewall or the ingress. `config` owns
PKI, which is the closest thing, and it is not close.
@@ -162,8 +162,8 @@ So the evidence and the gap point the same way. `connectivity` owns:
- certificates for both name spaces
This is an addition to an accepted record, so it is a decision, not a drafting choice. It
belongs in a new record that extends ADR 0015 the way
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) does —
belongs in a new record that extends ADR 0001 the way
[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) does —
not written here.
### Where the networking actually lives
@@ -195,7 +195,7 @@ The invitation was to check whether the concept survives. It does not, in one pi
Today a **feature** means both *a thing built once* and *a thing selected per node*, and the
delivery pipeline is hard to reason about precisely because those have different cardinality
and one word ([ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md)
and one word ([ADR 0010](../../02-DECISIONS/0010-delivery.md)
is the pipeline half of the same confusion).
Split it:
@@ -230,7 +230,7 @@ the fact that it runs its own development on them is dogfooding, not architectur
## What agents are, structurally
Self-improvement and self-healing are not a tier. Agents are participants
([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)) that hold identity in
([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)) that hold identity in
tier 2, act through tier 3 like any other caller, and run as workloads in tier 4.
This matters for one reason: **an agent must not have a privileged path**. Anything an agent
@@ -241,7 +241,7 @@ no human checkpoint.
## How this is tested
The lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) raises the
The lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) raises the
tree above on one machine: virtual machines as nodes, a real overlay between them, the real
substrate bundle, the real control plane, the real delivery path.
@@ -254,7 +254,7 @@ being the one thing nobody exercises until it breaks.
- Where the record lives. It is infrastructure by shape and domain by content, and putting it
in the substrate risks recreating a circularity in the one place the design just removed one.
- Whether tier 2's contexts are one repository or several. Open from ADR 0015 already.
- Whether tier 2's contexts are one repository or several. Open from ADR 0001 already.
- Whether an `edge` node is in the inventory or merely present — which decides whether "node"
is one concept or two.
- The migration. Nothing here says how today's mesh becomes this, and the skeleton is worth
@@ -266,13 +266,13 @@ The tier-0 binary was first called `mesh-agent`, because "node agent" is the ref
else in the industry. That is wrong here, and wrong in the specific way
[`how-we-build.md`](../../00-META/how-we-build.md) §4 exists to catch: **Agent** is a
first-class concept in this mesh — a participant, some of whom are human, holding identity and
memory ([ADR 0012](../../02-DECISIONS/0012-agents-are-persistent-employees.md)). One document
memory ([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)). One document
carried both meanings.
It is the same failure as the anatomy naming in the current runtime: an evocative domain word
pointing at infrastructure.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) supplies the fix in
[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) supplies the fix in
its own title — *the mesh brokers capabilities; nodes host; agents think.* Three verbs, three
components: the control plane **brokers** (`mesh-control`), the tier-0 binary **hosts**
(`mesh-host`), the participant **thinks** (`agents`, untouched).
@@ -3,7 +3,7 @@ status: active
initiated: 2026-08-23
touches:
- 03-DESIGN/00-as-is/03-provisioning.md
- 02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
became: []
---
@@ -14,7 +14,7 @@ became: []
Provisioning is the mechanism the whole mesh rests on: a module declares what it needs, and the
mesh makes it exist, generates the credential, records the grant, and puts the values where the
module will read them. [ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md)
module will read them. [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)
calls it the mesh's core concern rather than its plumbing.
[Research 006](../006-mesh-from-scratch/code-skeleton.md) then asks it to carry **more**: the
@@ -1,12 +1,14 @@
---
status: active
status: graduated
initiated: 2026-08-23
touches:
- 03-DESIGN/00-as-is/04-delivery.md
- 02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md
- 02-DECISIONS/0013-an-artifact-is-build-output.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0010-delivery.md
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
became: []
became:
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0010-delivery.md
---
# 008 — The coordinator: a change checked in becomes a deployed state
@@ -41,13 +43,59 @@ Research 006 adds a requirement the current design does not have: the coordinato
**before the mesh is self-hosting**, when source and artifacts come from outside, and keep
working across the transition to self-hosted providers.
## The questions
## What it became
*Closed 2026-08-28.* All six questions are answered, by two records, and the second exists
because the first was honest about what it did not fix.
**Does the coordinator dispatch stages, or converge nodes on a declaration?** — *Converge.*
[ADR 0010](../../02-DECISIONS/0010-delivery.md): a pipeline ends when the
declaration is updated, and the host applies it and reads back — so the reporter is the applier.
**Does the three-silo split survive?** — *Yes, with the third redefined.* The cardinality
observation holds; the third silo is not a stage any more.
**How does a change become a pipeline, reliably?** — *It does not become a pipeline at all.*
[ADR 0010](../../02-DECISIONS/0010-delivery.md) applies 0058's
move one level up: the control plane holds what source exists and what has been built, and builds
the difference. **An event makes it fast; nothing makes it necessary.** The failures this effort
catalogued — a truncated commit list, a broken path match — become latency rather than silence.
**What is a deployed state?** — *Two comparisons, not an event.* Does every node's reported state
match what is declared, and is what is declared built from current source? A milestone can be
claimed by something that did not check; a comparison cannot.
**What produces a verdict, and what is it about?** — *An artifact, and it gates eligibility.* The
mesh must not converge onto something broken, so an artifact may be declared only once the lab
has judged it fit. Sharper than the question expected: a verdict is a property an artifact has,
not a report about a run.
**How does delivery work before self-hosting?** — *It mostly stops being a question.* A reconciler
reads source and writes artifacts; where those live is a binding, external first and internal
later. The transition looked hard because a pipeline's stages name their targets.
## What this effort was right about
Its first question — *what is a deployed state, and how does the mesh know it is in one* — was
marked "everything follows from this", and everything did. Both records above are answers to it:
0058 makes the applier the reporter, and 0063 makes currency a comparison. The effort put the
load-bearing question first.
## What is NOT closed by this
[ADR 0010](../../02-DECISIONS/0010-delivery.md) names four costs
and one of them is a real risk rather than a trade: **a reconciler that cannot reach its target
retries forever, and without something that notices, the failure is silence** — which is the
fault this effort exists to catalogue, reintroduced in a new place. That belongs to observability
and it is not designed.
## The questions (all answered above)
| Question | Why it matters |
|---|---|
| What is a **deployed state**, and how does the mesh know it is in one? | Everything follows from this. If a stage reports transport, "deployed" is a claim nobody checked. A desired-state model with reconciliation gives a different answer from a job-completion model. |
| Does the coordinator dispatch **stages**, or converge nodes on a **declaration**? | The current model is a state machine over stages. The alternative is that a node is told what should be true and reports what is. The second makes drift visible; the first cannot see it. |
| How does a change **become** a pipeline, reliably? | Detection has failed for reasons unrelated to the change, silently. |
| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)) and to a module carrying its own assertions. |
| What produces a **verdict**, and what is it a verdict about? | Ties to the lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) and to a module carrying its own assertions. |
| How does delivery work **before self-hosting**, and across the transition? | From research 006: source and artifacts start external and are re-bound to internal providers. The coordinator has to be indifferent to which. |
| Does the **three-silo** split survive the artifact/part split? | [ADR 0014](../../02-DECISIONS/0014-build-publish-and-deploy-are-three-silos.md) is cardinality-driven, and research 006 renames the thing the cardinality is about. |
| Does the **three-silo** split survive the artifact/part split? | [ADR 0010](../../02-DECISIONS/0010-delivery.md) is cardinality-driven, and research 006 renames the thing the cardinality is about. |
+3 -3
View File
@@ -4,7 +4,7 @@ initiated: 2026-08-23
touches:
- 01-RESEARCH/006-mesh-from-scratch/code-skeleton.md
- 03-DESIGN/01-to-be/01-end-to-end-testing.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0016-the-lab.md
- 03-DESIGN/00-as-is/00-overview.md
became: []
---
@@ -28,7 +28,7 @@ Recorded because incremental is the reflex answer and it is wrong in this case.
requirements — none of these can half-apply. Running both models at once means the old one's
assumptions keep constraining the new one, which is how a migration becomes permanent.
- **Nothing external depends on it.** No users outside the operator, no service level to hold.
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md)).
- **The lab exists precisely for this** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
A big-bang that has been rehearsed end to end, repeatedly, on identical machines is not the
same risk as one performed for the first time on the real mesh. This is also why the lab is
phase 0 rather than a verification step later: the new mesh is *developed* inside it, so by
@@ -70,7 +70,7 @@ than a discovery.
| Phase | What | Done when |
|---|---|---|
| **0** | **Build the lab's bootstrap scenario** ([ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| **0** | **Build the lab's bootstrap scenario** ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) — virtual machines, a network, a way to place a binary, snapshot and reset. No forge, no coordinator, no pipeline. | A machine can be raised from nothing, reset, and raised again, repeatably. |
| A | Build tier 0, **inside the lab**. The host's interface first — it carries the skeleton's biggest unproven claim. | A bare machine becomes a managed node with no mesh present. |
| B | Build tier 1 and 2. The bootstrap scenario grows into the full one by addition — the same machines, with more placed inside them. | The lab raises a full mesh from nothing, repeatedly, from pinned external artifacts. |
| C | Enough of tier 3 to operate it. | The mesh can be driven without direct database access. |
@@ -3,8 +3,8 @@ status: active
initiated: 2026-08-24
touches:
- 03-DESIGN/01-to-be/03-scenario-lifecycle.md
- 02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md
- 02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md
- 02-DECISIONS/0016-the-lab.md
- 02-DECISIONS/0016-the-lab.md
became: []
---
@@ -23,7 +23,7 @@ Measured on a workstation, 2026-08-24. Numbers in [`measurements.md`](measuremen
## Why it matters
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) makes the
[ADR 0016](../../02-DECISIONS/0016-the-lab.md) makes the
bootstrap scenario the inner development loop for tiers 0 and 1 — the argument being that
raising a node from nothing stops being the least-exercised path and becomes the most-exercised
one. **That argument is only true if raising and resetting are cheap.** A loop that costs
@@ -13,7 +13,7 @@ image.
| Fact | Value | Consequence |
|---|---|---|
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-a-lab-node-is-a-virtual-machine.md) is not paying an emulation penalty |
| Hardware virtualisation | present | virtual machines run at native speed; the choice in [ADR 0016](../../02-DECISIONS/0016-the-lab.md) is not paying an emulation penalty |
| Storage drivers the daemon offers | **`dir` only** | no copy-on-write, therefore no cheap snapshot |
| Host filesystems | ext4 throughout | nothing copy-on-write to put a pool on |
| btrfs kernel module | **available** | the kernel can do it |
@@ -72,7 +72,7 @@ worst, before any of the mesh's own work begins.
**This is too slow for an inner loop**, and the reason is not the design.
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md) argues that
[ADR 0016](../../02-DECISIONS/0016-the-lab.md) argues that
making the bootstrap path the inner development loop turns the least-exercised code in the
system into the most-exercised. That argument holds only while resetting is cheap. At a minute
and a half a cycle, with occasional multi-minute stalls, the loop is one a person works around
@@ -120,7 +120,7 @@ A four-machine reset-and-rerun cycle, the operation the inner loop repeats most:
| **cycle** | **~90 s, unbounded at worst** | **~15 s, dominated by boot** |
At fifteen seconds, dominated by a boot that cannot be avoided, the inner loop is viable and
[ADR 0029](../../02-DECISIONS/0029-the-labs-first-scenario-has-no-pipeline.md)'s argument holds.
[ADR 0016](../../02-DECISIONS/0016-the-lab.md)'s argument holds.
At ninety it did not.
### One honest counter-observation
+101 -15
View File
@@ -1,10 +1,15 @@
---
status: active
status: graduated
initiated: 2026-08-25
became:
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0008-a-context-owns-its-store.md
- 03-DESIGN/01-to-be/06-the-control-plane.md
- 03-DESIGN/01-to-be/07-the-substrate.md
touches:
- 02-DECISIONS/0002-everything-is-a-module.md
- 02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
- 02-DECISIONS/0036-a-node-is-a-managed-machine.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
- 03-DESIGN/00-as-is/02-modules-and-manifests.md
- 03-DESIGN/00-as-is/10-module-catalogue.md
- 04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md
@@ -12,15 +17,69 @@ touches:
# 011 — The module graph
> **A third edge was found after this graduated.** This effort established *presence* and
> *instantiation*, and both are **runtime** edges — they answer *what does this need in order to
> run*. Delivery needs a different question answered — *what has to be rebuilt when this changes*
> — and that is a **build** edge, fixed inside an artifact rather than negotiated when it runs.
> Recorded by [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md), which also
> notes what this effort's three entities turn out to be good for
> ([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)).
## What is being investigated
Whether the catalogue's missing structure is a **graph** — modules declaring what they need,
what they offer, and what they exclude — and what that replaces.
[ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) proposes
**[`worked-provider.md`](worked-provider.md) works one module through completely**, and breaks
the tidy version. A database is nine things, not one — and a small game asking the mesh for its
own database shows there are **two kinds of edge**: *presence*, where the thing must exist, and
*instantiation*, where a provider makes something for a consumer and hands back credentials.
Instantiation implies presence and not the reverse. The current system already had this split
and the design had collapsed it.
**[`features.md`](features.md) answers what happens to `feature`.** It is one word for four
things spanning three tiers — artifacts built once per version, resources applied to a machine,
actions run against something that is not this machine, and checks that are requirements in
disguise. Measured: every one of the twenty-one handlers implements all six stages, so `configs`
has a build stage with nothing to build and `npm` has a start stage with nothing to start. That
emptiness is the conflation, and it is why a stage that did nothing and a stage that failed look
alike. Nothing replaces it, because it was never one concept. **One property is worth keeping:
content is detected, relationships are declared.**
**[`cases.md`](cases.md) enumerates what a module can be** — twenty kinds of thing the mesh has
to install, run, own or know about — and extracts the axes a manifest must express. Two of those
axes appear in no current thinking: **how many instances** a thing may have, and **whether two
can coexist**.
**The design is in [`proposal.md`](proposal.md): one kind of edge.** A module provides names
and requires names, and that single relation absorbs requiring a module, requiring a resource,
and the interface-and-adapter idea. An abstract name is legitimate **only where providers are
genuinely substitutable** — `terminal` passes, `database` does not, because a consumer speaking
Postgres does not speak MongoDB. The adapter is what creates an interface; without one there is
a **tag**, which describes and does not bind.
A **node provides names too**, which makes capability checking stop being a separate mechanism
and makes the host's own capability report an input to resolution rather than something a
person reads.
[`analysis.md`](analysis.md) measured the current catalogue. Its value to the design is two
lessons rather than its machinery — a field that means *depends on* should say so, and
placement does not belong in a manifest — and the rest is recorded as as-is evidence.
**Measured, and the premise was wrong for the existing system: the graph is not missing there.**
[`analysis.md`](analysis.md) — 126 manifests, 103 edges, no cycles, nothing dangling, and a
resolver that topologically sorts them, already called by the tool loader, the installer and
the delivery coordinator. What the effort assumed would need building is a thing to call.
What survives is narrower: three declarations that do not exist (`excludes`, a required node
capability, an interface with adapters), and two defects worth fixing whatever else is
concluded — `provider:` is a dependency edge that is not read as one, which makes the closure
for a working mesh come out without a database; and the resolver continues past a cycle and
past a missing dependency, contrary to ADR 0001.
[ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) proposes
grouping modules by domain. [Research 005](../005-domain-grouping/analysis.md) measured that
proposal and found its evidence holds in exactly one place — reachability — which
[ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) has since absorbed
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) has since absorbed
into the host. The measured case for domain grouping has therefore been consumed by a decision
taken for unrelated reasons, and what remains is fifty modules that co-change with nothing.
@@ -43,7 +102,7 @@ and abandoned in favour of one concept with facets, for a reason worth keeping:
Filing decisions that follow from nothing are the disease research 005 measured. A second
taxonomy would reproduce it.
So [ADR 0002](../../02-DECISIONS/0002-everything-is-a-module.md) survives, and the question
So [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) survives, and the question
becomes what a module must be able to **declare**.
## The shape being investigated
@@ -52,7 +111,7 @@ Five declarations, of which two exist today.
| Declaration | Today | Notes |
|---|---|---|
| **requires a resource** — a database, a bucket | yes | provisioning, [ADR 0005](../../02-DECISIONS/0005-capabilities-are-provisioned-on-declaration.md) |
| **requires a resource** — a database, a bucket | yes | provisioning, [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) |
| **provides a resource** | yes | as above |
| **requires another module** | **no** | the dependency edge — the graph's substance |
| **excludes another module** | **no** | installing A makes B unavailable |
@@ -106,12 +165,39 @@ integration being wrong looks like from the outside.
## Open questions
Struck-through rows are answered, with where. The rest are live.
### Settled
| Question | Answer |
|---|---|
| ~~What does the graph **delete**?~~ | For the existing system: nothing, it is already there ([`analysis.md`](analysis.md)). For the design: the module/resource distinction, the interface as a kind of thing, capability checking as a separate mechanism, domain grouping, and — the first clear deletion — **grant kinds**, once a module may only be granted what it exclusively owns ([`worked-provider.md`](worked-provider.md)). |
| ~~Is an interface a module, or a name?~~ | A **name**, and only where providers are genuinely substitutable. The adapter is what creates one; without an adapter there is a **tag**, which describes and does not bind ([`proposal.md`](proposal.md)). |
| ~~Where do domain modules fit?~~ | They do not. There is core infrastructure — concrete modules named individually, not flavourable, nothing standing in front of them. |
| ~~What happens to domain grouping?~~ | Superseded. Folders assert relationships; edges record them. What grouping was for is a tag and a query. [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) is `proposed` and should be superseded rather than narrowed. |
| ~~Is there one kind of edge?~~ | **No — two.** *Presence*, where a thing must exist, and *instantiation*, where a provider makes something for a consumer and hands back credentials. Instantiation implies presence, not the reverse. |
| ~~When two modules provide one name, who chooses?~~ | Neither the consumer naming a node nor the consumer not caring. The consumer declares the **scope of its own need** — shared across its instances, or one each — the mesh binds, and the binding is written down and sticky. Where it is written follows the scope. |
| ~~Can several modules share one database?~~ | **No.** A module is granted only what it exclusively owns — no shared writes and no read role on another's store, because reading couples you to its layout just as firmly. |
| ~~What about a dashboard reading a dozen stores?~~ | **The rule is about contexts, not processes.** The mesh's own board reading the mesh's own store is the mesh showing its own data — not a boundary crossing. Everything inside a context reads its store freely; what is forbidden is a *different* context reading it. An earlier answer here was wrong. |
| ~~Can every registry consumer be served another way?~~ | **Largely dissolves.** Of eighteen direct consumers, the owner keeps its database, node appliers are already stopped by ADR 0005, and the bulk are **foreign tenants** — thirteen tables across three contexts — who need to move out rather than read differently. |
### Live
| Question | Why it is open |
|---|---|
| What does the graph **delete**? | If modules gain declarations and lose nothing, this is motion rather than progress. The effort has not finished until it names what stops existing. |
| Where does resolution happen — mesh or platform package manager? | The mesh must model mesh-level edges. Whether it also resolves operating-system packages, or delegates, decides whether a solver has to be written. |
| Is an interface a module, or a name? | Arch makes it a name that packages claim. Making it a module gives it a manifest, an owner and a place to document the contract — and a thing with no implementation to install. |
| What does an exclusion mean for something already installed? | Refuse the install, or make the conflict visible and let it be decided. The second is a policy surface; the first is a package manager. |
| Does node adoption scan for capabilities, applications, or both? | The operator proposes scanning an adopted node and enabling what it finds. Under the working position above, the scan is for capabilities — but a machine with a terminal already installed is also a module already satisfied, and whether that is adoption or drift is undecided. |
| One installation image, or several? | Proposed: pre-built images carrying different capability sets, so a machine is adopted quickly. Several images bake capability sets at image time, which is the filing problem in a new form and reintroduces what detection exists to avoid. One image carrying the host and nothing else is [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md)'s *one binary installed by hand*, automated. The effort should settle which. |
| What happens to domain grouping? | [ADR 0017](../../02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md) is still `proposed`. If the graph is the answer, 0017 is superseded rather than narrowed — its text is never edited. |
| How many instances should a module have? | Not derivable and not global: one-per-mesh is right for the broker and wrong for a store a disconnected node needs. It is per-module and nothing in the schema says it. |
| Can two of something coexist? | `excludes` covers part of it. Two terminals are fine, two things wanting one port are not, two brokers might be either. |
| What does a provider hand back? | Credentials and an address for a store; a command for a terminal. Same relation, different shape crossing it. |
| Is provisioning one mechanism or two? | The mesh's own registry is provisioned **before there is a mesh**, so provisioning is part of the bootstrap and part of what the carried bundle expresses. At bootstrap the store is local; afterwards it is on another node. Same operation, both sides of a tier boundary. |
| Is a tool surface one relation with two audiences, or two? | 56 of 126 modules carry tools — more than carry a service — and what consumes them is an **agent**, not a module. |
| ~~Do the remaining cross-context reads want an interface or events?~~ | **Derived, not chosen.** Neither is SQL — that only ever runs against your own store. ADR 0004 makes disconnection ordinary, so anything that must work while disconnected cannot use a request and needs a local copy: a subscription. Anything where a stale answer is worse than none cannot use a subscription. |
| What does a consumer do about events it missed while disconnected? | Replay from a point, ask once for a full picture and resume, or rebuild. The question every projection has, and unanswered here. |
| What happens to a grant when its consumer is removed? | Dropping is data loss; keeping is a leak. ADR 0005's removal rule does not obviously carry, because the thing lives inside another module's state. |
| Is a declaration composed per node, from what that node reported? | Some configuration follows the hardware. Either the host fills a blank — deciding, against ADR 0005 — or the control plane composes from the node's inventory first. |
| Would `excludes` and capability requirements actually be used? | Zero manifests declare either, which is equally consistent with *nobody needs them* and *nobody can express them*. |
| What does an exclusion mean for something already installed? | Refuse the install, or surface the conflict and let it be decided. |
| Are tiers a view of the graph, or a constraint on it? | If a tier is a computed level the word is a convenience. If *a tier may depend only on tiers below it* is to be enforced, it is a constraint and must be stated as one. |
| Where does resolution happen — the mesh, or the platform's package manager? | The mesh must model mesh-level edges. Whether it also resolves operating-system packages decides whether a solver gets written. |
| How far may the control plane be split? | A single board over several contexts works because they are contexts *inside* one control plane with one interface. If a context becomes its own deployable with its own interface, the board is coupled to N of them and the composition has nowhere to live that tier 3 permits. A constraint on splitting, worth knowing before splitting. |
| What does the pipeline schedule, once features are gone? | It schedules features today. The four categories they split into have different lifecycles, so the unit of work differs for each and needs naming. |
| Should placement leave the catalogue? | A provision pins itself to a named node in the manifest. Placement is an inventory decision, and having it in the catalogue means a second node cannot provide the mesh's store without editing its consumer. |
@@ -0,0 +1,138 @@
# The graph is not missing
Measured against `origin/main` of the code repository, 2026-08-26. Every manifest, read through
git refs rather than a checkout.
The effort was opened to ask whether the catalogue's missing structure is a graph. It is not
missing. **It exists, it is healthy, and three separate parts of the system already use it.**
That is the finding, and it changes what is worth asking.
## Finding 1 — the graph is already declared, and it is clean
| | |
|---|---|
| manifests | 126 |
| declare a dependency on another module | 64 |
| declare a requirement on a provision | 11 |
| **edges** | **103** |
| declare neither | 62 — 49% |
| **cycles** | **0** |
| **dependencies declared but absent** | **0** |
| deepest chain | 5 |
Half the catalogue is unconnected, which matches
[research 005](../005-domain-grouping/analysis.md)'s finding that fifty modules co-change with
nothing. The connected half is well formed: no cycles, nothing dangling.
## Finding 2 — it is already resolved, and already used
The platform SDK carries a dependency resolver that topologically sorts modules, and it is
called from three places: the tool loader at startup, the installer when syncing modules onto a
node, and the delivery coordinator when expanding what a change affects.
It also already does something the effort assumed would need designing: **a requirement on
another module's provision is treated as an implicit edge to that module**, so a consumer does
not have to declare the same relationship twice.
So *ordering by the graph* — which
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) says
the control plane will do — is not a thing to build. It is a thing to call.
## Finding 3 — the most important edges in the mesh are invisible
The one place the graph is wrong, and it is wrong about the substrate.
A module that needs a database declares it like this:
```yaml
provisions:
- name: mesh-db
provider: postgres # ← a dependency on the postgres module
node: <a named node> # ← and where it must run
```
`provider:` names a module. It is a dependency, declared, in the manifest — and it sits inside
`provisions:`, which is what a module *offers*. The resolver reads `dependencies:` and
`requires:`, so it never sees it.
| | |
|---|---|
| provider references that name a real module | 4 |
| **invisible to the resolver** | **3, across 2 modules** |
Three edges is nothing, and they are the mesh's own database, the mesh's own broker, and the
work engine's database. The most load-bearing relationships in the system are the ones the
graph cannot see.
**The consequence is measurable.** Computing what a working mesh needs, from the declared
graph:
```
registry → sdk → mesh → meshware 4 modules, 4 levels
```
No database. No broker. A closure that is arithmetically correct and obviously wrong, and wrong
for exactly one reason: a field that means *depends on* is not read as one.
This is [04-ISSUES/003](../../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md)
again, in a new place — not a key nothing reads, but a key read as something other than what it
means.
## Finding 4 — the resolver continues past faults it should stop on
Two behaviours, both contrary to
[ADR 0010](../../02-DECISIONS/0010-delivery.md):
- **A cycle warns and falls back to input order.** A cycle means no correct order exists; the
resolver proceeds with an arbitrary one and logs a line.
- **A dependency that does not exist warns and continues.** The validation is documented as
*non-fatal, logged as warnings*.
Neither has fired in the current catalogue — there are no cycles and nothing dangling — which
is why nobody has noticed. They are latent, and they are in the component that
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) makes
responsible for the ordering a host will apply without question.
## Finding 5 — placement is decided in the catalogue
`node:` in a provision pins it to a named node, in the manifest. Two modules do this today,
and they are the substrate ones.
Which node runs what is an inventory and placement decision — tier 2 by the skeleton's own
test. Having it in a manifest means the catalogue decides placement, and a second node cannot
provide the mesh's database without editing the module that consumes it.
## What this means for the effort
**The opening question — "what does the graph delete?" — has an answer: nothing, because the
graph is already there.** The premise was wrong, and finding that out is the effort's first
result rather than a setback.
The questions that survive are narrower and answerable:
| Missing declaration | Manifests using it today |
|---|---|
| `excludes` — installing A makes B unavailable | **0** |
| a required node capability | **0** |
| an interface, with adapters providing it | **0** |
Those are what a graph would *add*. What it would delete is a different and smaller list, and
the honest version of it is: nothing yet.
**And two defects worth fixing regardless of what else this effort concludes:**
1. `provider:` is a dependency edge and is not read as one. Fixing it makes the closure correct
— which is what [research 012](../012-the-minimum-viable-node/00-overview.md) needs in order
to answer what a one-node mesh requires.
2. The resolver continues past a cycle and past a missing dependency. Both should refuse.
## What was not measured
- **Whether the missing declarations would be used.** Zero manifests declare exclusions or
capabilities, and that is equally consistent with *nobody needs them* and *nobody can express
them*. Nothing here separates those.
- **Whether the interface-and-adapter idea has a consumer.** It is a good shape, and it is
argued for rather than measured.
- **What the closure should be.** Finding 3 says the computed one is wrong. It does not say
what the right one is; that needs the fix first.
+128
View File
@@ -0,0 +1,128 @@
# What a module can be
Every kind of thing the mesh has to install, run, own or know about, before deciding what a
manifest says. Written to be argued with: a case here that turns out not to exist should be
struck, and one that is missing is a hole in whatever schema follows.
The hard cases are at the end, and they are the point.
## The ordinary cases
**1 — A supervised service.** A container the mesh runs and keeps running. Nobody starts it;
it is simply up. *A relational store, a message broker, an object store, a dashboard.*
**2 — A system package with configuration.** Not a container. Installed into the machine,
configured through files, run by the service manager. *A firewall, a resolver, an overlay.*
Note: [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) says applying
these is the host's job — so what the module contributes is the *deciding*, not the doing.
**3 — An application a person launches.** Installed on a node, started by a human, running only
while they use it. *An editor, a chat client, a file manager, a terminal.*
**4 — A command-line tool.** Installed, on the path, run when invoked. No service, no window.
*A formatter, a query client, a backup utility.*
**5 — A library.** Never runs at all. Consumed at build time by other modules. *An SDK.*
**6 — A one-shot task.** Runs once, changes something, exits. *A schema migration, a data
import, a seed.*
**7 — A scheduled task.** Runs repeatedly on a timer, exits each time. *A backup, a prune, a
report.*
**8 — An adapter.** Exists to make several unlike things look alike behind one name. *A model
provider behind an assistant interface.*
**9 — A standalone application in its own repository.** Same shape as any of the above; the
difference is only where its source lives
([ADR 0015](../../02-DECISIONS/0015-applications-live-in-their-own-repository.md)). Worth
listing because a schema that assumes a monorepo path would exclude it.
## The cases that break a naive schema
**10 — Something that is a service *and* an application.** A git forge is consumed by other
modules as a remote and a registry, *and* operated by a person through a web interface. A web
analytics service grants a tracking identity *and* is a dashboard somebody reads. Neither is a
service-or-application choice; both are true simultaneously.
**11 — Something that provides to others *and* consumes from others.** The store provides
databases and needs a filesystem. The forge provides a registry and needs a database. Provider
and consumer are not kinds of module; they are ends of edges.
**12 — Something the mesh installs that then becomes a node capability.** The container runtime
is installed *by* the mesh, and once it works, the node **provides** `container-runtime` to
everything else. So a module can change what its node provides. The node's provides-list is
therefore partly derived from what is installed on it, not only detected from what was already
there — and the two have to agree.
**13 — Something that must be adopted rather than installed.** The machine already has the
package manager, the container runtime, possibly the version control system, each with
configuration somebody chose. The module does not install it; it takes it over
([research 012](../012-the-minimum-viable-node/00-overview.md)).
**14 — Something that is a set, not a thing.** *Core infrastructure* is not installable — it is
a name for the concrete modules a working node needs. Whether that is a module whose only
content is `requires`, or a query over the graph, or a pinned list outside the catalogue, is
undecided and is one of the sharper questions here.
**15 — Something with one instance for the whole mesh.** There is one mesh database, not one per
node. Assigning it to two nodes is not redundancy, it is two meshes. Contrast with a terminal,
where per-node is the only sensible reading.
**16 — Something that may be installed several times over.** Two terminals coexist happily. Two
things wanting port 443 do not. Two message brokers might be fine or might be a split brain,
and nothing in *provides* and *requires* distinguishes those.
**17 — Something that is not software at all.** A firewall policy. A DNS record. A certificate.
It owns no binary, runs nothing, and is entirely *desired state* — which is the one case that
fits the host's declaration model exactly and fits an installable-package model not at all.
**18 — An agent.** The mesh's own premise is that agents are participants. An agent has an
identity, a licence, a node it runs on, and work it does. Whether that is a module, a record in
the control plane, or something else is not obvious, and getting it wrong shapes everything
about how agents are assigned.
**19 — The host itself.** Tier 0 installs everything else and is installed by hand. It is not a
module, and a schema that cannot say so has a bootstrap problem hiding in it.
**20 — Something the mesh depends on and does not control.** A domain registrar, an upstream
resolver, a certificate authority, an electricity supply. Almost certainly not modules — but
the mesh's health depends on them, and they are the reason a node can be perfectly configured
and still not work.
## What varies across the cases
The list above matters less than this. These are the axes a manifest has to express, and each
one is a question the schema must answer or deliberately refuse.
| Axis | Range | Sharpest case |
|---|---|---|
| **Does it run?** | supervised · launched by a person · once · on a timer · never | 5, 6, 7, 17 |
| **Who starts it?** | the mesh · a human · nothing | 1 vs 3 |
| **Where does it come from?** | container image · system package · our source · already on the machine | 12, 13 |
| **Does it provide to other modules?** | a resource · an abstract name · nothing | 8, 11 |
| **Does it hold state?** | yes, and it matters where · no | 1 vs 3 |
| **How many instances?** | one per mesh · one per node · many per node | 15, 16 |
| **Can two coexist?** | yes · no · only with different settings | 16 |
| **Is it installed or adopted?** | installed · adopted · either, depending on the machine | 13 |
| **Does installing it change what the node provides?** | yes · no | 12 |
**Two of these are not in any current thinking**, and both come from the hard cases:
*how many instances* (15) and *can two coexist* (16). `excludes` covers part of the second and
nothing covers the first.
## Questions the cases raise
- **Is "runs" a property or a kind?** The axes suggest a property — one schema, with a field
saying how it runs, `never` included. The alternative is several kinds of module with
different schemas, which is the taxonomy [research 011](00-overview.md) already rejected once
for services and applications.
- **Is an agent a module?** (18) If yes, the schema carries identity and licensing. If no, the
mesh has two catalogues.
- **Is core infrastructure a module?** (14) A module whose only content is `requires` is either
elegant or a category pretending to be a thing — the same trap `database` was.
- **What names one-per-mesh?** (15) Nothing in provides, requires or excludes says it, and
getting it wrong means two of something that must be one.
- **Where does a policy live?** (17) Pure desired state fits the host's declaration exactly.
Whether it is a module at all, or something the control plane derives and no catalogue entry
exists for, is open.
@@ -0,0 +1,131 @@
# What a feature is, and what it splits into
Measured against `origin/main` of the code repository, 2026-08-26.
The operator wants the feature concept gone.
[Research 006](../006-mesh-from-scratch/00-overview.md) left it open — *"does `feature`
survive? The skeleton splits it in two and argues the conflation is what makes the delivery
pipeline hard to reason about."*
This is what it actually is, and what it turns into.
## What it is today
A **feature** is a kind of content a module can carry, detected from what its directory
contains rather than declared. Twenty-one of them, each with a handler that owns its whole
lifecycle:
```
configs dist-assets events hooks migrations migration-artifact npm npm-install
prerequisite-env prerequisite-packages prerequisite-provision provision-migrations
provision-seeds seeds service systemd tools verifiers verify vhost detect
```
Each moves through **six stages**: build, publish, install, configure, start, verify.
## Finding — every handler implements every stage
The structural evidence, and it is not a style problem.
| handler | stages it implements |
|---|---|
| `configs` | build · configure · install · start · verify |
| `service` | build · configure · install · start · verify |
| `systemd` | build · configure · install · start · verify |
| `tools` | build · configure · install · start · verify |
| `vhost` | build · configure · install · start · verify |
| `migrations` | build · configure · install · start · verify |
| `npm` | build · publish · install · configure · start |
`configs` writes files onto a node. It has nothing to build, and it has a build stage.
`npm` publishes a package to a registry. It has nothing to start, and it has a start stage.
**One interface spans build-time and apply-time, so every kind of content must implement both
halves and most of them do nothing in one.** That is the conflation, and the emptiness is what
makes the pipeline hard to reason about: a stage that does nothing and a stage that failed to
do anything look identical from outside.
## What it splits into
The twenty-one are not one kind of thing. They are four, and they belong to different tiers.
**Artifacts — built once per version, then published.** `npm`, `dist-assets`,
`migration-artifact`. Nothing about a node is involved; the output is a thing that exists in a
registry. **Tier 2, delivery.**
**Resources — desired state on a machine.** `configs`, `service`, `systemd`, `vhost`, `tools`.
Applied, converged, idempotent — which is exactly what
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md)
already describes and what the host already does. **Tier 0.**
**Actions — run once, against something that is not this machine.** `migrations`, `seeds`,
`provision-migrations`, `provision-seeds`, `hooks`. A migration runs against a database, and the
database may be on another node entirely. Neither an artifact nor node state, which is why they
sit awkwardly in a scheme built for both. **Tier 2, and the operator wants seeds gone.**
**Checks — assertions, not changes.** `prerequisite-env`, `prerequisite-packages`,
`prerequisite-provision`, `verify`, `verifiers`, `detect`. The prerequisites are **requirements
in disguise** — a module saying what must be true before it can be installed, which is precisely
what an edge in the graph says. The verifiers are the read-back the host already performs.
## So the answer is: it splits, and the split is a tier boundary
**`feature` is one word for four things spanning three tiers.** That is why every handler
implements every stage, why half of them are empty, and why the pipeline is hard to reason
about.
Nothing replaces it, because it was never one concept:
| Was a feature | Becomes | Whose |
|---|---|---|
| npm, dist-assets, migration-artifact | an **artifact** | delivery |
| configs, service, systemd, vhost, tools | a **resource** in a declaration | the host |
| migrations, hooks | an **action** against something else | delivery |
| seeds | *nothing* — the operator wants them gone | — |
| prerequisite-* | an **edge** in the graph | the catalogue |
| verify, verifiers | the host's **read-back** | the host |
## Tools are a facet, and their consumer is not a module
The most common content in the catalogue: 56 of 126 modules carry a tool surface, more than
carry a service.
It survives the split, and it does not fit either half cleanly. A tool is not an artifact and
not node state — it is a **contract the mesh publishes on a module's behalf**, and what consumes
it is an **agent**, not another module.
That is a second audience, and the design has only described one. A module `provides` things
other modules require; a module also `provides` things agents call. Same word, different
consumer, different lifecycle — a tool appears when the module is assigned somewhere and
disappears when it is not, and nothing in the graph edges says so.
Whether that is one relation with two audiences or two relations is undecided, and it is the
kind of question that is cheap now and expensive later.
## What is lost, and should not be
**Detection.** Features are detected from the directory rather than declared, and the reason is
good: *a declared list and the directory it describes drift, and the directory is the one that
is true.* That property is worth keeping whatever the concept is called — a module that says it
has migrations and has none, or has them and does not say so, is a fault nobody sees until it
matters.
Under the split, detection still applies: what a module **contains** is read from what is there.
What it **requires**, **provides** and **excludes** is declared, because none of that is visible
in a directory.
That line is worth stating precisely, because it is the one the current design got right:
**content is detected, relationships are declared.**
## Open
- **Is an action a resource?** A migration is not node state and not an artifact. It could be a
resource type the host applies with the target being a database rather than a machine — which
would collapse the third category into the second, at the cost of the host reaching something
that is not the machine it is on. That cost looks too high, but it has not been argued.
- **Where do hooks go?** They are arbitrary code a module runs at a stage — the escape hatch. A
design with no escape hatch is either very good or has not met reality yet, and this one has
not.
- **What does the pipeline schedule, once features are gone?** It currently schedules features.
If the four categories have different lifecycles, the unit of work is different for each, and
what the coordinator orchestrates needs naming.
@@ -0,0 +1,167 @@
# One kind of edge
> **Superseded in part by [`worked-provider.md`](worked-provider.md).** Working postgres through
> completely shows there are **two** kinds of edge, not one: *presence* — the thing exists and is
> reachable, nothing created — and *instantiation* — the provider is asked to make something for
> this consumer and hands back credentials. Instantiation implies presence; presence does not
> imply instantiation. Everything else below stands; the claim in the title does not.
A design, not an account of what exists. [`analysis.md`](analysis.md) measured the current
catalogue and its value here is two lessons rather than its machinery: a field that means
*depends on* should say so, and placement does not belong in a manifest.
## The idea
**A module provides names. A module requires names. That is the only edge.**
Everything the effort listed as separate declarations turns out to be one relation with
different kinds of name on either end.
```
postgres provides postgres
kitty provides kitty, terminal
xterm provides xterm, terminal
anthropic provides anthropic, ai-assistant
meshboard requires postgres, lavinmq
vscode requires terminal, display-server
```
A name is either **concrete** — a module's own name, so `requires: postgres` means that module
and no other — or **abstract**, so `requires: terminal` means whatever provides it.
That single relation absorbs three things the effort had listed separately:
| Was going to be | Is |
|---|---|
| requires another module | requires a concrete name |
| requires a resource | requires the module that provides it, by name |
| an interface, with adapters providing it | an abstract name, legitimate only where an adapter makes providers substitutable |
The interface stops being a kind of module. It is a name with more than one provider *and a
contract they all satisfy*, and nothing has to declare that it is one.
### An abstract name is only legitimate when providers are actually substitutable
The test, and it is a strict one: **can a consumer be switched from one provider to another
without changing?**
`terminal` passes. Anything that runs a command in a terminal works, and a consumer never
learns which one it got.
**`database` fails, and it is the example worth keeping.** A module connecting to Postgres does
not connect to MongoDB, or to SQL Server, or to MariaDB. Different wire protocol, different
dialect, different driver. A consumer that declared `requires: database` and was handed any of
them would break — so the name promises something no provider can deliver, and the resolver
would satisfy a requirement that is not satisfied.
That failure has a shape this repository already knows: something declared, accepted, and not
true. It is worse here than usual because the graph would report success.
**The adapter is what creates an interface.** `ai-assistant` is a legitimate abstract name
exactly when adapters exist to normalise the providers behind it. Without an adapter there is
no interface — there is a category, and a category is not an edge.
So `postgres` is required by name, and a module that could genuinely work with several stores
requires whichever it actually speaks to.
### Categories are catalogue metadata, not structure
*Database* is still a useful word — for finding things, for a person browsing what the mesh can
host, for grouping in an interface. It is a **tag**.
Tags describe. Edges bind. Keeping them apart is what stops the catalogue acquiring a second
kind of relationship that looks like a dependency and is not — which is what a folder named
after a domain already was.
## The node provides names too
The move that makes capabilities stop being a separate system.
A node's profile is a set of provided names. `display-server`. `container-runtime`.
`amd64`. A module requiring `display-server` is satisfied by **the node**, exactly as a module
requiring `postgres` is satisfied by another module.
So there is one resolution rather than two: *is this name provided by anything available here?*
A graphical application cannot be installed on a node with no display server for the same
reason, and through the same code, that it cannot be installed without its libraries.
And it means the host's capability detection — which already reports what a machine can be
asked to do, with a reason for each verdict — **is the node's provides-list**. It was built to
be read by a person; it turns out to be an input to resolution.
## What a module declares
```yaml
name: meshboard
provides: [meshboard]
requires: [postgres, lavinmq, container-runtime] # by name: it speaks their protocols
excludes: []
tags: [observability] # for finding it, never for resolving it
runs: # what applying it means
- container: ...
constraints: # what must be true of a node, never which node
- architecture: amd64
```
`postgres` and `lavinmq` are concrete because this module speaks their wire protocols and would
break against anything else. `container-runtime` is abstract and satisfied by the **node**.
**`excludes`** is the one genuinely new relation: naming something that cannot coexist with
this. It is not derivable from requires and provides, and without it two modules that both
provide `terminal` — or two that both want port 443 — look independent right up until
installing the second breaks the first.
**`constraints` are not placement.** They say what must be true of a node, never which node.
Which node runs what is an inventory decision, and the measurement found the current catalogue
deciding it in the manifest — a module pinning its database to a named node, so that a second
node cannot provide it without editing the module that consumes it. That is the mistake this
separation exists to avoid.
## What it deletes
The question the effort opened with, answered for the design rather than for what exists.
- **The distinction between requiring a module and requiring a resource.** One relation.
- **The interface as a kind of thing.** A name with several providers.
- **Capability checking as a separate mechanism.** The node is a provider.
- **The domain module** — a `networking` module that exists to gather a firewall, a resolver and
a proxy under one name. It came from an older shape and does not fit: there is no such thing
to install. There is **core infrastructure**, which is a set of concrete modules named
individually — a firewall, a store, a resolver — with no flavour and no grouping module
standing in front of them.
- **Domain grouping as structure** ([ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md)).
Folders assert relationships; edges record them. What grouping was for — finding things,
seeing what belongs together — is a **tag** and a *query* over the graph, neither of which
anybody has to keep true by hand.
- **Tiers as a separate concept**, possibly. A tier is a level in the graph, and levels are
computed. Whether the coarse boundary is still worth naming is below.
## What it does not delete, and should not
**A resolver still has to exist**, and it will need version constraints, conflict handling and
a story for when two modules provide one name and nothing says which to use. That is the
long-solved and easily-botched part, and the design should say what it **delegates** to the
platform's package manager rather than reimplementing.
## Open
- **When two modules provide one abstract name, who chooses?** A node with both `kitty` and
`xterm` satisfies `terminal` twice. Whether the choice is a node setting, a mesh setting, or
an explicit pin is the substance of the interface idea, and this proposal does not settle it.
It is a smaller question than it was: the strict substitutability test means the consumer
genuinely does not care which it gets, so the choice is about preference rather than
correctness.
- **Are tiers a view of the graph, or a boundary that survives it?** If tiers are levels, the
concept is derived and the word is a convenience. If the tier rule — *a tier may depend only
on tiers below it* — is meant to be enforceable, it is a constraint on the graph rather than
a description of it, and it has to be stated as one.
- **What does a provider hand back?** A consumer requiring `postgres` needs credentials and an
address; one requiring `terminal` needs a command. The requirement is satisfied by the same
relation in both cases, and what flows across it is not the same shape. Whether that lives in
the name, beside it, or in what the provider returns is undecided.
- **What does a node provide that is not a capability?** Its architecture is a provided name
under this scheme, and so is its operating system. That may be elegant or may be one
abstraction too far; nothing here tests it.
@@ -0,0 +1,491 @@
# A provider, all the way through
One module, worked out completely, because it is the case that breaks the tidy version. It
looks like *a container that runs a database* and it is at least nine things.
Postgres is the example. **The shape is not specific to it** — see the end.
## What it carries
**1 — A supervised container.** An image, a version, a data volume, and configuration. Easy,
and the only part the phrase "a docker service" describes.
**2 — Persistent state, and where it lives matters.** The volume is the database. Moving this
module between nodes is not rescheduling; it is a migration. Almost nothing else in the
catalogue has this property, and nothing in `provides` / `requires` expresses it.
**3 — Configuration that is partly the machine's.** Tuning follows the hardware — memory,
storage. A declaration generated centrally cannot know those, and
[research 012](../012-the-minimum-viable-node/00-overview.md) says the machine's own values win
on conflict. So some of this module's configuration is *derived from the node it lands on*.
**4 — A tool surface.** It exposes capabilities agents can call — query, list, provision. That
is not a resource on a machine and not an artifact; it is a contract the mesh publishes on the
module's behalf.
**5 — A provisioner.** The part that matters, and the one below.
**6 — Its own bookkeeping.** The provisioner must remember what it granted to whom, or it
cannot revoke, rotate or clean up. So a module that provides state to others *also* holds state
about its providing — and that state is not the database's data.
**7 — An exposure decision, per node it runs on.** Reachable from the machine only, from the
local network, or publicly. That is a property of *this assignment*, not of the module: the
same module on two nodes may answer differently.
**8 — Credentials it generates.** Per consumer, and they have to reach the consumer. Which
means a provisioning edge carries a payload, and the payload is a secret.
**9 — Health that is not "the container is up".** A container running and a database accepting
connections are different facts, and the second is the one anything cares about. This is the
host's read-back rule, at a distance.
## The provisioner is a second kind of edge
The tidy version of this effort said: *a module provides names, a module requires names, that is
the only edge.* A small game wanting to store data shows it is not.
```
my-cool-game requires postgres # I speak its protocol, it must exist
my-cool-game requires a database FROM postgres, called my-cool-game
```
The first is **presence**: the thing exists and is reachable. Nothing is created; nothing flows
back. `vscode requires terminal` is this, and so is `requires container-runtime`.
The second is **instantiation**: the provider is asked to make something *for this consumer*,
and hands back what the consumer needs to use it. A database, a user, a password, an address.
They differ in every way that matters:
| | presence | instantiation |
|---|---|---|
| creates something | no | yes, one per consumer |
| carries a payload back | no | credentials, an address |
| can be revoked | — | yes, and must be when the consumer goes |
| provider holds state about it | no | yes — who was granted what |
| satisfied by | anything providing the name | that provider, specifically |
**This is the mesh's actual power**, in the operator's words: a small game declares it wants a
database and the mesh makes one. Nobody creates a user by hand, nobody pastes a connection
string. That is worth being precise about rather than folding into a single relation because
one relation is prettier.
## What that costs the design
**The proposal's "one kind of edge" is wrong**, and the current system already knew: it has
`dependencies` for presence and `requires: provision:` for instantiation, with the resolver
deriving a presence edge from every instantiation edge. [`analysis.md`](analysis.md) recorded
that derivation as a convenience. It is not — it is the correct relationship between two
genuinely different relations.
So: **two kinds of edge, one graph.** Instantiation implies presence. Presence does not imply
instantiation.
## Which provider, when two nodes run one
Asked directly, because two nodes can each run a relational store and a consumer has to be
served by one of them. Neither obvious answer is right.
**Not "the consumer names the node."** That is placement in the consumer's manifest — a small
game edited because a database moved, which is the fault
[`proposal.md`](proposal.md) separates constraints from placement to avoid.
**Not "the consumer does not care" either.** For presence it genuinely does not: a terminal is a
terminal. For instantiation it cares permanently, because the data lands in exactly one store
and the wrong choice is discovered long afterwards.
**What the consumer does know is the scope of its own need.** Not which node — how many of the
thing it wants, relative to itself:
| Scope | Means | Example |
|---|---|---|
| **shared** | one instance serves every instance of this consumer | the mesh's own registry: every node reads the same rows |
| **per instance** | each instance of this consumer gets its own | a local cache, a per-node queue |
That is a property of the consumer, expressible without naming anything. And it is the thing
that actually decides: a shared need cannot be satisfied by a provider each node runs
separately, and a per-instance need should not be satisfied by a shared one.
**Then the mesh binds, and the binding is written down.** Not recomputed: a resolver that
re-derives which store serves a consumer will one day derive a different answer and relocate a
database, so the binding is made once and changed only deliberately.
**Where it is written down follows the scope.** A shared grant belongs to the *module* and every
assignment of it references the same one — which is the answer for one module installed on two
nodes wanting one database between them. A per-instance grant belongs to the *assignment*. Same
relation, two homes, and which home is not a detail: it is what makes two instances share
something or not.
The pieces that follow, none of them settled here:
- **When several providers satisfy the scope**, something chooses — most plausibly locality,
preferring a provider on the same node. That is a default, and it must be overridable, because
the reason to override it is exactly the reason nobody anticipated it.
- **A binding is a thing that can be wrong.** Once recorded it can be inspected, and a consumer
bound to a store on a node that no longer exists is a question somebody can be asked rather
than a failure at connect time.
- **Moving a binding moves data.** Whatever the mechanism, changing it is a migration and not a
configuration change, and a design that lets it look like the latter will lose something.
## Provisioning is early, not late
An assumption worth killing: that provisioning is something the control plane does for
consumers once a mesh is running.
**The mesh's own registry is a provisioned database.** So is its virtual host on the broker.
Neither exists until something creates them, and nothing in the mesh works until they do. The
order is:
```
1 the store runs from the bundle the host carries
2 a database is created in it a provisioning step
3 the mesh's own schema is applied a migration, against that database
4 the control plane starts and only now is there a mesh
5 everything else is provisioned the ordinary path
```
Steps 2 and 3 happen **before there is a mesh to do them**. So provisioning is not a
control-plane service that consumers use; it is part of the bootstrap, and part of what the
carried bundle has to be able to express.
**Which strains what a declaration is.** [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)
has the host applying *declared state on this machine*. A database inside a running store is not
a file or a unit — and at bootstrap it is, at least, local: the store is on the same machine as
the host applying the bundle.
Later it is not. A consumer on one node provisioned from a store on another is the ordinary
case, and reaching it is not the host's job.
**Resolved as two mechanisms, which is the answer rather than a compromise**
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). The host
runs bootstrap actions locally from the bundle; the control plane provisions across the mesh
afterwards. Different actors, different scopes, different trust paths — so there is no single
operation with a tier boundary running through it.
## Several modules, one database — and the case for refusing
Asked, then reconsidered by the operator: *maybe we should not allow it.*
The permissive version was a per-consumer **schema** inside a shared database — its own
namespace, its own migrations, revocable by dropping the schema, with a cross-context join
possible but deliberate.
**The stricter version is better, and it goes further than schemas.**
> **A module is only ever granted a resource it exclusively owns.**
No shared writes. And **no read-only role on another module's database either** — reading
another context's tables couples you to its layout exactly as firmly as writing them does, and
the coupling is harder to see because nothing breaks until the owner changes a column.
That is what [`how-we-build`](../../00-META/how-we-build.md) §4 already says: *contexts integrate
through the record, never through a shared schema.* The permissive version kept the letter of it
and left the temptation in place, and the path of least resistance wins eventually. A boundary
that is merely inconvenient to cross is a boundary that gets crossed.
### What it costs
**Cross-module reporting.** Anything wanting to know what several modules hold can no longer
join across them. It consumes their events, or calls their interface, and neither is as
immediate as a query.
That cost is the point rather than a regrettable side effect — it is §4's whole argument, and
the mesh already has both mechanisms: an event stream every context publishes to, and a tool
surface every module exposes. What gets harder is the thing that was making work belonging to
one context keep having to be implemented in another.
**One more connection per consumer.** A dozen modules means a dozen databases rather than a
dozen schemas in one. For a relational store this is unremarkable; it is worth stating only so
nobody discovers it as a surprise.
### What it deletes
The effort has been looking for what the design *removes* rather than adds, and this is the
first clear instance:
- **Grant kinds.** There is one — an exclusive resource. No schema grants, no read roles, no
scoping rules for who may see what inside a shared thing.
- **The question of who owns which table**, and with it the guessing at revocation time.
- **Cross-module migration ordering.** Two modules migrating one database need their migrations
ordered against each other. Exclusive ownership means a module's migrations are ordered only
against itself.
- **A whole class of permission modelling** that a shared store would otherwise need.
### The rule is about contexts, not processes
An earlier version of this file argued that a dashboard reading a dozen stores was caught by the
rule, because a dashboard is a surface and surfaces speak to an interface. **That was wrong, and
it drew the line in the wrong place.**
The mesh's own board showing nodes, modules and deployments is not a separate context reaching
across a boundary. It is the mesh showing its own data. Reading that store is not a violation
however it is done, and requiring it to go through an interface to reach facts its own context
owns would be ceremony.
**The rule is: a context is granted what it exclusively owns.** Everything inside that context —
its service, its surface, its tools — reads it freely. What is forbidden is a *different*
context reading it.
Which is exactly what the count below shows: the problem was never surfaces. It was three other
contexts keeping their tables in the mesh's database.
**And one surface over several contexts is normal.** The board visualises the mesh, the work
engine, the knowledge base and more, and the alternative — a separate web application per
context — is worse for everyone who uses it. That is not a compromise with the rule; composing
several sources into one view is what a surface *is*.
What it changes is only where it reads from: each context's **interface**, not each context's
**store**. Most of that already exists — 56 of 126 modules carry a tool surface, more than carry
a service.
**And the unified board is what keeps those interfaces honest.** If a view cannot be built from
a context's interface, that interface is inadequate — discovered in the one place where it is
cheap to notice, rather than the first time something else needs the same data and quietly
reaches for the store instead.
**But "the board calls each context's interface" is not quite it either**, and the objection is
right: if every context runs its own service with its own interface, the board is coupled to N
of them instead of N schemas, something has to compose them, and composition is logic — which
tier 3 says a surface does not hold. That moves the problem up a layer rather than solving it.
**The skeleton already answers this and the argument above talked past it.** `work` and
`knowledge` are not separate services; they are **contexts inside the control plane**, alongside
the record, inventory, delivery and the rest — and `api` is listed there as *the one interface
every surface speaks to*.
So the board speaks to **one** interface. Behind it the contexts stay separate: separate stores,
integrating through the record. But they are one tier, one repository, one deployable, and
coupling *within* a tier is not what the tier rule forbids.
Which resolves the objection rather than deflecting it: the problem does move up a layer, and
the layer it moves to already exists and has this as its job.
**The caveat is load-bearing.** This holds only while the contexts are not separate deployables.
The moment one becomes its own service with its own interface, the board is back to N clients,
something must compose them, and the composition has nowhere to live that tier 3 permits. That
is a real constraint on how far the control plane may be split, and it is worth knowing now
rather than discovering it by splitting.
If composing even one interface turns out too slow, the answer is a projection the board owns
and keeps current from events — not access to somebody else's tables.
### Checked against the real consumers
The rule's survival turned on whether every reader of the mesh's own registry could be served
some other way. **Eighteen consumers open a direct connection to it.** Four groups, and only one
of them is work.
| Group | What happens under the rule |
|---|---|
| **The owner and its machinery** — the mesh module, the SDK, the environment and configuration synchronisers, secrets | Nothing. It owns the database. |
| **Node appliers** — the overlay, the shell daemon, the resolver | **Already resolved.** [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) stops the host querying the mesh database, decided for tier reasons with nothing to do with this. |
| **Foreign tenants** — the work engine (10 tables), the knowledge base (2), pipeline logs (1) | They need **their own database**. They are not reading the registry; they are storing their own data in it. |
| **Genuine cross-context reads** — the work engine reads `nodes`; two others read a handful | The only ones needing an interface or events. |
**Thirteen foreign tables live in the mesh's registry database**, belonging to three separate
contexts. That is [`how-we-build`](../../00-META/how-we-build.md) §4's shared schema, counted.
**So the question was the wrong shape.** The bulk of the problem is not readers needing a new
route to data — it is **tenants needing to move out**. Tasks, agents and teams have nothing to do
with nodes and modules; they are co-located by history. Give that context its own database and
its dependency on the registry shrinks to a single table.
What remains is a handful of genuine cross-context reads, small enough to enumerate rather than
estimate. **The rule holds.**
### Request or subscription, and what decides
The remaining cross-context reads need one or the other. **Neither is SQL** — under exclusive
ownership a module runs SQL against its own database and nothing else, whatever transport a
query might travel over. Both options are the mesh's own channel, and both ride the broker
([ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md)), so the transport is
not the distinction.
**The distinction is where the answer lives when you need it.**
| | request | subscription |
|---|---|---|
| you ask | at the moment you need to know | never — you are told |
| the answer lives | on the other side | in your own store |
| freshness | always current | as current as the last event you received |
| when the other side is down | you cannot answer | you answer from your copy |
| what you must handle | a round trip that can fail | events you missed while you were down |
**What decides is not taste.** [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
makes disconnection an ordinary situation rather than an exception. So:
> **Anything that must keep working while disconnected cannot use a request** — there is nobody
> to ask. It needs a local copy, which means a subscription.
And the converse: anything that must be *correct at the instant of asking*, where a stale answer
is worse than no answer, cannot use a subscription. A display can lag. A decision about whether
a grant is still valid cannot.
That turns an apparently open question into a derived one. What remains genuinely open is
narrower: **what a consumer does about the events it missed** while it was disconnected — replay
from a point, ask once for a full picture and resume, or rebuild from scratch. That is the same
question every projection has, and nothing in the record answers it yet.
## A migration belongs to the consumer and runs on the provider
A game defines migrations. They run against the database the store granted **it**. So the
migration is:
- **owned** by the consumer — it is that module's schema, versioned with that module;
- **hosted** by the provider — it runs inside something the consumer does not control;
- **ordered** after the provisioning edge — there is nothing to migrate until the grant exists;
- **scoped** to the grant — the consumer's migrations touch its database and no other.
Ownership crosses the edge, which nothing in *provides* and *requires* expresses. And it is the
*action* category from [`features.md`](features.md) made concrete: not an artifact, not node
state, and not something the host can apply, because the thing it changes is not the machine.
It also gives a consumer's own install an internal order — **provisioned, then migrated, then
started** — that depends on an edge rather than on the module's contents.
## What still has no answer
**How many instances of postgres should exist?** One per mesh is wrong — a node that must work
while disconnected cannot depend on a database elsewhere. One per node is wrong — the mesh's own
registry is one thing, not one per node. So the answer is per-module, and nothing in the schema
says it. This is [`cases.md`](cases.md) axis *how many instances*, and postgres is the case that
proves it cannot be a global rule.
**What happens to a grant when the consumer is removed?** The game is uninstalled. Its database
still exists, holding its data. Dropping it silently is data loss; keeping it forever is a leak.
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) says
the host removes what it applied and no longer declares — but this is not on the host, it is
inside another module's state, and the same reasoning does not obviously carry.
**Where does node-derived configuration come from?** (3) The control plane composes a
declaration, and cannot know this machine's memory. Either the host fills in a blank the
declaration leaves — which makes the host decide something, against
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) — or the control
plane reads the node's inventory first and composes with it. The second is consistent and means
a declaration is composed *per node from what the node reported*, which is a stronger claim than
anything recorded so far.
## Providing is not a substrate thing
The four pinned services are the obvious providers, and they are not a category.
| Service | What a consumer asks it for |
|---|---|
| a relational store | a database, a user, credentials |
| another relational store, different vendor | a database — **and not the same one** |
| a message broker | a virtual host, a user, permissions |
| an object store | a bucket and keys |
| an image registry | a repository |
| an identity provider | a client, a realm, a secret |
| an analytics service | a site, and a tracking identity |
| a low-code data platform | a base, and a token |
| an application platform | a project, which is several of the above at once |
| a mail server | a mailbox, an alias, credentials |
**Any hosted service can be a factory.** Providing is a facet a module may have, not a kind of
module it is — which is the same conclusion the effort reached about services and applications,
arriving from the other direction.
That kills the last reason to keep *provider* as a category. A module runs something, or grants
something, or both, or neither.
### Two stores, and why `database` still is not a name
Two relational stores from different vendors both grant *a database*. They are the sharpest
possible test of the substitutability rule from [`proposal.md`](proposal.md), and they fail it
completely: different wire protocol, different dialect, different driver, different client
library compiled into the consumer.
A consumer declaring `requires: database` and being handed either would break against one of
them. So the name promises what no provider delivers — and now with two real providers in the
catalogue rather than a thought experiment.
*Database* remains a **tag**. It is how a person finds both. It is not how a consumer names what
it needs.
## The same shape, three more times
The message broker has all nine. So does the object store, and so does the image registry. They
differ in what a consumer asks for — a database, a virtual host, a bucket, a repository — and in
nothing structural.
**A substrate service is a service plus a factory.** That is the whole pattern, and there are
four of them. It generalises past the substrate too: anything that grants something per consumer
has this shape, and anything that does not is the simpler case.
But two things differ *between* them, and both matter more than the similarity.
### The broker cannot be managed over the broker
[ADR 0002](../../02-DECISIONS/0002-nodes-communicate-over-a-broker.md) makes the broker the
channel every node takes work from, and
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) makes it the security
boundary — everything a node applies arrives through it.
So the module providing the broker is also **the way modules are managed**. A declaration cannot
be delivered to it over itself, and reconfiguring it is done through the thing being
reconfigured. Nothing else in the catalogue has that property; the store is consumed by the
control plane but is not how the control plane *reaches* anything.
This is exactly what the carried bundle exists for
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)): the broker is raised from
what the host carries, before there is a channel, because there is no other way to raise it.
Recorded here because it is a constraint on *one module*, not a general rule, and a schema with
no way to say so hides it.
### Two modules of identical shape want different instance counts
The broker is one per mesh — a single point of failure and a single point of trust, by decision
rather than by accident. The store cannot be: a node that must keep working while disconnected
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) cannot depend on a database
somewhere else.
Same nine properties, opposite answers. Which settles something the cases file left open: **how
many instances is not derivable from what a module is.** It is a decision per module, it has to
be declared, and nothing in `provides`, `requires` or `excludes` says it.
### And revocation differs in consequence
Revoking a database leaves data behind until something drops it — a leak, and recoverable.
Revoking a virtual host drops whatever had not been delivered — not recoverable, and silent.
The relation is the same and the blast radius is not, which is an argument for the provider
deciding what revocation means rather than the mesh applying one rule to all of them.
## The assignment is a third thing
Recorded because the operator tried the alternative and abandoned it: **modules were once
node-agnostic**, and it did not survive contact.
The worked example says why. Several of a provider's nine properties are not properties of the
module at all:
- **where its state lives** — a volume on a particular machine;
- **how it is reached** — the same module on two nodes may answer locally, on the network, or
publicly, and that is a per-assignment decision;
- **configuration derived from the hardware** — tuning follows the memory and storage of the
machine it landed on;
- **whether this instance is the one** a given consumer is provisioned from.
None of those belong in the catalogue, because they differ per node. None belong in the node,
because they are about this module. **They belong to the pairing**, and a design with only
modules and nodes has nowhere to put them — which is what "node-agnostic" ran out of.
So there are three entities, not two:
> **a module** · **a node** · **an assignment**, which is a module on a node and carries its own
> configuration
The current system already has this, arrived at the same way: environment values are stored per
module *and per node*, so a module's settings differ between the machines running it.
**This does not put placement back in the manifest.** A module still says what must be true of a
node and never which node ([`proposal.md`](proposal.md)). What changes is that the *result* of
placing it is a thing with its own state, rather than a fact recorded on one of the two ends.
And it makes the composed declaration question from above answerable: a declaration is built
from the module, the node's inventory, and the assignment between them. Three inputs, which is
why two were never enough.
@@ -0,0 +1,199 @@
---
status: active
initiated: 2026-08-26
touches:
- 02-DECISIONS/0005-the-node-host.md
- 02-DECISIONS/0011-managed-files-are-generated-never-edited.md
- 03-DESIGN/01-to-be/05-the-node-host.md
- 03-DESIGN/00-as-is/05-runtime-and-installation.md
- 01-RESEARCH/011-the-module-graph/00-overview.md
---
# 012 — The minimum viable node, and adopting what is already there
## What is being investigated
Two questions that turn out to be one:
**What is the bare minimum to run a one-node mesh?** Not the tiers as asserted, but the actual
closure — take the thing that must run, walk what it needs, and the set that comes back is the
answer.
**And how does a machine that is already in use become that?** A candidate node is not empty. It
has a package manager, probably a container runtime, possibly a git installation, each with
configuration somebody chose. The mesh must **own** those, and owning is not the same as finding
them present.
## Why
Building tier 0 reached a wall that looked like a packaging problem and is not.
The host can be told to run a container or install a package. Both need a file — an image, an
archive — and the question was where the host gets it. That framing produced a bad trilemma:
carry everything in the bundle, download at apply time, or have something push the files in
first. Downloading fails on the first node, which cannot fetch the image registry from the image
registry it is trying to start.
> **Qualified by [ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md).** The
> reframing below still holds for what a *tailored installer* contains — the missing pieces for a
> given machine. It does **not** have to hold for container images: the installer fetches those
> by digest, because a real machine has a network and the sealed case is the lab.
**The reframing:** the machine is not offline. What matters is *when* the fetching happens. Move
it from apply time to **build time** — build the installer on a machine that has a network,
tailored to the target, and apply it on a target that then needs nothing. That is the same move
the lab already made for its router image, and the same property the delivery design already
claims: what ships is self-contained and a deploy touches no network.
Which makes the interesting question not *where do artifacts come from* but **what is missing
from this particular machine**, and that needs both of the questions above answered.
## What tailoring implies
- **The binary stays generic; the payload is tailored.** One static host per architecture. What
is machine-specific is the bundle it applies. Rebuilding the host per machine would buy
nothing.
- **Detection is the input, not a report.** What the host already reports about a machine —
its profile and inventory — is what the difference is computed against. This is the first use
of stage 1 by something other than a person reading it.
## Adoption
**Simply having a package installed is not enough.** If the mesh manages the container runtime,
it decides that runtime's configuration; a runtime found already installed carries settings
somebody chose, and those cannot be discovered by noticing the binary exists.
So a machine already in use is **adopted**: what is there is read, taken over, and thereafter
generated.
**This was the original path.** [`00-as-is/05`](../../03-DESIGN/00-as-is/05-runtime-and-installation.md)
records adoption of a pre-existing machine's configuration as the original mechanism, since made
legacy and explicitly out of scope for the lab. It returns here for a different reason than it
was dropped for, which is a thing to notice rather than to gloss.
**It creates a state that does not exist today.** [ADR 0011](../../02-DECISIONS/0011-managed-files-are-generated-never-edited.md)
has managed files generated and never edited; adoption needs a one-time import before that rule
starts applying. Three states, and the middle one is new:
> unmanaged → **adopted once** → generated
**And it crosses a boundary just drawn.** [ADR 0005](../../02-DECISIONS/0005-the-node-host.md)
says the host never touches what it did not create — the rule that stops a converger deleting
what the mesh never put there. Adoption is the deliberate act of taking ownership of exactly
that. The rule needs a companion rather than an exception: *never, unless adoption made it the
host's*, with adoption being explicit, recorded, and visible in what the host says it owns.
## Nothing is taken over without keeping what was there
**Before adoption touches a file, the original is kept.** Adoption happens on machines somebody
is already using, and the configuration being taken over is configuration somebody chose. A
one-way door on a working machine is not an installation, it is a risk nobody agreed to.
This is a *never* rule rather than a courtesy, and it earns that by the same incident the mesh's
strongest rule already carries: the worst loss in this record came from a tool acting on a path
it did not own. Adoption is that act, made deliberate — which makes the safeguard obligatory
rather than optional.
What that requires, and what remains open: where the copy lives, whether it is recorded in what
the node knows about itself so that adoption is *visibly* reversible, and whether the mesh keeps
it forever or hands it back when it stops managing the thing.
## Adoption produces a briefing, not just a result
Proposed by the operator, and it answers a question this effort had open with two bad answers.
Adoption meets things a script cannot decide. A container runtime configured with one storage
driver and a mesh wanting another. A package pinned to a version somebody chose for a reason.
Local settings the mesh has no opinion about and no business discarding. Silently winning is
wrong in both directions; refusing outright makes a machine in use unadoptable.
**So adoption has two outputs.** What it did — mechanical, recorded, in the node's state. And a
**briefing**: what it found, what it took over, and what it could not resolve, written to be
read by a person or an agent, which is the first thing a session on that node has to work with.
Conflicts are **flagged, not resolved**. That is the same principle the declaration parser
already applies — name every problem at once, to somebody who can act on it — at a larger
scale, and applied to a case where refusing wholesale would be worse than proceeding.
### On conflict, the machine's configuration is kept
Decided by the operator, after first deciding the opposite — recorded that way because the
reasoning for each direction is the useful part.
Where the existing configuration and the mesh's disagree, **what is already on the machine
stays**, the conflict is flagged, and it is reconciled afterwards. Adoption always **completes**
— flags inform, they do not block — and *adopted with open questions* prevents nothing. The
node is a node.
**What this buys.** Adoption becomes non-destructive by construction. The class of conflict that
made the opposite rule dangerous — a storage driver against the filesystem it is actually on, a
data directory pointing at a mount that exists — cannot arise, because nothing tied to the
machine's physical reality is ever overwritten. A machine in use keeps working exactly as it
did.
**What it exposes, which is the mirror of what it fixes.** The mesh's configuration is not only
preference. Some of it is what a module needs in order to function at all. Keeping the machine's
version there produces a module that is installed and does not work — *an installed package is
not a capability*
([04-ISSUES/007](../../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md))
arriving from a direction that issue did not anticipate. And a fleet where every node kept its
own settings is a fleet where a module works on one node and fails on another with nothing in
the mesh able to say why.
**The distinction that dissolves both rules.** Neither direction is right as a blanket, because
the question is not *whose configuration wins*. It is whether the module **requires** the
setting or merely **prefers** it — required contradictions cannot be kept without breaking the
module, and preferences should always yield to what is already there.
That is a property of the module's own declaration rather than of the adoption algorithm, which
makes it one more thing the graph would carry
([research 011](../011-the-module-graph/00-overview.md)). Until modules can say which of their
settings are load-bearing, adoption is choosing a default in the dark, and the default chosen
here is the one that does not break the machine it is adopting.
### The briefing carries an outcome, and the outcome is derived
Proposed by the operator: the report states plainly whether adoption succeeded, partly
succeeded or failed, and each line carries its own severity.
**Each line is marked, and the overall is the worst mark present.** Derived rather than stated
alongside, because two fields written independently drift — and a briefing reading *full
success* while carrying a failed line is exactly the fault this record keeps cataloguing. An
outcome computed from its lines cannot disagree with them.
| Mark | Means |
|---|---|
| `ok` | done, and verified |
| `kept` | a disagreement; the machine's value was kept and somebody should look |
| `unknown` | could not be determined |
| `failed` | could not be done — the node is not what was asked for |
**`unknown` is not a shade of success.** Adoption will meet configuration it cannot parse and
state it cannot read, and folding those into *fine* is the same move as reporting an installed
package as a capability. A thing nobody could determine is a thing nobody can rely on, and it
gets its own mark for the same reason a capability detector reports *why*.
**And this opens something the earlier rule did not cover.** *Flags inform, they do not block*
was decided about **conflicts** — where the mesh chose, deliberately, and the machine still
works. A **failure** is different in kind: not *we chose* but *we could not*. Treating both the
same makes a node where something the mesh needed never happened indistinguishable from one
where a log level differed. Whether a failed line still lets adoption complete is therefore
reopened by adding severity, and is not decided here.
## Open questions
| Question | Why it is open |
|---|---|
| What is the closure for a one-node mesh? | The skeleton asserts four pinned services. [Research 006](../006-mesh-from-scratch/00-overview.md) already asks whether it is four or five and does not answer. A graph gives a computed answer instead of an asserted one, which is [research 011](../011-the-module-graph/00-overview.md). |
| Is "tier" the same thing as a graph level? | Tiers were named as a bootstrap order. If the closure is computed, tiers may be a derived view of the graph rather than a separate concept — or they may be a coarser boundary that survives for a different reason. |
| ~~What happens when existing configuration contradicts what the mesh needs?~~ | **Answered** — the machine's configuration is kept, the conflict is flagged, and it is reconciled afterwards. |
| ~~Do flags block, or only inform?~~ | **Answered** — they inform. Adoption always completes, and the node is a node. |
| Can a module say which of its settings are load-bearing? | The question that dissolves the conflict rule rather than choosing a side. A setting the module *requires* cannot be kept from the machine without producing something installed and broken; a setting it merely *prefers* should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph. |
| Does a `failed` line still let adoption complete? | *Flags inform, they do not block* was decided about conflicts, where the mesh chose and the machine works. A failure is *we could not*, which is different in kind — and treating them alike hides the worse one behind the commoner one. |
| How is a flagged conflict reconciled, and by whom? | The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided. |
| Where does the kept original live, and for how long? | Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing. |
| What shape is a briefing? | Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log. |
| Does owning a package mean owning its version? | Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it. |
| How does a bundle stay true between building and applying? | It is built against a scan of the target. The machine can move between the scan and the apply, so the host has to verify rather than assume — and fail plainly when the bundle no longer fits. |
| What cannot be precomputed at all? | Anything built from source on the target still needs a toolchain and a network at that moment. Tailoring moves that cost rather than removing it, and *minimal viable* has to be honest about what it cannot ship ahead. |
| Does presence differencing understate the gap? | Knowing a package manager is installed does not say it is the version the mesh needs. A difference computed on presence alone is optimistic. |
@@ -1,11 +1,12 @@
---
topic: the mesh
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 15. The mesh brokers capabilities; nodes host; agents think
# 1. The mesh brokers capabilities; nodes host; agents think
## Context
@@ -69,6 +70,52 @@ invariants were found violated simultaneously (see Consequences).
**The mesh brokers capabilities. Nodes are places where work runs. Agents are personas
that think and act.** Everything else supports one of those three.
### What this is, plainly — and what "mesh" does not mean
*Written 2026-08-29, from working through connectivity and asking whether the word still fits.*
Four layers. Naming them honestly is worth more than the word on the tin:
| | |
|---|---|
| **machines are linked by a private network** | and every machine reaches every other over it |
| **one node holds knowledge of all of them** | the control plane, and only it |
| **modules are how anything is built and delivered** | this *is* the CI/CD, not something beside it ([ADR 0010](0010-delivery.md)) |
| **every node runs a session you can message** | a feature of the node, remembering across callers; any node can message any other ([ADR 0004](0004-a-node-and-how-it-joins.md)) |
| **workers are hired onto nodes to do tasks** | employees, with a lifecycle — a different thing from the row above ([ADR 0003](0003-agents-are-persistent-employees.md)) |
**The last two rows are not the same thing and the vocabulary of one does not describe the other.**
A node's session comes with the machine: nobody hires it, it holds no tasks, it is never
reassigned, and it goes when the node leaves. A worker is an employee — named, hired, drained,
retired, movable. They are built from the same parts and run on entirely different terms, and
collapsing them is how the employee vocabulary ends up stretched over something it does not fit.
**Neither makes the node itself a thinking thing.** A node is a machine; both of these run *on*
one, which is why *a node does not authenticate to a model provider, agents do* is unaffected by
either.
**Where the value is** is that both can reach across the whole set: a shell, a service, a file, or
simply a question to another node. Not machines that can be configured centrally, which is
ordinary, but a set of machines that can be worked across as though they were one.
**This is not a mesh in the peer-to-peer sense and will not become one.** The word describes what
machines can reach, not how they are governed:
| | a mesh? |
|---|---|
| what a machine can reach | **yes** — genuinely any to any |
| how the traffic travels | no — anything crossing sites transits the hub |
| who decides | no. One node, declared |
**And *master* overstates it in the other direction.** A master implies the others need it in
order to function. They do not: every node holds what it was last told and runs from that copy
*always* — not as a fallback, as the only mode it has. So the control plane being gone is every
node in the ordinary disconnected situation at once, and **what is lost is change, not
operation.**
The accurate phrase is **one authority, no failover**, and both halves are deliberate
([ADR 0006](0006-the-substrate-and-the-control-plane.md)).
### Nodes and agents are decoupled
A node is a place where an agent can run — that is the entire relationship. There is no
@@ -182,7 +229,7 @@ existing pipeline. Nothing here requires a flag day, and nothing here is cheap.
- `modules/hal/sdk/src/feature-handlers/index.ts` — `FEATURE_HANDLERS`, the fixed handler
array that makes a feature a singleton per module
- `modules/postgres/tools/index.ts` — the adoption path that rotates a shared credential
- Mediahuis `papa-hq`, ADR 0009 *Composable, independently-shippable modules* — the
- Mediahuis `papa-hq`, ADR 0020 *Composable, independently-shippable modules* — the
constraints that make a unit independently shippable, applicable unchanged to features
- impire.io / soulstream — *the record* as integration substrate, personas over services,
and "cheap awareness and expensive thinking"
@@ -1,73 +0,0 @@
---
status: accepted
date: 2026-03-14
deciders: jochen
reconstructed: true
---
# 2. Everything is a module, and one manifest describes all of them
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh carries several kinds of thing: containerised services with data and ports, pure
capability providers with no service at all, and bare markers whose only content is that a
node has them. Before this decision these were separate concepts with separate handling —
the earlier vocabulary was *capabilities*, and services were installed by a different path
than tools.
Every distinct kind of thing needs its own install path, its own change detection, its own
place in the delivery pipeline, and its own documentation. Three kinds means three of each,
and every new feature has to be built three times or, more commonly, once — leaving two kinds
quietly unsupported.
## Considered options
1. **Separate concepts per kind** — a service registry, a tool registry, a node feature flag
list. Rejected: it is what existed, and the cost was paid in every cross-cutting change.
2. **One manifest, kind inferred from directory contents.** Chosen.
3. **One manifest with an explicit `type:` field on every module.** Partly adopted — a service
still declares itself — but the general rule became inference, because a declared list and
the directory it describes drift, and the directory is the one that is true.
## Decision
Everything the mesh installs is a **module**: a directory with a manifest. The manifest
declares identity, environment variables, what the module provides, what it requires, and how
it is exposed. What kind of module it is follows from what the directory contains:
| Contains | Is |
|---|---|
| a compose definition | a service |
| a tools directory | a capability provider |
| a daemon or unit directory | a long-running process |
| a configs directory | a source of managed files |
| nothing but a manifest | a flag — presence is the whole content |
A module may be several of these at once. Each is a **feature**, and the delivery pipeline
addresses features, not modules.
The mesh's own components are modules on exactly these terms. They get no privileged install
path, no separate registry, and no exemption from the pipeline.
## Consequences
- One mechanism to learn, one to document, one to fix. A pipeline improvement reaches
everything the mesh carries.
- Dogfooding stops being a discipline and becomes structural: if the mesh's own components
need an exception, the machinery is unfinished, and that is visible immediately.
- Feature detection from directory contents means a directory rename silently changes what a
module *is*. This has bitten repeatedly — a hook named for a feature the module does not
have is skipped without complaint.
- The manifest becomes load-bearing and grows. It is now the largest single point of
coupling in the mesh.
## References
- `Rename capabilities → modules across the entire codebase`, 2026-03-14.
- `Merge fail2ban, ufw, firewall apps into modules`, 2026-03-15 — the first modules to arrive
by conversion rather than by creation.
- Knowledge base: `modules`, `modules/manifest-reference`, `conventions/modules`.
- The rename-breaks-detection shape: `troubleshooting/hooks-named-for-missing-feature`,
`troubleshooting/health-check-tools-index-false-positive`.
@@ -1,11 +1,12 @@
---
topic: the mesh
status: accepted
date: 2026-02-25
deciders: jochen
reconstructed: true
---
# 1. Nodes communicate over a message broker, not over HTTP
# 2. Nodes communicate over a message broker, not over HTTP
> Reconstructed after the fact from the evidence cited below. The decision was taken in
> implementation, not in a record; this document states what was decided and why, not a
@@ -1,11 +1,12 @@
---
topic: the mesh
status: accepted
date: 2026-07-12
deciders: jochen
reconstructed: true
---
# 12. An agent is a persistent employee, not an instance of a pool
# 3. An agent is a persistent employee, not an instance of a pool
> Reconstructed after the fact from the evidence cited below.
@@ -59,7 +60,7 @@ itself is an agent of a kind exempt from the hiring lifecycle.
or another agent is hired — both deliberate acts.
- The transition was not free. Lifecycle columns had to reach every query that selects an
agent, and the ones that were missed failed at the moment of hiring rather than at startup.
- This is the decision [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) generalises:
- This is the decision [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) generalises:
one kind of participant, differing only in modality.
## References
@@ -1,65 +0,0 @@
---
status: accepted
date: 2026-04-02
deciders: jochen
reconstructed: true
---
# 3. The mesh database is the source of truth; the repository is node-agnostic
> Reconstructed after the fact from the evidence cited below.
## Context
Two things must be known to run the mesh: **what exists** — which modules there are, what each
declares, how each is built — and **what runs where** — which node hosts which module, with
which settings, at which version.
The repository is the natural home of the first. It was initially also the home of the second:
per-node directories held that node's configuration, and adopting a machine meant committing
its files. That has three costs. A node cannot be changed without a commit, so runtime state
and source share a review cadence they do not share a rhythm with. Two nodes cannot be
reconciled, because nothing holds both. And the repository becomes an inventory of the
installation, which is exactly the content that cannot be made public.
## Considered options
1. **Per-node directories in the repository.** Rejected — it is what existed. Every binding
change is a commit and a deploy, and the repository accumulates an inventory of one
particular mesh.
2. **Configuration files distributed to nodes and edited there.** Rejected. There is then no
authority: two nodes disagreeing have no arbiter, and drift is invisible until something
breaks.
3. **A mesh database as the single authority, cached locally for resilience.** Chosen.
## Decision
A single database holds every binding: which node hosts which module, at which selection, with
which environment overrides, plus mesh-level settings that all nodes read. The runtime loads
its configuration from that database at startup and falls back to a local cache when the
database is unreachable.
**The repository defines what exists. The database defines what runs where.** No node-to-module
mapping is ever committed.
A node is therefore not described anywhere in source. Bringing one into the mesh is a database
operation.
## Consequences
- The repository becomes node-agnostic, and can be published without disclosing an
installation. This repository's public stance rests on that property.
- A binding changes without a commit, a build, or a deploy.
- The local cache means a node survives losing the database, but a node running from cache is
running from a snapshot with no indication of its age. Divergence is silent by construction.
- The database is the hardest dependency in the mesh. It is also a module, provisioned like
any other, which makes its bootstrap circular — resolved by the first-node initialisation
script, and the reason such a script exists.
- Nothing on a node is authoritative. That is what makes the next decision necessary.
## References
- `Phase 3: rename core modules to hal/ namespace`, 2026-04-02, and the mesh configuration
tables that landed with it.
- Knowledge base: `mesh` — "The repo is node-agnostic. It contains no per-node assignments."
- The stale-cache shape: `troubleshooting/installed-version-and-deployments-are-stale`.
@@ -0,0 +1,263 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 4. A node, and how it joins
*Consolidated 2026-08-28 from four records.*
## What a node is
**A managed machine inside the mesh.** Not a device that is merely known about, not an
unprivileged something. If the mesh does not manage it, it is not a node — it is a client, a
peer, or a thing on the network, and those want their own names rather than a weakened version of
this one.
**A disconnected node is still a node, in a different situation.** Reachability is **state, not
class**. A machine switched off, roaming, or behind a connection that has dropped has not become
a lesser kind of thing; it has a last-known state and a pending set of declarations.
The distinction people reach for is real, but it is **capability** — what this machine can be
asked to do — and that belongs in the host's profile rather than in the definition of a node.
**This is the rule that does the most work elsewhere.** A single control plane is tolerable
because its absence is every node in the ordinary disconnected situation at once. An episodic
host on a phone is that situation more often. Neither needed a new mechanism.
### A node runs one agent session
*Written 2026-08-29. It runs on every node today and appeared in no record, which is how something
deliberate comes to look accidental.*
**A node is a machine. The session is a feature of it** — one of the things running there, like the
host, like any workload. The node does not think; something on the node does. Which is why
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)'s *a node does not authenticate to a model
provider, agents do* holds unchanged: the session authenticates, and it is not the machine.
**It is permanent, and it remembers.** Anything in the mesh can send it a message; it replies; and
what it was asked ten minutes ago is still there next week, alongside what everything else asked in
between — the same way both sides of any conversation remember it.
Its system prompt is the node's **engram** — what makes one node's replies recognisably its own
rather than generic.
**It has its own tools**, and fewer than a session a person is driving directly. So a question can
be answered by going and looking: *what is in our forge*, not only *what is your battery*.
**Messages travel the broker like everything else** ([ADR 0002](0002-nodes-communicate-over-a-broker.md)).
There is no second transport and nothing is dialled.
**Any node can message any node, and this is the one part of the system that is genuinely a mesh**
— symmetric, with no centre. A node that is asked something it does not have can ask another, and
how it passes the question on is its own business: it may say who wants to know, or simply ask. A
person relaying a question makes the same choice, and it follows from the engram rather than from a
message format.
**There is no authorisation between nodes.** Every node is the operator's own, so a message from
one is a message from them, and asking a node something is asking a colleague rather than
presenting credentials. Stated once so it is not discovered later: **the mesh boundary is therefore
the security boundary** — anything inside can reach what any node can reach, which is what puts the
whole perimeter on the token and the overlay
([ADR 0007](0007-connectivity.md)).
**It can be switched off, and switched off it still answers.** A node whose session is disabled
replies saying so, at once, with no model involved — the queue is still read, and the state is the
reply. That is deliberate and it is the same rule the host follows about a service that does not
exist: **absence must never be indistinguishable from a failure to answer.** A node with nothing
there is a silence somebody has to go and diagnose; a node that says *I am switched off* is not.
**One per node, always, and it cannot be moved to another machine.** Two and nothing decides which
replies; none and the node is mute; moved, and one machine is answering as another.
**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it,
it holds no tasks, it drains nothing and it is never reassigned — that vocabulary was written for
workers and does not describe this. It exists because the node does, and it is gone when the node
leaves.
## How it joins
**The host has one behaviour and two sources of declaration.** What differs between the first
node and the fiftieth is not what the host does but where the declaration comes from — and, as
above, that is a situation rather than a class.
| | declaration comes from |
|---|---|
| no mesh reachable | the pinned bundle the host carries |
| mesh reachable | the control plane, over the link |
**The first node is not a different kind of node.** It is a node whose mesh is not up yet. It
applies the bundle it carries, the control plane comes up on top of it, and from that moment it
takes declarations like everything else. **Its specialness is temporary and self-erasing**, which
is what the hand-run bootstrap scripts never were.
**A joining node does the minimum to be reachable and nothing else** — an identity, an address,
and one peer to reach. It does **not** compute the overlay: the whole peer set is derived
centrally and pushed down.
That is also why the migration is smaller than it looked. The hard part of the overlay — every
node's key, address, site and reachability — is only needed to compute the *whole* mesh, and a
joining node needs one peer.
## The link is the security boundary
**Everything reaching a node arrives one way**, and four properties make that a boundary rather
than a pipe.
**It is outbound and node-initiated.** The node dials the control plane; nothing dials a node. Not
only defensive — most nodes sit behind a household connection with no forwarded port, so an
inbound control channel would work for one node and not the rest, and the difference would be
invisible until it mattered. **A node has no listening control surface at all.**
**A node holds its own identity and nothing else.** No shared secret, no credential to anything it
does not own. **Compromise of a node is compromise of that node** — which the current arrangement
does not have, because every node permanently holds the same database and object-store
credentials, and there is no mechanism that rotates one and informs everything holding it.
### What that identity is: a keypair the node generates
*Written 2026-08-29. This is the same rule as the sentence above, and it had been treated as an
open question for weeks because of a word.*
**The node generates a keypair. The private half never leaves the machine. The mesh records the
public half.** Ed25519, the same as the control plane's signing key, in the other direction:
the mesh proves itself to a node by signing, and a node proves itself to the mesh by signing.
**This was never open.** [`08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md) already
says it of the overlay keys, in these words: *each node generates its own keypair, the private key
never leaves the machine, the public key is published to the mesh* — and adds that this **is**
ADR 0004's *a node holds its own identity*, applied. What was missing was applying it to the thing
this record is about.
**The word that caused it:** the lifecycle says a joining node *receives* its own durable identity,
which reads as the mesh issuing something, and then the question becomes *issuing what*. It does
not issue anything. The node arrives holding its identity; what it receives is **being known**.
Enrolment is the moment the mesh writes down a public key it will believe, and the one-time secret
is what buys the right to have it written down.
**Everything above then holds literally.** Nothing is stored that could be stolen and replayed: the
mesh's copy is a public key, so a copy of the mesh's database grants nothing. *Compromise of a node
is compromise of that node* becomes true rather than aspirational, because the only secret on a
machine is the one that identifies it.
### What "connecting to the mesh" is, concretely
*Written 2026-08-29, because it was asked and this record had never said it.*
**One outbound AMQP connection from the node to the broker, held open.** That is all of it. There
is no second connection and nothing is ever dialled *at* a node. Being in the mesh, operationally,
means that connection is up; being disconnected means it is not
([`09-the-node-lifecycle.md`](../03-DESIGN/01-to-be/09-the-node-lifecycle.md)).
**Two different things ride on it, and conflating them is what made this confusing:**
| | what it answers | who issues it |
|---|---|---|
| **an AMQP account** | may this connection be accepted at all | **the mesh, at enrolment** |
| **the node's keypair** | which node is speaking, on every message | **the node**, above |
**The account is the mesh's to issue**, and per node. The broker has to authenticate somebody
before a connection exists, and a shared account would let any node consume another's queue —
which is the shared-credential fault this record exists to remove, reappearing at the transport.
So enrolment creates that node's account and hands it over, and it is rotatable without touching
the node's identity.
**The keypair is not made redundant by it.** With only an account, the control plane would know
which node is speaking *because the broker says so* — and that is the same transitive authority
this record refuses in the other direction. A compromised broker could then attribute reports to
whichever node it liked, and the control plane would act on them. Signing is what removes the
broker from the question in both directions.
**So a node holds two things after enrolment**: a credential the mesh issued for reaching the
broker, and a key it generated itself that the mesh only ever sees the public half of. Both are
its own, neither reaches anything else, and *compromise of a node is compromise of that node*
still holds.
### Its own key, not the machine's SSH host key
Reusing the host key is the obvious economy and it is refused, for reasons that are operational
rather than fastidious:
- **It is regenerated by ordinary events.** A reinstall, an image cloned, `ssh-keygen -A` on a
rebuild — each silently un-enrols the node, and the failure appears as an authentication problem
with no cause anybody changed.
- **It is managed by something else.** Its lifecycle belongs to the machine's SSH daemon, and an
identity the mesh depends on should not rotate on a schedule the mesh does not know about.
- **Not every node has one.** A partial host has no SSH daemon
([ADR 0005](0005-the-node-host.md)), and an identity scheme that excludes a supported kind of
node is not one.
**The mesh should still know the host key** — it knows every node, so it can distribute host keys
the same way it distributes authorised keys
([ADR 0006](0006-the-substrate-and-the-control-plane.md)), and node-to-node SSH stops depending on
trust-on-first-use. That is the good half of the idea, kept.
**Authority is mutual.** The node proves it may join, and the control plane proves it is the
mesh. One-way is not enough: the host applies whatever the link delivers, so a node that cannot
tell the mesh from something impersonating it will apply that something's declarations.
**What may be pushed is bounded by form, not by trust.** Declarations of known shape, never a
command to run. Stated honestly, **this bounds form and not impact**: a compromised control plane
can declare harmful state and the host will apply it faithfully, because that is what it is for.
What the property buys is that the blast radius is *describable* — exactly what the declaration
language can express, which can be reviewed. An arbitrary command channel has no such bound.
## The enrolment token carries the mesh
Mutual authority needs the node to verify something before it trusts anything, and that is a
circle: verifying the mesh needs the mesh's certificate authority, and obtaining one means
trusting whoever hands it over. There is a second circle beside it — a node must reach the mesh
before the mesh has configured it, so it can resolve no mesh name.
**Both are the same shape: a node needs a fact about the mesh before it has any trustworthy way
to obtain one.** So that fact arrives by a path other than the mesh.
**The token carries five things**, and it is the only thing a joining node needs:
| | |
|---|---|
| **who it is** | the name the mesh calls this machine |
| **where** | the broker's **address**, not a name — there is no resolution yet, and this is why none is needed |
| **what it is connecting to** | the fingerprint of the broker's certificate |
| **who it will believe** | the control plane's signing identity |
| **the right to join** | a one-time secret, useless once used and useless after it expires |
*The first row was added 2026-08-30, from raising a mesh end to end for the first time.* It reads
like an oversight and is not: **the node cannot work its own name out.** The name is the mesh's,
chosen when the record was created, and the broker account the node authenticates as is named
after it — so it must be known *before* the mesh can tell the node anything. It is not a secret,
and whoever issues the token already has it.
Without it, enrolment fails at the broker with an empty username and a message about credentials,
which points at everything except the cause. **A missing fact that surfaces as an authentication
error is worse than one that surfaces as a missing fact.**
**Carried out of band**, by the person adopting the machine. That is what breaks both circles:
its authenticity comes from the channel it travelled, not from anything the node can check
afterwards. **Trust on first use, with the first use moved out of band** — the difference between
a pin and a guess.
**The endpoint and the authority are two identities.** A node connects to the broker and takes
instruction from the control plane behind it. Pinning only the broker would make the control
plane's authority *transitive*, and a compromised broker could then forge declarations — which,
since the host applies whatever the link delivers, is the whole machine. So the transport is
verified once at connect, and **each declaration is verified by its signature, every time**.
**What this settles:** the mesh's certificate authority is not a bootstrap concern — it certifies
internal names once a node is a member. Nothing needs name resolution before the link. And
nothing is placed on disk beforehand except the token, which is the first moment *a node holds
only its own identity* becomes true rather than aspirational.
## Consequences
- **Declarations must be signed**, and the host must tell *this is not from the mesh I joined*
apart from *this is malformed*. Rotating the signing identity is a fleet-wide operation with an
overlapping rollover, and that is the cost of not trusting the broker.
- **The token becomes security-critical**, because it carries the pin. Tampering substitutes the
mesh — which is strictly better than the alternative, where there is nothing to tamper with and
the node trusts the first answer unconditionally.
- **A rejoining node is ordinary.** There is no long-lived secret to recover, so a node that lost
its identity gets a new token.
@@ -1,67 +0,0 @@
---
status: accepted
date: 2026-04-06
deciders: jochen
reconstructed: true
---
# 5. Capabilities are provisioned on declaration, not configured by hand
> Reconstructed after the fact from the evidence cited below.
## Context
Most modules need something another module holds — a database, a cache, a bucket, a message
vhost, an identity client. Wiring that by hand means creating the resource, creating a user,
generating a credential, putting it in the consumer's configuration, and repeating all of it
on every node the consumer runs on.
Every step is a place to make a mistake that surfaces much later, and the credential ends up
written somewhere it can be read.
## Considered options
1. **Manual setup, documented.** Rejected. Documentation of a manual procedure is a
description of the mistakes people will make.
2. **A shared credential per resource type**, distributed to all consumers. Rejected: no
isolation, and rotation becomes a mesh-wide outage.
3. **Declared requirements, satisfied by the provider module.** Chosen.
## Decision
A module declares what it **provides** and what it **requires**. A requirement names the
provider, the resource type, optionally a name and a target node, and a mapping from the
resource's connection fields to the consumer's environment variables.
The mesh satisfies it: a provisioner belonging to the provider creates the resource and its
credential, records the grant, and writes the mapped values as database overrides. The
synchroniser from [ADR 0004](0004-managed-files-are-generated-never-edited.md) then
materialises them. Neither the credential nor the topology is ever written by hand.
A requirement may name a provider on another node. The grant records consumer and provider
nodes separately, so cross-node wiring is the same declaration.
## Consequences
- **Provisioning becomes a core concern of the mesh, not plumbing.** A module asks for a
capability; where it lives is the mesh's problem. This is the property
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) later builds the whole domain model
around.
- Credentials are never authored, so they are never authored badly, and they are never in the
repository.
- Each consumer gets its own credential, so revocation is per-consumer.
- Rotation is where this bites. A shared secret rotated for a new consumer invalidates the
peers holding the old one, and this has taken the mesh down. The declaration model makes
granting easy and says nothing about fan-out.
- A module with no requirements skips the stage entirely, which is correct and also means the
absence of provisioning is indistinguishable from provisioning that did not run.
## References
- `Remove shell/ helper library; split brain into independent workspaces`, 2026-04-06 — the
provisioner daemon becomes its own component.
- `Coordinator refactor: centralize pipeline orchestration`, 2026-04-04 — the provision-then-
environment-then-start sequence becomes the coordinator's.
- Knowledge base: `provisioning`, `provisioning/requires`.
- The rotation failure: `troubleshooting/provision-rotation-invalidates-peers`,
`troubleshooting/provision-adoption-rotates-live-credential`.
+155
View File
@@ -0,0 +1,155 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 5. The node host
*Consolidated 2026-08-28 from eight records. Tier 0 is one component and was decided over a
week; the reasoning is kept, the fragmentation is not.*
## It applies; it does not decide
**The host makes a machine match what it was told, and never works out what that should be.**
This is the line the whole tier rests on, and it is not about privilege — it is about what a
single machine can *know*. Deciding needs knowledge the machine does not have: which nodes should
run a store, which peers belong in an overlay, whether a node has been unreachable for a week.
Anything needing a second node is the control plane's.
The practical form: **the host never queries the mesh's database and holds no credential to it.**
Two modules in the current mesh do, and they are the reason every node permanently carries a
database credential.
## It depends on nothing that must be installed first
**A statically linked binary. Copy it onto a machine and run it — that is the whole
installation.** Written in Go, because the job is system-level and because a runtime that must be
installed first would make the host depend on the thing it exists to install.
**What it needs from the machine is not a dependency in this sense.** An init is not installed;
it is what the machine already is. A package manager is the distribution. Those are what a
machine *is*, not what must be put on it before the host works.
## It is built per operating system
**`systemd` and `pacman` are the Arch host's implementation, not abstractions the mesh grows.**
They are not independent choices: a machine has pacman *because* it is Arch, and the package
manager, service manager and packaging format arrive together as one decision somebody made at
install time.
```
mesh-host-arch pacman · systemctl · a container runtime
mesh-host-alpine apk · rc-service
mesh-host-android neither — a partial host
```
**Abstracting them was rejected on correctness, not effort.** The service applier reads systemd's
`LoadState` to tell *not installed* apart from *stopped* — which is what stops it reporting
absence as success — and OpenRC has no equivalent. An interface spanning both must drop it, and
the lowest common denominator is exactly where that fault lives.
**Almost all of it is shared.** The declaration vocabulary, the store, the apply loop, the
read-back discipline, the refusal model and the link are portable. Two appliers differ.
**A host that cannot implement a shape refuses it.** Android has no package manager it may drive
and no init it may register with, so it implements `file`, `directory` and `action` and refuses
the rest — the same refusal an unknown type gets, with a different reason. Those three are the
portable floor, and they are what makes a partial host a real thing rather than a broken one.
## It is a root service, and it never manages its own unit
**Root**, because no useful part of the job is unprivileged: it writes under `/etc`, installs
packages, manages units and runs containers.
**It cannot run in a container**, and the reason is decisive rather than stylistic: installing the
container runtime is a step of the bootstrap, so a host inside a container would need the thing
it exists to install. Everything above tier 0 is a container; the host is not. That split is the
tier boundary made concrete.
**The installation owns the host; the host owns everything else.** It manages `service` resources
and its own unit is one — the temptation is obvious and it ends with a host stopping itself half
way through an apply, leaving a machine with nothing running to fix it.
## An init is asked for one thing
**Start this at boot.** That is all, and every init can express it — systemd, OpenRC, runit, s6.
**Everything else is a launcher the host ships**, which supervises it: restart it when it exits,
count consecutive failures, roll back after too many, halt after that. Policy in a unit file can
only be read and hoped for; a script with a counter can be tested, and this is the one piece that
must work on a machine where the host does not.
**The launcher does not exec the host, it supervises it** — so restarting is ours rather than the
init's. The cost is signals: a supervisor that exits while its child runs leaves the host to be
*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine
this design is about. So it traps the shutdown signal, passes it down, and waits.
**A clean exit is the upgrade path**, and it is the easiest thing to get wrong — twice now. The
host stands aside for a new binary by exiting zero, so anything supervising must restart on a
zero exit and must not count it as a failure.
**Recovery is local, and detection is the mesh's.** Nothing dials a node and a host that cannot
start cannot report, so the node must recover itself. But a local supervisor sees one process
failing and cannot tell a broken machine from a broken release — only something watching every
node can, which is why a host rollout is staged and stops when nodes go quiet.
## A host may be episodic
**Resident or episodic, and both are hosts.** A phone has no init to register with and nothing
worth supervising, because a supervisor would be killed alongside what it supervises. So it runs
when the platform allows and is killed when the platform wants the memory — **and that is
disconnection**, which is already an ordinary situation.
It needs no keep-alive and no new mechanism: the store is already authoritative while
disconnected, reconcile already happens on start, and *last heard from* is already reported
rather than alarmed on. An episodic host cannot be the first node, because every bootstrap step
is a shape it refuses.
## What a declaration is
**An ordered list of resources the host owns.** JSON, because Go's standard library carries a
JSON parser and no YAML, and the one binary whose argument is that it needs nothing must not
gain a parser to buy authoring comfort in a machine-written document.
**Ordered, because ordering is a decision.** The host does not sort and does not resolve
dependencies — that would be deciding, and deciding the thing most likely to differ between what
the control plane intended and what the machine does.
**Every resource has a stable identity** — a name the control plane keeps across declarations, not
a position and not a hash of content. It is what lets the store say *this is the same resource I
applied last time*, which is what makes removal possible at all.
**Unknown is refused, whole.** A field the host does not know is something the control plane
believes it asked for. A declaration naming one is rejected entirely, naming every problem at
once — a host that applied the parts it understood would leave a machine that looks configured
and is not.
**Six shapes:** `file`, `directory`, `service`, `package`, `container`, `action`. Every addition
widens what a compromised control plane can express, so the list is a security artefact and grows
deliberately.
### The bundle may carry actions; the link may not
An `action` runs a command, and the host never learns what it means. It is needed because the
bootstrap creates a database before there is any mesh to ask for one, and the host must not learn
what a database is.
**Permitted from the bundle, refused from the link**, and the asymmetry is the whole point: a
bundle arrives *with* the binary, so anyone able to put a hostile action there could have put it
in the host itself — refusing it buys nothing and costs the bootstrap. The link is a separate
party, reachable separately, and an action there is an unbounded blast radius.
**An action must carry its own verification**, which is also its idempotency check. The host does
not know what a database is, so *is it already there* is a question only the declaration can ask.
## Consequences
- **The migration is smaller than it looks.** A joining node never needs mesh-wide state — it
needs an identity, an address and one peer, and the rest arrives as declarations.
- **What is applied is recorded after it works, never before.** A failed apply leaves the machine
in whatever state it reached, and nothing must claim otherwise.
- **A second operating system is additive**: two appliers and a four-line init file.
@@ -0,0 +1,279 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 6. The substrate and the control plane
*Consolidated 2026-08-28 from six records. Extended 2026-08-29, by building it: the language, and
what must be running before the control plane starts — which this record had left not established
and could not have settled the way it was asking.*
## The control plane is what needs to know about more than one node
That is the whole test, and it follows from the host applying rather than deciding: **deciding
needs knowledge a single machine does not have.**
| question | whose |
|---|---|
| write this file, with this content, with this mode | the **host** |
| which nodes should run the store | the **control plane** |
| is this unit running | the **host** |
| which peers belong in this node's overlay | the **control plane** |
| has this node been unreachable for a week | the **control plane** — nobody else is watching |
**Anything a single machine could answer alone is not the control plane's.**
### Seven contexts and one interface
**inventory, config, connectivity, provisioning, delivery, observability, identity** — plus
`api`, the one interface every surface speaks to. Each earns its place by the test above rather
than by being ours.
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
not need to know a node exists. *Being ours does not make something infrastructure.*
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
stated nowhere.* SSH appears three times across this design and every time as something that
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
thing that knows which humans and agents exist and which nodes they may reach, which is
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
It is three questions wearing one name, and only two of them are the mesh's:
| | |
|---|---|
| **humans** | their key, on the nodes they are allowed on |
| **agents** | the same, with a lifetime — and revocation that has to actually work |
| **an agent reaching another node** | **this is the point of the mesh, not an exception to it** — see below |
**There is no such thing as node-to-node SSH here, and that is a clarification rather than a
restriction.** The actor is always an **agent**; a node is only where it happens to be running —
[ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md): *a node is a place where an agent can
run, that is the entire relationship.* An agent hired onto one node reaching another to do work is
the capability the whole arrangement exists to provide.
**The credential is the agent's, never the node's.** It lives in the agent's own credential
directory ([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)), so a node's
`authorized_keys` lists **agents** and never nodes. Three things follow, and they are why this
shape is better rather than merely allowed:
- [ADR 0004](0004-a-node-and-how-it-joins.md)'s *a node holds its own identity and nothing else*
stays true — no node holds a key that reaches another node;
- a compromised node costs the credentials of the agents that were on it, not a way into
everything;
- **who may reach what is a mesh-wide fact**, which is exactly why it is `identity`'s and not
something arranged locally.
**And it does not conflict with the host having no inbound control surface.** That rule is about
how a node's *declared state* changes: over the broker, never by being dialled. An agent with a
shell is not the mesh reconfiguring a machine — it is what a person with a terminal has always
been, and this design already depends on it working
([ADR 0007](0007-connectivity.md): the overlay is *the way back in*). What such a session can leave
behind is drift, and drift is what reconciliation is for.
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
removed. Listing it as an eighth context would settle by naming what has not been settled by
arguing.
### One node runs it, and nothing takes over
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
built, so none of it can be subtly wrong.
#### The option that would make it a real mesh, and why not
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
to reject anything.*
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
- every node holds the **whole inventory**, so there is a replication process between them;
- replication needs a writer, so one node is elected **master**, and something promotes a new one
when it drops — Redis Sentinel and its whole family of problems;
- and it still would not deliver what the name promises, because **application databases are not
replicated.** A workload's store lives where the workload lives.
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
to become **a replicated database system for everything running on it** — not for our own
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
it would be supporting.
**So there are three central roles, not one**, and it is worth seeing them separately because
only the third costs operation:
| | its loss costs |
|---|---|
| **the control plane** | nothing can be *changed*. Nothing stops running |
| **the broker** | nothing can be told anything, or report anything |
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
**Whether these are one node is not decided here.** All three must be dialable by every node, which
pushes toward one; nothing says they must be.
**That is sound rather than merely cheap**, because the design already tolerates its absence by
construction: a node reconciles from its own store and never needed to ask anybody to hold the
state it was last given. **The control plane being down is not a new failure mode — it is every
node in the ordinary disconnected situation at once.** What is lost is *change*, not *operation*.
The honest half: this node is a single point of failure, recovery is **restore rather than
failover** — which makes backup the availability mechanism rather than hygiene — and
**certificate renewal is the clock.** An outage outlasting a renewal window expires every public
name, which turns an inconvenience into an outage on a timer. Nothing measures that today.
## The authority is the control plane, not a database
**There is no single mesh database.** Each context owns its store exclusively, and *the mesh
database* names a thing that will not exist.
**No node reads any of them** — not for writes, not for reads. A node is *told* what to own, over
the link, in a bounded vocabulary; it **states** what it applied, and the owning context writes.
The difference is the security boundary: something that can write cannot be prevented from
writing anything.
**A node runs from its own store always, not as a fallback.** The current arrangement's nastiest
property is that *a node running from cache looks identical to a node running from the database*,
with no age on the cache and nothing reporting divergence. Under this there is no second mode to
be mistaken for the first.
**What survives from the original decision:** the repository defines what exists, the mesh defines
what runs where, and no node-to-module mapping is ever committed. That is what makes the
repositories node-agnostic and why anything about the mesh can be published at all.
**The error underneath was a category error**: *source of truth* named a storage location when it
meant an **authority**. Once the store is the answer, *which database* becomes the question, and
shared schemas follow.
## The substrate is what the control plane consumes and cannot grant itself
Every module needing a database asks provisioning for one. The control plane needs a database too
and cannot ask itself, because it is not running yet. **That circularity is the definition**, and
anything on the wrong side of it is raised from the bundle the host carries.
| role | product | |
|---|---|---|
| relational store | **PostgreSQL** | its own state lives there |
| message bus | **LavinMQ** | it cannot grant itself a virtual host — and **precedes it**, below |
| object store | **MinIO** | it cannot grant itself a bucket |
| image registry | **an OCI registry** | it cannot grant itself a repository |
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
**The role and the product are both written.** The role is what the argument turns on; the product
is what gets installed and pinned, and a design that names only the role does not record that the
choice was made. **The dependency is on the protocol** — AMQP, S3, OCI — which is what keeps
naming them safe. The store is the exception: the provisioning model uses databases, roles and
schemas as PostgreSQL means them.
**A container runtime is detected, not chosen** — docker or podman, because a machine that
already has one keeps it. Only the version probe differs between them; the behavioural difference
(podman has no daemon, so containers do not return after a reboot unless a unit is enabled)
belongs in the declaration rather than the host.
**Being substrate and being in the bundle are different questions.** PostgreSQL and LavinMQ must
precede the control plane; the object store and the registry are substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.
### Why the broker precedes it too
Written after the fact, because this record first left it *not established* and framed it as
turning on whether the control plane's own contexts talk to each other over the bus.
**They do not** — they are one process and dispatch internally. Under that framing the broker is
provisioned like anything else and the bundle stays at one image.
**The framing cannot answer the question.** What decides it is not how the contexts reach each
other. It is how the control plane reaches a *node* — and that is settled above: never except over
the link, and the link is AMQP ([ADR 0002](0002-nodes-communicate-over-a-broker.md),
[ADR 0004](0004-a-node-and-how-it-joins.md)). So:
```
the bundle raises the control plane
the control plane provisions the broker ← by telling a host to run it
telling a host happens over the link
the link is the broker
```
**And it is not avoided by the first node being local.** `enrol` dials the broker at the address in
its token, which is the first node's own third step. A machine that raised the mesh still joins it
the ordinary way, and that was deliberate — its specialness lasts two commands. Making it join by
some other route would buy a smaller bundle by giving up the property the design was built to have.
The broker precedes the control plane for the same reason PostgreSQL does: **the control plane
cannot grant itself the thing it would need in order to grant it.**
**What it costs.** Two images rather than one, against the wish above to keep the bundle at roughly
one so a person can read it — two is still readable, four would not be. And three things that are
not images, each an action the bundle declares and the host runs, the way the database already is:
a virtual host, a credential on it, and **a certificate**. That last is the awkward one: a token
pins the fingerprint a host must expect *before it sends anything*, so the broker needs a
certificate at a moment when there is no mesh to issue one and no public name to obtain one for.
Self-signed and pinned is the shape that fits; how it is later replaced by the certificates in
[ADR 0007](0007-connectivity.md) is not decided here.
## The control plane is written in Go
The same language as the host, so tiers 0 and 2 are one language and not two.
The reason that decides it is not familiarity. **Its image is pinned by digest in the bundle**,
which means it is fetched and run on a machine where no mesh exists yet — nothing to check it
against, nothing watching, and a person expected to have read the bundle and believed it. A
statically linked binary makes that image the program and nothing else: no interpreter, no package
tree, no transitive dependency that arrived because something needed a date library. Everything
under that line is something somebody would have to audit, on the one image the whole mesh is
raised from.
A second reason, smaller and still real: the control plane runs a reconcile loop of its own
([ADR 0010](0010-delivery.md) — artifacts against source, as the host reconciles machine state
against declarations). Two loops of the same shape are cheaper to hold in one head when they are
also the same language.
**The option rejected** is TypeScript, matching the lab and the surfaces that will speak to this.
The argument for it is that tier 3 is web and CLI, so a TypeScript control plane would share types
with its callers rather than generating a contract. True, and it does not reach far enough:
`mesh-sdk` is *contracts shared across tiers* and **tier 0 is Go**, so the contracts cross a
language boundary whatever tier 2 is written in. The choice is between generating them for one
consumer or for two.
**What it costs, plainly:** the control plane can import nothing that exists today, and a person
moving between tier 2 and tier 3 changes language. Neither is recovered later — the language is the
most expensive thing in this record to reverse.
## The installer fetches what it pins
`substrate.lock` carries **references, not payload** — an image name and a **digest**, fetched at
apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of
a thing rather than carrying its bytes.
The assumption that a machine might have no network came from the lab and was wrong: a machine
being adopted has one, and the sealed case is the lab.
**The lab places images by raising a registry inside the scenario**, which is what a real node
pulls from anyway — so it tests the real path rather than a stand-in for it. The digests that
registry serves are its own, and that satisfies this rule: what is required is a reference that
is **exact and cannot move**, and one it assigned is both. Assuming an upstream digest had to be
preserved is what made this look impossible for a while
([04-ISSUES/009](../04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md)).
**Its contents are per operating system** even though its mechanism is not — package names, unit
names and service names all differ, so an Arch host embeds an Arch bundle.
## Consequences
- **The bundle stays small and reviewable.** A list of pinned references is something a person can
read; a bundle containing images is not.
- **An apply can fail because something is unreachable**, which a self-contained artifact could
not. That must fail *legibly*, naming what could not be fetched and from where.
- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across
contexts consumes their events or calls their interfaces.
- **A queue with no limit grows until the broker's disk is full**, and the broker is what every
node depends on. The bound is per queue and is not decided.
- **The bundle carries two images and four actions**, and the substrate bootstrap grows a step.
- **Nothing in the first node's path is special-cased.** Enrolment is walked on node one.
- **The broker's certificate at bootstrap has no answer yet**, and is named as unfinished rather
than assumed. It is the first thing that will be wanted when the link is built.
- **The language cannot be revisited cheaply.** It is the one line here close to irreversible.
+191
View File
@@ -0,0 +1,191 @@
---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 7. Connectivity
*Consolidated 2026-08-28 from three records. Overlay, resolution, exposure, filtering and
certificates are one design.*
## Why it is control-plane work
Apply the test — *everything that needs to know about more than one node* — and not one of the
five can be answered by a machine on its own:
| | needs to know |
|---|---|
| **overlay** — who peers with whom | every node, and which can be dialled |
| **resolution** — which name is which node | every node |
| **exposure** — which public name reaches which container | which node is publicly reachable |
| **filtering** — which port is open, to whom | what is assigned here, and the overlay's shape |
| **certificates** — who may present which name | which name belongs to which node |
That is exactly what the current arrangement gets wrong, by computing all five on the node from a
direct database connection. Two modules do this, and they are the only two left holding a
credential to the control plane's database.
**The shape of the fix, once for all five:** the connectivity context computes the configuration;
it arrives over the link as `file` resources; the service reads files and knows nothing about the
mesh. **This costs no new host vocabulary.**
## A route is a grant
**Ingress is not substrate.** The control plane does not need a route to start — it listens
locally — and no node needs one to reach it, because the node dials out and has no listening
control surface. It grants itself a route afterwards, the way it grants itself a bucket.
The strongest objection deserves stating: the `api` is the one interface every surface speaks to,
so eventually it *does* want a public name. But **wanting one later is not needing one to
start**, and that distinction is the entire substrate test.
**A module that must be reachable declares it needs a route; the proxy provides one.** Ordinary
instantiation, with the direction mirrored — the consumer supplies a target and receives a name.
**Exposure is three facts at two scopes**, which is why it cannot live on the node:
| the fact | scope |
|---|---|
| the public name resolves to an address | **mesh** — which node is publicly reachable |
| a certificate valid for that name exists | **mesh** — issued once, used on one node |
| the proxy maps that name to that container | **node** |
**A node without a public address is proxied by one that has**, across the overlay. Most nodes sit
behind a connection with no forwarded port, so exposure cannot assume the workload's node is
reachable.
## Reachability is declared, not inferred
The overlay's peer graph is computed from whether a node can be dialled, and that was inferred
from a regular expression over the address. **The address is evidence of reachability; it is not
the fact**, and the gap has already cost:
| address | the regex says | actually |
|---|---|---|
| `100.64.0.0/10` — carrier-grade NAT | **public** | **not reachable.** An endpoint is written to an address nothing can reach |
| any IPv6 address | public | the test is v4 shapes only |
| a routable address behind a closed firewall | public | not reachable |
| a documentation range standing in for a public segment | private | reachable — this is the lab bug |
**A test environment having to choose its addresses to satisfy a regex is the regex telling us it
is not a fact.**
So: **an endpoint, or none** — declared. And **the hub is declared, never derived from an address
prefix**, because an election decided by the first four characters of an address fails silently,
cannot be queried, and makes a renumbering an outage.
**The address remains evidence and stops being the fact.** Where an observed endpoint disagrees
with a declared one, the disagreement is a **reportable condition**, not a silent correction.
**What does not change** is the lesson underneath: role does not imply reachability — a
home-hosted node is a server that cannot be dialled. This keeps that and stops encoding it as a
pattern match.
### Some nodes must be reachable, and this had not been said
*Written 2026-08-29, on being asked and finding no answer.*
Everything above treats reachability as a **fact to record** — which node can be dialled, so that
exposure and certificate issuance can be placed. It never said the converse, and the converse is a
hard requirement:
| role | dialled by | so it needs |
|---|---|---|
| **the node running the broker** | every node, outbound ([ADR 0002](0002-nodes-communicate-over-a-broker.md)) | to be reachable from wherever nodes are, at a **stable address** |
| **the hub** | every node not co-located with its peer | the same |
**Across the internet, "reachable from wherever nodes are" means publicly reachable.** For a mesh
confined to one network it does not — the requirement is about the nodes that exist, not about the
public internet.
**Stable is the sharper half.** A token carries the broker's *address, not a name*, because there
is no resolution before joining ([ADR 0004](0004-a-node-and-how-it-joins.md)). A broker node whose
address moves invalidates every token issued for it, and a node that was disconnected across the
change cannot get back.
**A mesh whose nodes are all behind NAT cannot be raised.** That is a real precondition and it
belongs with the others rather than being discovered.
### The link stays on the underlay, and that is a repair channel
The obvious objection is that all traffic should run over the overlay. Nearly all of it does — SSH,
services, node to node — and the exception is each node's own outbound link to the broker.
**At join time it is forced**: a node has no overlay yet, so it cannot use one to ask for one.
**Afterwards it is a choice**, and the reason is that the link is how a broken node is fixed. A
repair channel carried over the thing being repaired is not a repair channel: a node whose only
path home was the overlay is gone the moment an overlay declaration is wrong.
**What it does not cost is confidentiality.** The link is already authenticated and encrypted
against a pinned fingerprint ([ADR 0004](0004-a-node-and-how-it-joins.md)), so moving it onto the
overlay would not protect traffic that is unprotected today.
## A filter rule names its source
`scope: public` is declared in five manifests, is part of no rule type, and is **referenced by no
code**. So five manifests appear to restrict a port and restrict nothing — on the modules most
worth restricting.
**A rule names its source. `from:` is the only way to scope one, and a rule without one is open**
— which it must say plainly rather than appear to deny.
**`scope:` is removed rather than implemented**, because giving it meaning would leave two ways to
express one thing. And the general fix is that **an unknown key is refused**: the host's
declaration parser already works this way, and manifests are the layer where that discipline is
missing. `scope:` survived because nothing rejected it, and it spread by copying to five
manifests.
## Order, and what it costs
**The link runs on the underlay and never on the overlay.** The overlay is configured by the mesh,
so a link requiring it could never be established on a new node.
**The first declaration is the overlay and nothing else** — because a node's address and peers are
*assigned* so it cannot come earlier, and because it is the way back in. A node reachable over the
overlay can be fixed by hand if a later declaration breaks the machine; **a large first
declaration risks a node that is broken and unreachable at once.**
**Reachable is not the same as having a control surface.** Every node reaches every other over the
overlay — SSH, services, ordinary traffic — and every node consumes from the broker. What is
forbidden is a listening thing that accepts instructions and changes the machine.
## Consequences
- **The last two direct database connections leave the nodes**, and with them the database
credential every node carries.
- **The `/etc/hosts` floor goes**, along with the bootstrap circularity it patched.
- **Two certificate authorities stay separate on purpose**: a public one for public names, the
mesh's own for internal ones. A single-CA lab would hide any bug living in the split.
- **What happens when the hub is down**: nothing takes over. Non-co-located paths stop; co-located
peers and every assigned workload keep running.
## Open — the link over the overlay, with a fallback
*Raised 2026-08-29 and deliberately left open, because the honest gain is smaller than it looks
and it is a decision rather than a derivation.*
The proposal: a node prefers the overlay for its link and drops to the underlay when the overlay
is not working — so ordinary operation is private and the underlay stays as the way back.
**Two things it would have to get right:**
- **The trigger cannot be "is the overlay up".** A WireGuard interface has no link state; once
configured it is up whether or not the far end exists. So there is no flag to read, and failing
over means *try, fail, time out, retry elsewhere*.
- **Running on the fallback has to be visible.** A node that quietly drops to the underlay is a
node whose overlay is broken with nothing to say so, and it will stay broken because everything
still works. That is this repository's recurring fault — a failure that reads as success — and a
fallback is the easiest place in the design to reintroduce it.
**What stops it being an obvious win:** if the underlay path must stay available for the fallback,
the broker stays exposed on it. So the exposure is unchanged and the traffic was already encrypted
— the gain is which network carries bytes, not what an attacker can reach.
**The version that would buy something is overlay-only**, with the broker firewalled to the overlay
in steady state, accepting that a node whose overlay breaks needs hands-on recovery. That is a real
trade: it exchanges the automatic way back in for a closed port.
Not decided either way here.
@@ -0,0 +1,114 @@
---
topic: the tiers
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 8. A context owns its store, exclusively
## Context
[ADR 0009](0009-modules-and-the-graph.md) settles what a module
declares. This settles what a grant may be, and it is the half that **removes** things.
`how-we-build` §4 already says *contexts integrate through the record, never through a shared
schema*, and states the cost: several domains share one forty-five-table schema, which is why
work belonging to one context keeps having to be implemented in another.
That was written as a principle. Counted, it is thirteen foreign tables belonging to three
separate contexts, living in the mesh's own registry database.
## Considered options
1. **A schema per consumer inside a shared database.** Namespaced, revocable by dropping the
schema, with a cross-context join possible but deliberate. Rejected: it keeps the letter of
§4 and leaves the temptation in place, and a boundary that is merely inconvenient to cross
gets crossed.
2. **Read-only roles on another context's store.** Rejected for the same reason and one worse:
reading another context's tables couples you to its layout exactly as firmly as writing them,
and the coupling is invisible until the owner changes a column.
3. **Exclusive ownership.** Chosen.
## Decision
> **A context is granted only what it exclusively owns.**
No shared writes. No read-only role on another context's store. If you need what another context
holds, you ask it or you subscribe to it.
**The unit is the context, not the process.** Everything inside a context — its service, its
surface, its tools — reads its own store freely. A board showing the mesh's own nodes and
modules is the mesh showing its own data, not a boundary crossing. What is forbidden is a
*different* context reading it.
### Asking or subscribing is derived, not chosen
[ADR 0004](0004-a-node-and-how-it-joins.md) makes disconnection an ordinary situation. So:
- **Anything that must keep working while disconnected cannot ask** — there is nobody to ask. It
keeps a local copy, which means subscribing.
- **Anything where a stale answer is worse than none cannot subscribe.** A display may lag; a
decision about whether a grant is still valid may not.
Neither is a query against another store, whatever transport it travels over.
### What that is, concretely
*Written 2026-08-29, on building the first one — the rule above was clear and what to type was not.*
**One PostgreSQL database per context, named for the context.** A separate database rather than a
separate schema is the whole point: a cross-schema join is a qualified name away, and a
cross-database join needs a foreign data wrapper somebody has to install and explain.
**And one credential per context, held only by it.** There is no mesh-wide connection setting and
no way to ask for one, so reaching another context's store is not a matter of restraint — a process
has no address for it and nothing to present. That is also how this rule is *checked*: what a
context can reach is the list of variables the declaration running it grants, and it is read there
rather than audited in code.
**Contexts that do not exist yet do not get a database.** The bootstrap creates the ones there are.
**How a context added later gets its database is open**, and it is a real question: by then there is
a control plane, but a control plane holding a credential that can create databases is holding
rather more than the thing it exclusively owns.
## What this removes
The first clear list of what the design deletes rather than adds:
- **Grant kinds.** There is one: an exclusive resource. No schema grants, no read roles, no
rules about who may see what inside a shared thing.
- **The question of who owns which table**, and the guessing at revocation time. Removing a
consumer drops what it was granted, whole.
- **Cross-context migration ordering.** Two contexts migrating one database must be ordered
against each other. Exclusive ownership means a context's migrations are ordered only against
itself.
- **A class of permission modelling** a shared store would otherwise need.
## Consequences
- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across
contexts consumes their events or calls their interfaces. That is §4's argument, and the cost
it names is the one already paid.
- **A single surface over several contexts still works** — that is what a surface is. It reads
interfaces, not stores. This holds while the contexts sit behind **one** interface; splitting
a context into its own deployable costs that, and the composition would have nowhere to live
that tier 3 permits. **A real constraint on how far the control plane may be split.**
- **Three contexts must move out of the registry database**, taking thirteen tables with them.
Their dependency on the registry then shrinks to almost nothing — one of them needs a single
table.
- **The node appliers were already handled.** [ADR 0005](0005-the-node-host.md)
stopped the host querying the mesh database for tier reasons unrelated to this, and it removes
most of the remaining direct readers as a side effect.
- **What a consumer does about events missed while disconnected is not decided** — replay from a
point, ask once and resume, or rebuild. The question every projection has.
## References
- [`how-we-build.md`](../00-META/how-we-build.md) §4 — the rule this makes enforceable.
- [Research 011](../01-RESEARCH/011-the-module-graph/worked-provider.md) — the count, the worked
provider, and the dashboard case.
- [ADR 0004](0004-a-node-and-how-it-joins.md) — why asking or subscribing is derived.
@@ -1,71 +0,0 @@
---
status: accepted
date: 2026-06-05
deciders: jochen
reconstructed: true
---
# 8. A step that fails must fail the job
> Reconstructed after the fact from the evidence cited below.
## Context
The mesh's expensive faults are not crashes. They are the operations that reported success and
did nothing: an artifact that partially downloaded and was extracted anyway, a package that
404ed from every mirror while the job went green, a hook that never ran because it was named
for a feature the module does not declare, a deploy that reported the transport succeeded
rather than that the effect happened.
Each of these was found long after it happened, by someone investigating an unrelated symptom.
The cost is not the failure; it is the interval between the failure and anyone learning of it,
during which decisions are made on the assumption that the thing worked.
## Considered options
1. **Continue on error and report at the end.** Rejected — it is largely what existed. A
summary nobody reads is not a report, and later steps run against the state the failed step
should have produced.
2. **Continue on error, and let health checks catch the divergence.** Rejected. It converts a
precise, located failure into a vague one discovered elsewhere, and requires a health check
for every possible partial state.
3. **Fail the step, fail the job, say which step.** Chosen.
## Decision
A step that fails stops the sequence it is part of, and the failure is surfaced where the work
was requested — not only in a log.
Concretely, and these are the forms it takes:
- A scripted sequence gates each step on the previous one. A directory change that fails must
stop the commands that assumed it.
- An artifact that does not fully download is not extracted.
- A stage reports the **effect** it achieved, not that it dispatched a message. "Started" must
mean the thing is running, not that a command returned.
- A template that cannot resolve a variable is not written half-rendered.
**Prefer failing to lying.** A green result that is not true costs more than a red one.
## Consequences
- Failures are noisier and land earlier, on the person who caused them.
- Some jobs that used to complete now stop. In every case examined so far, that job was
producing a partial result that something downstream trusted.
- This is a rule the mesh has adopted repeatedly rather than once, because each instance is
written in a different place — a shell hook, a download path, a deploy stage. It is not
enforced by a mechanism, and cannot currently be checked in general. New instances are still
being found; the package-install case remains open as
[`04-ISSUES/001`](../04-ISSUES/001-failed-package-install-reports-success/00-report.md).
## References
- `fix(installer): fail loudly when feature artifact download fails` (#244), 2026-06-05.
- `A flavor template with an unresolved variable is written to disk instead of failing`
(#710), 2026-08-08.
- Knowledge base: `troubleshooting/deploy-reports-transport-not-effect`,
`troubleshooting/service-started-is-not-ready`,
`troubleshooting/green-pipeline-means-transport-not-effect`,
`troubleshooting/silent-failures-and-stale-state`.
- The core value it became: [`00-META/mission.md`](../00-META/mission.md), "Failure must
be loud."
+441
View File
@@ -0,0 +1,441 @@
---
topic: what runs on it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 9. Modules and the graph
*Consolidated 2026-08-28 from six records.*
## Everything is a module
One kind of thing, one manifest describing all of them. A database, a web application, a window
manager and a firewall rule set are all modules — not because they are alike, but because
**anything else means a second kind of thing with its own rules, and then a third.**
**A module is the unit of delivery**: assignable to a node, versionable, replaceable on its own.
## There are no domain modules
An earlier decision grouped modules by domain — four things constituting *how a node is
reachable* becoming one `networking` module. **That was wrong, and the correction is worth
keeping** because the observation behind it was right.
The measurement holds: reachability is the **only** place in the catalogue where modules
genuinely change together under one intent. What did not hold is the conclusion. Tight coupling
means they share an **authority** — one place that decides for all of them — and not that they
should be one artifact. `wireguard` and the proxy are deployed to different sets of nodes, so a
module containing both would be assigned where half of it is unwanted.
> **Coherence is a context. Delivery is a module.**
**Folders assert relationships; edges record them.** What grouping was for — finding things,
seeing what belongs together — is a tag and a query, neither of which anybody has to keep true by
hand.
### What a domain module turns out to be, and why it is not the one refused above
*Written 2026-08-29, from building it. The heading above reads as a contradiction of what now
exists and is not one — but only if the difference is stated, so it is stated here.*
**What was refused contains things. What exists contains nothing.**
| | `networking` as refused | `networking` as built |
|---|---|---|
| what is in it | WireGuard, a proxy, a firewall — artifacts | nothing at all |
| what it says | *these ship together* | *I want a private network and names* |
| what is assigned | one module, half of it unwanted | whatever answers each requirement, each on its own |
The objection above is untouched by this and still correct: a module holding WireGuard and a
proxy is assigned where half of it is unwanted. **A module holding nothing cannot be, because
there is no half.** It is requirements and a name, and every artifact it leads to is still an
ordinary module assigned on its own terms.
**Why it is worth having.** Most people want the network working and do not want to choose a VPN.
`assign networking` finds one answer to each requirement and takes it without asking, because
with one candidate there was never a question — the rule below about refusing does the work.
Somebody who does care assigns the VPN they want, and *that is the whole of choosing*: there is no
flavor field, no variant syntax, and no second verb. **Picking an implementation is assigning a
module.**
**What it costs, stated because it is real.** Adding a second implementation to the catalogue
turns a settled question into an open one for **everyone using the bundle**, not only for whoever
wanted the alternative. Every node assigned `networking` refuses until somebody says which. That
is [the refusing rule](#a-requirement-with-several-answers-is-refused-never-guessed) applied
consistently, and the alternative is a default — which is the flavor field returning under a
better name. The cost is one assignment per node, and the message names the candidates.
**A consequence that had to be found by running it.** A bundle can drag an implementation in
through a requirement nobody looked at. Choosing a different VPN still installed WireGuard,
because the names module needed addresses only WireGuard hands out, and nobody was told. Two VPNs
on one machine is not always wrong — a machine may run one for another purpose — but being **the**
network the mesh runs over is singular, so that is a claim, and the collision is refused by name.
**The general rule: what a bundle pulls in is only as safe as the claims on what it pulls in
from.**
## Three edges
| edge | means | declared? | satisfied |
|---|---|---|---|
| **presence** | that thing must exist and be reachable here | yes | at provisioning |
| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes | at provisioning, and again whenever it must be |
| **build** | I was compiled against that artifact | **no — read from imports** | **at build, once** |
**Instantiation implies presence; presence does not imply instantiation.**
**A route is an instantiation edge**, and it is worth noticing because the direction is the mirror
of a database: the consumer supplies a target and receives a *name*, rather than supplying nothing
and receiving credentials. Same edge.
**Provider stops being a category.** Any hosted thing can be a factory — an identity provider
grants clients, a mail server grants mailboxes. It is a facet, not a kind.
**A module may also declare what it claims**, because some things cannot coexist and that is a
fact about the module rather than about a particular node. What that means precisely is below.
### Where the answer to a requirement is allowed to live
*Written 2026-08-29, from building it. The table above distinguishes **presence** from
**instantiation** and this is the half of that distinction nobody had noticed was missing: not
what the edge hands over, but **where the thing on the other end is.***
Two different things were both being written as a requirement:
| | *a shell*, *a display server*, *a private network* | *a database*, *an object store*, *an identity provider* |
|---|---|---|
| where the answer lives | **this machine** | **somewhere in the mesh** |
| how it is answered | install another module here | find the node already running it |
| what is missing if absent | a module to assign here | **a decision about where**, which is nobody's to make silently |
Answering the second like the first installs a database on every machine that uses one, which is
what it did.
**So a provided name carries a scope**, the same idea a claim already has, and written short in
the ordinary case so the few that are not node-scoped stand out rather than drowning. Scope is a
property of **the name, not of each provider**: two modules disagreeing about whether a database
is local would make one requirement mean two things depending on which happened to answer it, so
that is refused.
**A requirement answered from the mesh is never satisfied by installing it here.** Nothing, and
the mesh refuses and says which module to assign somewhere. Two, and it refuses and says how to
choose — the same rule as everywhere else, for the same reason: picking is guessing, and the wrong
guess puts somebody's data on a machine they did not choose.
**Choosing is recorded per node**, because that is the granularity the choice actually has — two
machines may reasonably use two different databases and a mesh-wide answer could not say so. A
choice pointing at a machine that does not provide the thing is refused rather than quietly
replaced by one that does, and a single available provider does not override a choice either.
**Both are the same rule: the mesh does not overrule a person, and it does not move data without
being told to.**
**What this is a prerequisite for.** Knowing *which node* answers is the first half of handing a
credential back — you cannot be given a database's password before it is settled whose database it
is. So a node's resolution now records what it takes from elsewhere, which is both the only part
of its set that stops working when a *different* machine goes away, and the place a credential
will hang.
### An edge has two directions, and only one of them is built
*Written 2026-08-29, from building it. The row above already says a consumer **supplies a target
and receives a name**; what it did not say is that those are two separate mechanisms, and that
having one without the other is what forced two modules outside the system entirely.*
| direction | the consumer says | who needs it |
|---|---|---|
| **contribution** | *publish me at this name, on this port* | the proxy, the DNS server, a firewall |
| **binding** | *and give me back a credential to it* | the database, the object store, the identity provider |
**Contribution is built.** A module declares what it contributes to a requirement; the control
plane collects every contribution on a node and writes them to a path the provider named, as a
file, in the mesh's own shape. **Contributing to something is requiring it** — asking to be
published means a publisher must exist, and a module that had to say both would eventually say
one, with the failure appearing as a machine where nothing serves the route.
**The control plane does not know what a reverse proxy is**, and does not write one's
configuration. It delivers the facts; the module turns them into whatever it runs. That boundary
is what makes swapping the proxy cost nothing in any module that publishes through it, and it is
[the same separation](0001-mesh-brokers-nodes-host-agents-think.md) that keeps third-party
software *on* the mesh rather than *of* it. It also costs the host nothing: a received file is a
file, which was checked by putting the control plane's output through the host's own parser rather
than by asserting it.
**Binding is built except for the secret**, and that turned out to be the useful way to cut it.
A provider says what a consumer needs in order to use it — a port, a driver, a realm — and a
consumer says where it wants to be told. The mesh adds the half only it has: **which machine, and
what that machine is called on the private network.** So an application on one node is handed the
address of its database on another, as a file, and reaches it by a name the mesh also created.
**The file states that it carries no credential, and why.** A missing field looks like a bug; a
stated absence looks like a boundary, and somebody wiring this up should not spend an afternoon
looking for a password that was never going to be there.
### And the secret, which is delivered without ever being held
*Written 2026-08-30, after looking at how the existing mesh does it. The design here is a reaction
to a measurement, not a preference.*
**The obvious arrangement is a credentials column, encrypted at rest.** It exists, and its own
tooling records what it bought:
| | |
|---|---|
| the tool for finding a secret matches **by value**, not by name | because one password is in the provisions table, the environment table, each node's environment file in plain text, and **inside every connection string composed from it** — copies its documentation calls *"often the only copies actually in use"* |
| a query against the encrypted column **returns zero rows and proves nothing** | so auditing moved to the decrypted copies on the machines |
**Two faults, and encryption at rest addresses neither.** The control plane can read what it
stores, so a copy of its database is a copy of every credential in the mesh. And one secret has
many homes with nothing tracking them — **composition is what mints the untracked ones**, because
building a connection string centrally creates a new secret-bearing value no rotation path knows
about.
**So the value is sealed to the node that will use it before it is stored.** With a key that node
generated and whose private half the mesh has never seen — a third key beside the identity and the
overlay, for the same reason those are two rather than one. What is stored is unusable by whoever
holds it, the mesh included, and the broker relays a blob it cannot read. This is what makes
[ADR 0004](0004-a-node-and-how-it-joins.md)'s *compromise of a node is compromise of that node*
true of secrets rather than true of identity and quietly false of everything that matters.
**And nothing is composed centrally.** A connection string is assembled on the machine that needs
one, if at all. The mesh delivers parts.
**What it costs, stated because it is real:** the mesh cannot audit by value. That is the right
trade rather than an oversight — a query over an encrypted column could not either, so the audit
was never real. What *is* answerable is which node holds what, which is the question rotation
actually asks.
**A consequence that shapes the mechanism.** The mesh discarded the plaintext, so it cannot
compose a file containing it. The credential is therefore **its own file**, holding the value and
nothing else, beside the readable one. That is better than the alternative it was forced into:
the readable half stays readable in the declaration, and the secret half changes only when the
secret does, so a service reloading on it reloads for a real reason.
**Rotation is generating a new one**, because reading the old one back is not possible. Both ends
are re-sealed and reach their machines in the same push — which removes the window where half the
mesh holds a dead credential, the failure
[recorded in ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) as consumers on three nodes
holding one for two days.
### The provisioner, which is where the mesh stops
*Written 2026-08-30, from building one and running it against a real database.*
**A password nothing was told to create authenticates nowhere.** The mesh generates one, seals it
to both ends and cannot read it — so it cannot tell the software to start accepting it either.
Something on the providing machine reads what arrived and makes it true. That is a provisioner.
**It belongs to the module, not to the mesh**, and the boundary is the same one that keeps
third-party software running *on* the mesh rather than being *of* it
([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)). The control plane decides and never
touches a machine. **What the mesh owns is the contract**, which is two files the host writes from
an ordinary declaration:
| | |
|---|---|
| the manifest | every consumer, what it asked for, and **where** its credential is |
| one file per consumer | that credential, alone in it |
Two files because the mesh discarded the value and cannot compose a document containing it. As
before, the constraint produces the better shape: the readable half stays readable and auditable
in the declaration, and the secret half changes only when the secret does.
**It reconciles; it is never told what changed.** It runs after every declaration and must reach
the same state from wherever it starts. Three consequences, and each of them is a fault that has
been shipped somewhere:
- **the password is set every time, not only on creation** — otherwise the role already exists,
nothing happens, and a rotation reports success while changing nothing
- **what it made and nobody asks for any more is removed** — otherwise a consumer that left keeps
a working login for ever and nothing ever says so. This is the same rule the host follows about
[removing what it declared and no longer declares](../04-ISSUES/010-the-first-declaration-destroys-the-substrate/00-report.md)
- **what it did not make is left alone** — otherwise it cannot be run on a system that predates
it, which is every system anybody would want to adopt
**A missing credential is refused rather than worked around.** A role created without one is a
login nothing can use, and nothing would report it until something tried to connect.
**This is where the mesh stops**, and saying so is the point of the section. It decides, delivers
and can prove what it delivered; the last inch belongs to whoever knows what `create role` means.
**One check that only became possible now.** Two machines wired together across no private network
is a mesh that reports itself configured and does not work, and the failure surfaces as a
connection timing out — the slowest place to find anything. It is refused, and it is only
*checkable* because the network became [something a machine is
given](#what-a-domain-module-turns-out-to-be-and-why-it-is-not-the-one-refused-above) rather than
something it has by virtue of holding an address.
**What the absence cost, measured.** Exactly two modules opened a direct connection to the control
plane's database — the proxy and the VPN — and they are the reason every node permanently holds a
credential to it. Both were doing by hand what this edge is for. The VPN's half is closed by being
[a module whose files are computed](../03-DESIGN/01-to-be/08-connectivity.md); the proxy's is
closed by contribution. **Neither needed a new kind of thing, and both had been outside the model
for as long as there was one.**
### Why the build edge is a different kind
It is fixed inside an artifact rather than negotiated when something runs, and **its only remedy
is a rebuild** — nothing can re-provision it.
It is also **derived rather than declared**, and the asymmetry is deliberate: a runtime edge is an
*intention* somebody has about how the mesh should be wired, and only a person can state it. A
build edge is a *fact about code that already exists*, and a declared list of dependencies drifts
from the imports it describes.
**An artifact is out of date when its source moved, or when anything it was built against moved.**
So what is recorded is a commit *and the identity of every artifact it was built against*, which
is what makes the rebuild set computable and *is this current?* answerable without building.
**The graph measures design quality, not just build order.** A module with many inbound build
edges is one whose every change is expensive — and that is readable before anything is built. The
current shared library is exactly that, and nobody could see it because nothing drew the edges.
## Provisioning is declared, never configured by hand
A module declares what it **provides** and what it **requires**. The mesh satisfies it: a
provisioner belonging to the provider creates the resource and its credential, records the grant,
and the values are derived onto the consumer. **Neither the credential nor the topology is ever
written by hand.** A requirement may name a provider on another node, so cross-node wiring is the
same declaration.
## What a module claims, and why it is not a list of rivals
*Written 2026-08-29, replacing pairwise exclusion.*
**Exclusivity is not a property of a module. It is a property of a singular resource the module
takes over.** Two shells do not compete for anything and any number may be installed. Two display
servers both want the seat, and only one may have it.
> **A module declares what it *claims*. Two modules claiming the same thing cannot both be
> assigned within that claim's scope.**
**Not "xorg conflicts with wayland".** Pairwise exclusion has a property that only shows up later:
adding a third display server means **editing xorg and wayland to know about it**. Every new
module requires changing modules nobody who wrote it owns, and the edits grow as the square of
the count. With a claim, the third one says `claims: the seat` and nothing else changes anywhere.
**The new module is the only thing that has to know anything** — which is the difference between
a catalogue that grows and one that calcifies.
The pattern is common enough to be worth listing, because seeing it is most of understanding it:
| these coexist | these claim one thing |
|---|---|
| shells — bash, zsh, fish | display servers — xorg, wayland (*the seat*) |
| editors — vim, emacs, helix | init — systemd, openrc (*pid 1*) |
| language runtimes | container runtime — docker, podman |
| terminal emulators | reverse proxies — nginx, caddy, traefik (*ports 80/443*) |
| browsers | time — chrony, timesyncd, ntpd (*the clock*) |
| | resolvers — resolved, dnsmasq, unbound (*`/etc/resolv.conf`*) |
| | network management — NetworkManager, networkd, netctl |
| | mail — postfix, exim, msmtp (*port 25*) |
| | audio — pipewire, pulseaudio (*the device*) |
**A claim has a scope**, because not everything singular is singular per machine:
| scope | example |
|---|---|
| **node** | the seat, pid 1, port 443 |
| **site** | a DHCP server on a segment |
| **mesh** | the hub, the control plane |
The last is not new — the mesh already enforces exactly one hub with a unique index
([ADR 0007](0007-connectivity.md)). Scope is that idea, said once rather than hard-coded per case.
**Some conflicts need no claim at all.** Two modules declaring the same file, or binding the same
port, are visible from *what they declare* — the mesh already holds every resource of every
declaration. So a claim is only written for the abstract ones, where nothing in the declaration
reveals the clash. That keeps the manifest small, which is worth protecting.
## A requirement with several answers is refused, never guessed
A module requiring *a shell* may be satisfied by three. The mesh does not pick.
| candidates | what happens |
|---|---|
| exactly one | assigned, silently — there was no choice to make |
| none | refused, naming what is missing |
| several | **refused, naming them**, and a person chooses |
**This is what makes a solver unnecessary.** Counting candidates is a few lines and has no
surprising behaviour; a solver that picks has to be understood before its answer can be trusted,
and it is understood by whoever is debugging it at the time. Nothing here is lost by waiting —
a solver can be added later without changing a single manifest, and the reverse is not true.
**Requiring a module and requiring a capability are different fields**, because the remedies
differ and the message should say which:
- *i3 needs xorg, which is not assigned here* — assign it.
- *this machine has no seat* — wrong machine; nothing can be installed to fix it.
### A capability may carry a value, and that is not a new idea
A capability is a named fact about a machine, **detected and never assumed**. Its presence gates
an assignment; its detail can also carry a value — `seat: card1-DP-1`, `panel: oled`, an
architecture, an amount of memory. Nothing new is needed for that: a verdict has always had a
detail beside its yes or no.
So *can this run here* and *what should it be configured as* are answered by the same fact, read
two ways. A module that must not be assigned without an OLED panel and one that dims itself
differently on one are reading the same line.
**What keeps the set from sprawling is the cost of adding one.** A capability must be detected,
and the detector must say how it knows — so nobody can add one they cannot check, which is the
whole of [04-ISSUES/007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md):
an installed package was treated as a capability and a node was assigned work it could not do.
**And detectors ship inside the host**, which is one statically linked binary. Adding a capability
means shipping a new host to every node that needs it. That is a real cost and it argues for
keeping the vocabulary small and general — `seat`, not `has-nvidia-with-two-outputs`.
## "Flavor" is retired
It was carrying three unrelated meanings — variants of a thing, a subset of one module a node
installs, and whatever the current system does, which earned two knowledge-base entries about
going wrong. **A word with three meanings cannot be reasoned about**, and every attempt to design
around it produced a rule that was right for one meaning and wrong for the others.
What it was reaching for is two ordinary things:
- **Different modules that provide the same thing.** `zsh` and `fish` both provide *a shell*. They
are two modules, not one module with a switch: they share a name and nothing else — different
packages, different configuration, different everything.
- **One module with a setting.** A monitoring module that is an agent here and a server there is
one module, configured. Nothing varies but a value.
If something is neither, it is probably two modules.
**A third thing it was reaching for, added 2026-08-29:** *I want this working and I do not care
which one.* That is a module with requirements and no files —
[a domain module](#what-a-domain-module-turns-out-to-be-and-why-it-is-not-the-one-refused-above) —
and it is what makes "different modules that provide the same thing" bearable for somebody who
does not want to know there is a choice.
## The core library is the mesh's domain
One module everything may depend on. It holds **what is true of the mesh regardless of which
context you are in**: a module, a node, an assignment.
The test: *would this still mean the same thing in a context that had never heard of the one it
came from?* A node would. A pipeline stage would not — that is delivery's.
**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types
depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub
where the relationship cannot be seen. **A library everything depends on is expensive to change
whether it holds types or code; the fan-in is what makes it expensive**, which is why *types, not
behaviour* was the wrong guard.
**It stays small on its own.** A domain model changes when what the mesh *is* changes, which is
rare. A drawer labelled *shared* changes whenever anybody writes something reusable, which is
constantly — and *who else might want this* always answers yes, which is how the current one grew.
## Consequences
- **Fewer things will be shared, and some code will be written twice.** That is the trade: the
current library exists because sharing felt free. Two similar functions in two modules is often
the better answer.
- **The check is a measurement rather than a prohibition.** Inbound build edges say when something
is becoming a hub, while it is happening rather than after.
- **Reading build edges needs a language-aware tool per language**, which is the real cost and the
reason declaring them looks tempting. It is still wrong.
+161
View File
@@ -0,0 +1,161 @@
---
topic: what runs on it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 10. Delivery
*Consolidated 2026-08-28 from five records.*
## Delivery is a comparison, not a pipeline
**The control plane holds what source exists and what has been built from it, and builds the
difference.** A change becomes a build because source is **ahead of artifacts** — answerable at
any moment — rather than because a message arrived.
**An event makes it fast. Nothing makes it necessary.** A missed notification costs latency and
cannot cost correctness.
That is the same shape the host uses on a machine, one layer up:
| | reconciles | against |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
**This is not the current coordinator repaired.** That is a state machine over stages; the value
of it here is as a catalogue of the ways this fails, and it has been used for exactly that.
### The module system is the CI/CD
*Written 2026-08-29, because this was the intention throughout and was never stated in one line.*
**There is no pipeline product beside the mesh, and there is not going to be one.** A module
declares what it is ([ADR 0009](0009-modules-and-the-graph.md)); the control plane notices its
source is ahead of its artifacts and builds it; the graph says what else that invalidates; the
node that should run it is told. Build, test, publish and deploy are the same reconciliation seen
at four points, not four stages wired together.
**Which is why the module system is the core of the setup rather than one component of it.** Every
other layer is carried by it: the substrate is modules the bundle raises before there is a mesh,
the control plane is a module, and an application is a module with a different manifest. A thing
that cannot be expressed as a module cannot be delivered at all — that is a real constraint, and it
is the one keeping a second delivery mechanism from growing beside this one.
**What disappears is the pipeline as a state machine** — no stage list something can be omitted
from, which is how a verify stage was built and never scheduled, and no run to lose.
### Currency is the whole input closure
**An artifact is out of date when its source moved, or anything it was built against moved.** So a
shared library changing invalidates everything with a transitive build edge to it, in dependency
order, because a module cannot be built against a new library until it exists.
**The module graph is a prerequisite of this, not an enabler of it.** Without it there is no
rebuild set and no ordering, and this cannot be implemented.
## An artifact is build output, never a source tree
Compiled and bundled with its dependency graph inlined. **A deploy is extract-and-run and touches
no network.**
The consequence is the whole cost of the decision: **anything not in the build output does not
ship.** Every file kind had to be brought into that rule separately, and each was discovered by
something silently not happening after a deploy — migrations reading a source layout,
provisioning scripts reading a source layout, selection files never packaged at all.
## Three silos, and the third is not a stage
The cardinality observation holds and is what the split is for:
| silo | runs | ends with |
|---|---|---|
| **build** | once per module | a self-contained artifact |
| **publish** | once per module | that artifact addressable — an image by digest, a package in the mesh's repository |
| **deploy** | **once, not once per node** | the affected nodes' **declarations updated** |
**Deploy stops sending commands to nodes.** It changes what the control plane says each node
should be, which is one write. What happens on the machines is the host's ordinary reconcile.
**Why this fixes the failure class rather than patching it.** Every recorded fault shares one
shape: *the thing that reported success was not the thing that did the work.* A coordinator
dispatching a command can only report on dispatch. Under this the reporter **is** the applier —
which already refuses to record a resource until it read it back, and already fails the whole
apply on one failed step.
**The verify stage disappears as a stage**, which is the strongest evidence for the shape:
verification stops being a step that can be omitted from a list and becomes a property of applying
at all.
**There is no fan-out**, so the defect class that came from the build node having passed through
two silos while others had not cannot arise.
## A step that fails must fail the job
A step that fails and lets the job continue **reports success for work that did not happen**.
Absence of an error is not evidence of an effect.
This is the mesh's most consistent failure shape, and it is not incidental — it is what stage
reporting measured. Documented instances: a service reported started when the container command
merely returned; an image pull failure that did not fail the deploy; a package install that 404'd
from every mirror while the job went green; a node left on old code after a failed download with
a version marker that had already advanced.
## The verdict is tiered
An artifact may not be declared until something has judged it fit. **Two tiers, because one gate
would be both slow and unreliable:**
| | judged by | when |
|---|---|---|
| the module's own tests | the build | **always** — this is most of it |
| the lab | a raised scenario | when an assertion genuinely needs a mesh |
A lab scenario takes tens of seconds and can fail for reasons that have nothing to do with the
artifact, and a shared-library change produces a cascade of dozens. One expensive
non-deterministic gate fails in both directions: a flaky run marks a good artifact unfit, a lucky
one marks a bad artifact fit, and **neither failure looks like itself.**
**A run that failed environmentally is not a verdict.** A machine that would not boot says nothing
about the artifact, and recording it as *unfit* is the same untruth as recording a dispatch as a
deploy.
## What a result means
> **The declaration is updated, and here is which nodes have applied it.**
A pipeline does not wait for every node — one may be legitimately switched off for a week, and a
delivery mechanism that blocks on a sleeping laptop is one nobody will use.
```
delivered declaration updated for 5 nodes
applied 3 of 5
outstanding 2 — last seen 4 days ago, 20 minutes ago
```
**Outstanding is not failure**, and conflating them is how the old system produced a stall with no
error anywhere.
## What must exist first
1. **The module graph, with build edges.** No graph, no rebuild set and no ordering.
2. **A recorded input closure per artifact**, so currency is answerable without building.
3. **Something that notices a reconciler is not converging.** Below.
## Open, and the first is the real risk
- **A loop that will not converge is harder to debug than a job that failed.** A failed job stops
and names its step; a reconciler retries forever. Without something that notices *this has been
trying for an hour*, the failure is **silence** — the fault this removes, reintroduced in a new
place.
- **The run identity people use is lost.** *Did my change go out?* is answerable today by opening
a pipeline. Something must replace that or this is worse to live with, whatever its properties.
- **Does a fit artifact declare itself?** If it does, merging to main deploys to production —
which may be wanted and is far too large a property to acquire by omission.
- **Rebuild storms are mostly behaviourally empty.** Reproducible builds would stop a cascade at
the first module whose output did not move; without them one commit redeploys the fleet for no
change in behaviour.
- **Detection stays the fragile input for latency**, though no longer for correctness.
@@ -1,17 +1,18 @@
---
topic: building it
status: accepted
date: 2026-04-03
deciders: jochen
reconstructed: true
---
# 4. Managed files are generated onto nodes and never edited there
# 11. Managed files are generated onto nodes and never edited there
> Reconstructed after the fact from the evidence cited below.
## Context
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) put every binding in the mesh
[ADR 0006](0006-the-substrate-and-the-control-plane.md) put every binding in the mesh
database. But the things that consume those bindings — environment files, service
definitions, daemon configuration, firewall rules — are files on a node's disk, because that
is what the software reading them requires.
@@ -26,7 +27,7 @@ directions at once.
1. **Bidirectional sync** — a node's edits flow back to the database. Rejected, and removed.
Two writers and no arbiter: whichever synced last wins, and neither is authority.
2. **Files are authoritative; the database is a cache of them.** Rejected — it inverts
ADR 0003 and returns to state that cannot be reconciled across nodes.
ADR 0013 and returns to state that cannot be reconciled across nodes.
3. **Strictly one-directional: the database is written, files are generated.** Chosen.
## Decision
@@ -1,60 +0,0 @@
---
status: superseded
superseded-by: 02-DECISIONS/0018-the-mesh-creates-no-symlinks.md
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 11. The installer owns linking; nothing else creates a symlink
> Reconstructed after the fact from the evidence cited below. The incident that earned the rule
> predates the record, and its date is not established here.
## Context
A service's definition lives in the module catalogue; its runtime directory and persistent data
live outside it. The mesh connects the two by linking the definition into the runtime location
— deliberately, so that runtime state and source stay separate while the running service reads
a current definition.
A link is also the easiest thing in the world to create by hand while fixing something, and a
container engine resolves a bind mount through it. A hand-made link pointed a volume somewhere
it should not have, and **production data was lost**.
## Considered options
1. **Copy instead of linking.** Rejected. A copy goes stale silently, which trades data loss
for a service running a definition nobody can find.
2. **Allow links, document the hazard.** Rejected. The hazard is not knowable at the moment of
the mistake — the link looks right and the resolution happens inside the container engine.
3. **One component owns linking; everyone else is forbidden.** Chosen.
## Decision
The installer creates and repairs every link the mesh needs. It reconciles them: a missing
link is created, a stale one is repointed, and a real file found where a link belongs is
adopted into the node's override location and replaced.
**Nothing else creates a symlink** — not a hook, not a fix, not an agent, not a person
debugging. The prohibition is absolute because the judgement required to make a safe exception
is exactly the judgement that was not available at the moment it mattered.
## Consequences
- The class of failure is closed, at the cost of a rule that reads as arbitrary to anyone who
has not seen the incident. That is why it is recorded here rather than only asserted.
- Links become reconcilable state rather than incidental filesystem facts.
- The rule is stated for humans and agents and is enforced by convention, not mechanism. A
check does not exist.
- The rule as written governs the mechanism rather than removing it. A link made by the
installer resolves the same way as one made by hand, so the hazard is narrowed and not
closed. [ADR 0018](0018-the-mesh-creates-no-symlinks.md) proposes widening this to "nothing
links, the installer included"; until that is accepted, this record governs.
## References
- Recorded as a non-negotiable in the governed constitution page, §2: *"Symlinks to repos or
service directories have caused production data loss via Docker volume path resolution. The
installer handles all linking. Never create symlinks manually."*
- Knowledge base: `services` — the reconciliation behaviour, including adoption of real files.
@@ -1,16 +1,16 @@
---
topic: building it
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0011-the-installer-owns-linking.md
---
# 18. The mesh creates no symlinks — a derived file is a copy
# 12. The mesh creates no symlinks — a derived file is a copy
## Context
[ADR 0011](0011-the-installer-owns-linking.md) responded to production data loss — a hand-made
[ADR 0012](0012-the-mesh-creates-no-symlinks.md) responded to production data loss — a hand-made
link, resolved through a container engine's volume handling, pointing a mount somewhere it
should not have — by centralising linking in the installer and forbidding it everywhere else.
@@ -25,7 +25,7 @@ Two things have changed since, and together they remove the argument that kept i
silently while the catalogue moves on, so a link was the cheap way to guarantee the running
node reads a current definition. That argument assumes the node's copy is unmanaged.
**It is not.** [ADR 0004](0004-managed-files-are-generated-never-edited.md) established that
**It is not.** [ADR 0011](0011-managed-files-are-generated-never-edited.md) established that
everything on a node's disk is derived from the mesh and regenerated when its inputs change,
and the installer already **reconciles** links rather than assuming them — repointing stale
ones, adopting real files it finds where a link belongs. Reconciling content is the same
@@ -34,12 +34,12 @@ operation as reconciling a pointer, plus a comparison.
So the mesh already has the machinery that makes a copy safe, and is using a link to solve a
problem that machinery solves better. Worse, a link is conceptually the wrong shape: it makes
the node's runtime state a *pointer into source*, which is the one thing
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) and ADR 0004 exist to prevent.
[ADR 0006](0006-the-substrate-and-the-control-plane.md) and ADR 0011 exist to prevent.
State is derived onto nodes; it does not reach back.
## Considered options
1. **Keep ADR 0011 as the final position** — centralised linking, forbidden elsewhere.
1. **Keep ADR 0019 as the final position** — centralised linking, forbidden elsewhere.
Rejected as the status quo. It governs the mechanism rather than removing it, and the
failure it was written for remains reachable by any code path the installer trusts.
2. **Keep links but harden them** — canonicalise before mounting, refuse a link that escapes
@@ -54,12 +54,12 @@ State is derived onto nodes; it does not reach back.
**The mesh creates no symlinks.** A file a node needs is placed on that node as a real file,
derived from the mesh and reconciled by the installer like every other managed file
([ADR 0004](0004-managed-files-are-generated-never-edited.md)).
([ADR 0011](0011-managed-files-are-generated-never-edited.md)).
The prohibition in ADR 0011 stands and widens: it ceases to be "only the installer may link"
The prohibition in ADR 0019 stands and widens: it ceases to be "only the installer may link"
and becomes "nothing links, the installer included".
When this is accepted, ADR 0011 becomes superseded rather than edited — its reasoning is why
When this is accepted, ADR 0019 becomes superseded rather than edited — its reasoning is why
the rule exists at all, and the incident behind it is the reason anyone believes either record.
## Consequences
@@ -73,7 +73,7 @@ the rule exists at all, and the incident behind it is the reason anyone believes
cost, and it is the whole cost: today a link cannot be stale, and a copy can. The answer has
to be detection — the installer comparing what is on disk against what the mesh says should
be — and it must be loud, because a silently stale definition is exactly the failure shape
this mesh keeps producing ([ADR 0008](0008-a-failed-step-fails-the-job.md)).
this mesh keeps producing ([ADR 0010](0010-delivery.md)).
- Reconciliation gets more expensive: comparing content rather than checking a pointer's
target, on every module, on every node.
- Disk usage rises, trivially, and is not a consideration.
@@ -90,12 +90,12 @@ the rule exists at all, and the incident behind it is the reason anyone believes
- **Migration order.** Converting a node's links is a change to how its services resolve their
own definitions, which is not a change to make everywhere at once.
Until those are answered this record stays `proposed`, and ADR 0011 remains the governing rule.
Until those are answered this record stays `proposed`, and ADR 0019 remains the governing rule.
## References
- [ADR 0011](0011-the-installer-owns-linking.md) — the incident, and the rule this widens.
- [ADR 0004](0004-managed-files-are-generated-never-edited.md) — the machinery that makes a
- [ADR 0012](0012-the-mesh-creates-no-symlinks.md) — the incident, and the rule this widens.
- [ADR 0011](0011-managed-files-are-generated-never-edited.md) — the machinery that makes a
copy safe.
- [`03-DESIGN/00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
— what the installer does today, including reconciliation and adoption.
@@ -1,63 +0,0 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 13. An artifact is build output, never a source tree
> Reconstructed after the fact from the evidence cited below.
## Context
A module is built once and deployed to every node assigned to it. What travels between those
two events is the artifact.
For a long time the artifact was a filtered copy of the module's source directory. Deploying it
therefore meant resolving and installing its dependencies **on the target node** — which
requires the target to reach a package registry, at deploy time, for every node, every deploy.
A node with no route to the registry could not deploy code that had already been built
successfully.
## Considered options
1. **Ship source, install dependencies on the target.** Rejected — it is what existed. Deploy
becomes a network operation with a failure mode per node, and the code that runs is
assembled independently on each one.
2. **Ship source plus its resolved dependency tree.** Rejected: large, slow, and it ships the
dependency resolution's platform assumptions along with it.
3. **Ship a self-contained build output; a failed bundle fails the build.** Chosen.
## Decision
The artifact is the module's **build output directory** — compiled and bundled, with its
dependency graph inlined. Deploy is extract-and-run and touches no network.
A build that cannot produce a self-contained output **fails**. It does not fall back to
shipping a dependency tree, because a fallback that works is a fallback that is never fixed —
an application of [ADR 0008](0008-a-failed-step-fails-the-job.md).
## Consequences
- A node can deploy without reaching a registry. What was built is what runs, identically, on
every node.
- Deploys are faster and their failure modes are local.
- **Everything not in the build output does not ship.** This is the decision's whole cost, and
it was paid several times before it was understood: migrations that read the source layout,
provisioning scripts that read the source layout, selection files never packaged at all. Each
worked in development, where the source is present, and silently did nothing after deploy.
- Any file a module needs at runtime must be deliberately placed into the build output. The
rule "the artifact is `dist/`" has to be applied to every file kind, not just compiled code,
and that generalisation was the expensive part.
- Bundling has its own failure modes that a compiler will not catch — a bundler can exit
successfully and produce output that cannot load.
## References
- `build: bundle artifacts so a deploy is extract-and-run` (#673), 2026-08-04.
- The consequences, in order: `Provision migrations and seeds read the source layout, not the
artifact` (#699), `Local migrations read the source layout too` (#700), both 2026-08-07.
- Knowledge base: `pipeline/artifacts-are-build-output`, `pipeline/bundling`,
`troubleshooting/shell-migrations-never-packaged`, `troubleshooting/flavors-never-packaged`,
`troubleshooting/esbuild-silent-tla-breakage`.
@@ -1,11 +1,12 @@
---
topic: building it
status: accepted
date: 2026-05-14
deciders: jochen
reconstructed: true
---
# 6. Schema and state changes are numbered migrations, in the same language as the code
# 13. Schema and state changes are numbered migrations, in the same language as the code
> Reconstructed after the fact from the evidence cited below.
@@ -1,78 +0,0 @@
---
status: accepted
date: 2026-08-04
deciders: jochen
reconstructed: true
---
# 14. Build, publish and deploy are three silos with different cardinality
> Reconstructed after the fact from the evidence cited below.
## Context
Delivery had been treated as one pipeline that a module passes through. It is not: its stages
run a different number of times.
- Compiling happens **once per module feature**, on the build node.
- Packaging and uploading happens **once per module feature**, on the build node.
- Installing, configuring, starting and verifying happens **once per module feature per node**.
Conflating them is what made earlier versions slow and hard to reason about. Work that should
happen once was being repeated per node, and the fan-out point was implicit rather than a
boundary anything could observe.
The split had been declared before it was real. Packaging still happened inside the build,
which meant the boundary existed in the documentation and not in the code.
## Considered options
1. **One pipeline, stages that know their own cardinality.** Rejected — it is what existed.
Cardinality is then a property of each stage's implementation, and nothing can reason about
the pipeline as a whole.
2. **Two silos: build-and-publish, then deploy.** Rejected. It leaves packaging inside build,
so build must know every module, every feature, and how each composes its artifact —
exactly the coupling the split exists to remove. A failed upload then retries by re-sending
a stale package instead of re-packaging.
3. **Three silos, with an explicit handover between each.** Chosen.
## Decision
Delivery is three silos, and the boundaries are real:
| Silo | Runs | Where |
|---|---|---|
| **build** | once per module feature | the build node |
| **publish** | once per module feature | the build node |
| **deploy** | once per module feature **per node** | every assigned node |
Commands and events are addressed **per feature**, not per module.
Build compiles and hands over a **staged tree** — not a package. Publish applies the module's
packaging rules, packages that tree, and uploads it. Publishing to a package registry *is*
publishing, so a module whose artifact is a package publishes in the publish silo, not the
build one.
Modules are resolved into dependency **levels**, and a level completes before the next begins,
so a module always builds against its dependencies' freshly published versions.
## Consequences
- Work that should happen once happens once. The fan-out point is explicit and observable.
- A failed upload retries by re-packaging, because packaging belongs to the stage that
uploads.
- The handover is a staged tree in a known location rather than the build's working directory,
which is reference-counted and cannot be assumed to still exist when a later stage runs.
- The build node is now the only node that has already passed through two silos when the
fan-out happens. Anything tracking a node's stage must account for **both** pre-fan-out
stages; code that knew only about the first parked the build node forever while every other
node deployed cleanly.
- A recovery mechanism that knows a subset of the stages it guards is worse than none — it
reports success over a stall it cannot see.
## References
- `publish owns packaging — the silos were not actually split` (#677), 2026-08-04.
- Knowledge base: `pipeline/three-silos` — including the note that the older architecture
documents claimed otherwise and were stale until 2026-08-06.
- The build-node stage-tracking failure was observed on pipeline #5557.
@@ -1,11 +1,12 @@
---
topic: building it
status: accepted
date: 2026-06-04
deciders: jochen
reconstructed: true
---
# 7. No workspace — each module is a standalone package consuming published dependencies
# 14. No workspace — each module is a standalone package consuming published dependencies
> Reconstructed after the fact from the evidence cited below.
@@ -47,7 +48,7 @@ its dependencies' freshly published versions.
- Development and the pipeline resolve imports identically. The divergence is gone by
construction rather than by discipline.
- A module in its own repository is not a special case. It builds exactly as a module in the
monorepo does — which is what makes [ADR 0010](0010-applications-live-in-their-own-repository.md)
monorepo does — which is what makes [ADR 0015](0015-applications-live-in-their-own-repository.md)
cheap.
- A cross-package change costs a publish-and-consume round trip. This is the real price, paid
on every shared-library change.
@@ -1,11 +1,12 @@
---
topic: building it
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 10. Applications live in their own repository; the monorepo is for the mesh
# 15. Applications live in their own repository; the monorepo is for the mesh
> Reconstructed after the fact from the evidence cited below.
@@ -46,7 +47,7 @@ reject it.
- An application's cadence is its own. It is not reviewed as mesh code and does not queue
behind mesh work.
- The separation is safe **only because** the pipeline and provisioning are identical either
side of it — which [ADR 0007](0007-no-npm-workspace.md) is what makes true. Without
side of it — which [ADR 0014](0014-no-npm-workspace.md) is what makes true. Without
standalone packages this decision would fork the build.
- The monorepo stops being an inventory of the installation, which is a precondition for
publishing anything about it.
@@ -59,6 +60,6 @@ reject it.
- The rule is stated in the governed constitution page authored 2026-07-10, §3, as a
convention violation reviewers must reject.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) decision 4 extends this from *new*
applications to the modules already in the monorepo.
- Knowledge base: `troubleshooting/unregistered-module-source`.
@@ -1,138 +0,0 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
---
# 16. A lab node is a virtual machine running the real install
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) makes a local mesh a prerequisite
rather than a convenience: *"everything that manifests between nodes is discoverable only in
production, which is where every fault of 2026-08-22 was found."*
Two things were measured while establishing what exists
([`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md)):
- **There is no local mesh.** The dev tooling starts providers through the host's own init
system and reads credentials from host paths (`modules/hal/developer/tools/dev-env.ts:146-178`).
It borrows the machine because there is nowhere else to put a mesh.
- **The one containerised node in the repository has been unable to build since 2026-06-04**,
when the npm workspace it depends on was removed. Nothing runs it, so nothing reported it.
So the question is not how to improve a local mesh. It is what a node *is* when it is not a
physical machine. Every subsequent question — how faithful is faithful enough, what may be
mocked, which failures remain reachable — follows from that one answer.
The hardware available is not a constraint: 125 GB of memory with 71 free, 24 threads, and
hardware virtualisation present.
## Considered Options
1. **An application container.** Rejected. **A node's job is to run containers**, so modelling
a node as one inverts the thing being modelled: module service stacks then require nested
containers through a privileged daemon, or a shared socket that makes isolation between
nodes cosmetic. Init is not PID 1, so units and timers need workarounds. Cheapest to start
and the least like a node.
2. **A system container.** Rejected, after first being recommended. It is genuinely good —
real init, properly nested containers, roughly a second to boot, cheap snapshots — and it
is the only option that makes a twenty-node run affordable. It was rejected because **the
scale requirement that justified it was invented rather than required**: the stated goal is
to run the real mesh, which is four nodes, on one computer. And a system container still
forces the question a virtual machine dissolves — *how faithful must a node be?* — which
then has to be answered again for every capability under test.
3. **`systemd-nspawn`.** Rejected. Already present, so nothing to install, but too primitive:
no storage pools, no snapshot management, no network management, no virtual machines.
Snapshots are what make the loop fast, so the saving is not worth what it costs.
4. **A virtual machine.** **Adopted.** A bare Arch Linux machine that the real install script
turns into a node.
## Decision
**A node in the mesh development lab is a virtual machine.** It boots a stock Linux image,
runs the real install, and becomes a node. It is not a model of a node, so no question arises
about how good the model is.
The environment is called **the lab**.
Three things follow directly and are decided here:
### The lab is driven by `incus`
Chosen for what it manages, not for what it is: virtual machines, their snapshots, and the
bridges between them, through one interface. It also manages system containers, so if a run
ever genuinely needs twenty nodes, that is a change of instance type rather than a rewrite.
Declared in `modules/hal/developer/module.yml`, so it installs the way every other package
does.
### The simulated public segment uses TEST-NET-3
`203.0.113.0/24`, reserved by RFC 5737, never routable.
This is not cosmetic. WireGuard decides per pair whether to write an `Endpoint` by testing the
peer's underlay address against an RFC1918 regex
(`modules/wireguard/hooks/index.ts:225-240`). A simulated public segment addressed from
private space makes the hub test as unreachable, so no spoke writes an endpoint for it,
nothing can initiate, **and the mesh silently never forms** — appearing as a WireGuard fault
rather than an addressing mistake.
The production LAN subnet and the entire overlay address plan are reproduced unchanged.
### The lab issues its own certificates
Public names are certified by an ACME server inside the lab; `.internal` names keep the mesh
CA. **The lab keeps production's two-authority split rather than collapsing it**, because a
single-authority lab would hide any fault living in that split.
This also makes the lab's port forward load-bearing: an HTTP-01 challenge must reach a
published-but-NATed node on port 80, so a broken forward becomes a reproducible certificate
failure rather than a mystery.
## Consequences
**The fidelity question disappears, and with it a class of argument.** There is no "how real
is this node" to litigate per capability, because the node is real. What remains not-real is a
short, enumerable list: the model provider, the public internet, and the public certificate
authority.
**The install becomes the thing under test.** A container-shaped lab would have had to skip
the bootstrap entirely. Here it runs, so it is exercised on every fresh lab.
**Reproducing the network is mostly a data problem.** The bootstrap performs no network
configuration at all; WireGuard, DNS, routing and internal TLS are generated by module hooks
from mesh-DB rows. The lab therefore exercises the same code production runs rather than a
reimplementation ([`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md)).
**Scale runs get expensive, and this is the real cost.** Four virtual machines are
comfortable; twenty are not, on a workstation. Faults that only appear at scale — a fan-out
reaching most consumers rather than all, a cascade that stalls with many modules — stay hard
to reproduce. The mitigation is that the same tooling runs system containers, so a scale run
remains possible at lower fidelity if one is ever genuinely needed.
**Boot is slower, and it does not matter.** Ten to twenty seconds against roughly one. A run
includes a full delivery — build, publish, install, migrate — measured in minutes, so boot
time is noise.
**One change is required before the lab can issue certificates.** The reverse proxy sets no
`caServer`, so it defaults to the public authority's *production* endpoint
(`modules/traefik/docker-compose.yml:17-19`). It must become configurable, defaulting to
production so real nodes are unaffected. Worth noting on its own: aiming at production rather
than staging means every certificate experiment on a real node consumes issuance quota.
## References
- [`01-RESEARCH/002-local-mesh`](../01-RESEARCH/002-local-mesh/analysis.md) — what exists, and
the four host couplings that only obstruct a container-shaped node
- [`01-RESEARCH/004-lab-network`](../01-RESEARCH/004-lab-network/analysis.md) — the topology
being reproduced and the endpoint constraint
- [`03-DESIGN/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) — what the lab
is for
- `modules/wireguard/hooks/index.ts:206-240` — the endpoint rule, and the incident comments
recording what it cost to get right
- RFC 5737 — reserved documentation address blocks
+84
View File
@@ -0,0 +1,84 @@
---
topic: building it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 16. The lab
*Consolidated 2026-08-28 from five records. The lab is one design and was split across five
decisions taken over three days; the reasoning is kept, the fragmentation is not.*
The environment a change is run against before it reaches real machines.
## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
a node**, so no question arises about how good the model is — which is the whole reason for
paying the cost of virtual machines rather than containers.
The lab is driven by **incus**, and a scenario is raised from a declaration.
## A router is scenery, and is therefore a container
**Nothing under test runs on a router.** It is not a participant, holds no identity, has nothing
installed on it by the mesh, and no assertion is ever made about its internals. It exists so that
packets between machines behave the way they behave in the world.
The fidelity argument that makes a node a virtual machine does not reach it: what a router *is*
does not matter, only what it *does to traffic*. So a router is a system container, and the lab
is cheaper for it.
## A scenario declares the underlay, and only the underlay
**What a hosting provider and a home router would have provided**, before any of our software
touched the machine:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and may be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering, no
names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
> A scenario provides what a hosting provider and a home router would provide, and nothing our
> software is responsible for.
## A scenario is a closed address space
Every segment materialises as its own isolated link belonging to one scenario instance. **Two
scenarios raised from the same declaration hold the same addresses and never meet**, because
nothing joins their links. The declaration therefore keeps its literal addresses and they mean
exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the way one would
use a console rather than the network. That is what makes two identical scenarios able to run at
once, and it is why placing anything inside a machine is a hypervisor operation rather than a
network one.
## Two scenario classes, and the first has no pipeline
| | **bootstrap** | **full** |
|---|---|---|
| contains | machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| verdict from | what the host reports about the state it reconciled | a delivery result ending in verification |
| exercises | tiers 0 and 1 | tiers 2 and above, and modules |
**The bootstrap class comes first**, because it is what develops the node host, and because a
full scenario needs tiers that do not exist yet. A lab that could only raise the larger class
would be a lab nobody could use until everything else was built.
## Consequences
- **The lab tests the real code path**, not a reimplementation of it. The network a scenario
produces is generated by the same code production runs.
- **Isolation is what makes it usable in parallel**, and it costs the ability to reach in over
IP. Everything the lab puts inside a machine — a binary, an image, a file — goes through the
hypervisor.
- **A sealed scenario cannot fetch anything**, which is a real limit rather than an inconvenience:
it is why images have to be placed and why a container runtime has to be in the base image.
@@ -1,11 +1,12 @@
---
topic: checking it
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
---
# 34. A test defends a decision
# 17. A test defends a decision
## Context
@@ -1,97 +0,0 @@
---
status: proposed
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0015-mesh-brokers-nodes-host-agents-think.md
---
# 17. Modules outside the platform core are grouped by domain, not by single function
## Context
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) recomposes the platform's own modules
into bounded contexts named after their aggregates, and sends the rest out of the monorepo on
the grounds that they run *on* the mesh rather than being *of* it.
That leaves the larger half unaddressed. Around three quarters of the catalogue are modules
that are neither part of the mesh's domain nor standalone applications: a firewall, a VPN, an
SSH daemon and a resolver; a file manager, a media player and a system monitor; a set of
media-library services. Today each is its own module, because one module is the unit of *one
piece of software*, and no other grouping exists.
The result is that the catalogue's shape records what was installed, not what anything is for.
Four modules that together constitute "how a node is reachable" have no relationship the mesh
can see: they cannot be assigned, versioned, reasoned about or replaced as one thing, and a
change to how the mesh handles connectivity has to be made four times.
This is the same failure ADR 0015 names for the core — *boundaries drawn by deployment accident
rather than by domain* — appearing outside it.
## Considered options
1. **Leave them as they are.** Rejected. The core gets domain boundaries and everything else
keeps accident boundaries, so the catalogue becomes harder to read after the refactor than
before it.
2. **One module per piece of software, with a tag or category field.** Rejected. A label is not
a boundary: it does not change what can be assigned, versioned or replaced as a unit, and it
drifts from the thing it labels.
3. **Group them into domain modules, each owning the software that serves one purpose.**
Proposed here.
4. **Extend ADR 0015's contexts to cover everything.** Rejected. Those contexts are named for
the mesh's own aggregates; a media library is not an aggregate of the mesh, and forcing it
into that model repeats the metaphor-naming mistake ADR 0015 exists to correct.
## Decision
*Proposed — the principle is settled; the domain list is not. See "Open" below.*
Modules that are not part of the platform core are grouped into **domain modules**. A domain
is named for the concern it serves, and owns the software that serves it. The unit stops being
one piece of software and becomes one purpose.
This extends ADR 0015 rather than replacing it. The eight bounded contexts for the mesh's own
domain stand unchanged. This decision covers what ADR 0015 leaves outside them.
Naming follows the same rule as the core: **name the domain for what it does, not for what it
is made of**. Connectivity, not a VPN implementation.
## Consequences
- A domain becomes assignable, versionable and replaceable as one thing. Changing how nodes
reach each other is a change to one module.
- The catalogue's shape starts describing purpose. A reader can tell what a mesh is *for* from
its module list.
- Swapping an implementation stops being a module replacement, with the data-volume and
provisioning consequences that carries, and becomes a change inside a domain.
- The count drops sharply, which is a symptom of the improvement rather than the point of it.
- **Grouping conceals.** A domain module hides which implementation is in use, and every
operational question — which port, which unit, which credential — gains an indirection.
- The migration is not free and has no obvious increments: a domain is only useful once
everything belonging to it has moved.
- Some modules genuinely serve one purpose and are already correctly sized. Grouping for its
own sake would be the same error in the other direction.
## Open
**The domain list is not settled and this record does not invent one.** What is decided is the
principle; what is not decided is the set. Candidate groupings are visible in the catalogue —
connectivity and reachability, node presentation and desktop, media libraries, observation and
metrics, storage and data services — but naming them here would be reconstructing a decision
that has not been taken.
Settling the list is a research effort, not an act of this record. Until it concludes, this
ADR stays `proposed`. That effort is
[`01-RESEARCH/005-domain-grouping`](../01-RESEARCH/005-domain-grouping/00-overview.md), and its
first measurement already narrows this record's scope: co-change analysis supports grouping for
reachability, argues against it for the provisioned infrastructure providers, and finds no
signal either way for the fifty modules that never change alongside anything.
## References
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — the core decomposition this
extends, and its rule about naming a context after its aggregate.
- [ADR 0010](0010-applications-live-in-their-own-repository.md) — standalone applications are
already out of scope here; they are not domains and do not group.
- [`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md)
— the catalogue's current shape, which is the evidence for the problem.
@@ -1,18 +1,19 @@
---
topic: checking it
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0034-a-test-defends-a-decision.md
extends: 0017-a-test-defends-a-decision.md
---
# 35. A picture of a system is read from the system, never from what asked for it
# 18. A picture of a system is read from the system, never from what asked for it
## Context
A scenario declaration is a file. A raised scenario is a set of machines, links and rulesets.
The two are supposed to correspond, and the entire value of the lab rests on noticing when
they do not — [ADR 0034](0034-a-test-defends-a-decision.md) says a claim nothing checks is a
they do not — [ADR 0017](0017-a-test-defends-a-decision.md) says a claim nothing checks is a
claim that will quietly stop being true.
Drawing a scenario makes that concrete, and forces a choice that looks cosmetic and is not.
@@ -91,8 +92,8 @@ difference read off directly.
## References
- [ADR 0034](0034-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — why the lab must not supply what the
- [ADR 0017](0017-a-test-defends-a-decision.md) — a claim nothing checks stops being true.
- [ADR 0016](0016-the-lab.md) — why the lab must not supply what the
mesh is responsible for; the same instinct, applied to facts rather than to configuration.
- [04-ISSUES/003](../04-ISSUES/003-firewall-scope-is-read-by-no-code/00-report.md) — the fault in production
form.
@@ -0,0 +1,132 @@
---
topic: how we work
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 19. How this repository works
*Consolidated 2026-08-28 from ten records that were one decision seen from ten angles. The
reasoning is kept; the fragmentation is not.*
## The repository
**`novox/hq` is Novox's headquarters, and it is public.**
Company-scoped, not the mesh's. Today almost everything in it is about the mesh, because the
mesh is what Novox is building — a fact about the present rather than a definition. A second
product would live here too.
**Public** means written for a reader who is not its author and has no access to the mesh it
describes. Nothing here may contain routable addresses, real domain names, node names, absolute
paths, usernames or credentials. The test: *would this paragraph still teach a stranger running
an entirely different mesh?*
**Separate from the code** because the cadence differs — a decision changes when thinking
changes, not when code changes — and because a public repository cannot be a private one's
subdirectory.
**The naming rule:** a repository belonging to a product carries that product's prefix. A
company-scoped one does not. So this is `hq` and the mesh's are `mesh-*`.
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the pinned tier-1 services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | the mesh's own domain ([ADR 0009](0009-modules-and-the-graph.md)) |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this one |
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths. **`Nox` is an identity of Novox**, an agent
participant within the mesh's own model, not a second system.
## The folders, and why they are numbered
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
Following the numbers walks the process in the order it happens.
| | |
|---|---|
| `01-RESEARCH` | an open question, while it is open |
| `02-DECISIONS` | what was decided, and why |
| `03-DESIGN` | what is being built |
| `04-ISSUES` | something wrong at the level of design or governance |
**`03-DESIGN` has two layers and they are never mixed.** `00-as-is/` describes the mesh that
exists, written from the implementation and the operational record. `01-to-be/` describes the one
being built toward. Every document says which it is. A statement about the future does not belong
in an as-is document, and an as-is document is never edited to describe an intention.
**`04-ISSUES` is for design-level faults** — a rule enforced by nothing, a stated invariant that
is false, a failure the design permits to be silent. Not an operational ticket queue.
## What a decision record is, and is not
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is not
recorded.**
That bar has been read too generously. A *finding* is not a decision. A bug is not a decision.
**A record is warranted when there is a genuine fork**: a direction reversed, an alternative
seriously considered and likely to be proposed again, or something contested that needs to stay
settled. Everything else belongs in the design document, where the reasoning is read.
**There is no ledger** — no separate document summarising, ranking or tracking decisions. A
chronological view is generated from frontmatter, which is what a ledger was actually for.
**A number identifies a record and never changes.** It is not a position, and it cannot be
both — a position moves when the set changes, and an identity that moves is not one.
That is not a preference. Records are referenced from **outside** this repository: code
comments, commit messages, the knowledge base. Renumbering once cost 96 references across two
code repositories, and nothing in either would have failed to compile — the comments would
simply have pointed at the wrong reasoning, which is worse than a broken link because nothing
reports it.
**So the reading order lives in a generated index**, from each record's `topic:` — what the mesh
is, then its tiers from the bottom up, then what runs on them and how it gets there, then how it
is built, how it is checked, and how we work.
**And the index is written, not only generated on demand.** A reader looking at the folder on a
forge sees the folder, not a command. The objection to a written index is that it drifts, and
that is answered by **checking** it rather than by refusing to write one — which is §5's own
rule: a rule states how it is checked. A record with no topic, or a topic nobody defined, fails
the same check, because the quiet failure is a record that vanishes from the order rather than
appearing in the wrong place.
**The design layer is what you read.** These records explain *why* a thing is as it is. They are
not a description of the system, and needing to read them to understand it would mean the design
documents had failed.
## Status, and views over it
**Every document carries its state in YAML frontmatter** — research overviews, design documents,
decision records, issue reports.
**There are no central status files.** Every cross-cutting view — a status matrix, a decision
index, an open-issue list — is generated from frontmatter when asked for, never written to disk.
Two places holding one fact drift, and the written one wins by being closer to hand.
**Prose does not restate status.** One place, and two is one too many.
## Workflows are playbooks
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs it,
steps, outputs. People and agents follow the same ones, and **agents do not act outside them**.
Each is wrapped by a thin skill that defers to the playbook as authoritative and adds only the
mechanical scaffolding — so the process has one definition rather than a document and an
implementation that disagree.
## Consequences
- **A reader has one place per thing.** The design layer describes the system; these records say
why; the playbooks say how work is done.
- **Records will accumulate more slowly**, because the bar is a fork rather than a finding. This
record is itself the correction: ten records became one because they were one decision.
- **The public rule constrains everything written here**, permanently and at every commit. It is
the reason research describes real observations without identifying the mesh it observed.
@@ -1,69 +0,0 @@
---
status: accepted
date: 2026-08-22
deciders: jochen
reconstructed: false
supersedes: none
---
# 19. HQ is its own repository, and it is public
## Context
The mesh's reasoning — mission, research, design, decisions — began inside the code
repository, under a folder there. The objection to moving it out was specific and good: the
mesh already has an operational memory and a structured archive, and adding a third store
repeats the mistake that consolidation was meant to fix.
## Considered options
1. **Keep it in the code repository.** Rejected, but the objection it rests on is correct and
is answered rather than dismissed — see Consequences.
2. **Put it in the structured archive**, alongside the governed documents. Rejected: the
archive is not reviewable as a diff, and a design argument is exactly the thing that needs
line-by-line review and a branch.
3. **Its own repository.** Chosen.
## Decision
HQ is its own repository, and it is **public** — written for a reader who is not its author
and has no access to the mesh it describes.
Three reasons it is separate:
- **The cadence differs.** A decision changes when thinking changes, not when code changes.
Tying documents to a code branch merges them on the code's schedule.
- **The reviewers differ.** A design argument is not reviewed the way an implementation is,
and should not queue behind a build.
- **The scope is wider than one repository.**
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) sends most modules out of the
monorepo; documentation governing several repositories cannot live inside one of them.
Being public is not incidental. It is enforceable only because
[ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) made the code repository
node-agnostic: there is no per-node content to leak. Nothing here may carry routable
addresses, real domain names, hosting providers, node names, absolute paths, usernames,
credentials, or operational detail useful only to an attacker.
The test is whether a paragraph would still teach a stranger running an entirely different
mesh.
## Consequences
- A document and the code it describes can no longer land in one commit. Keeping them honest
is a discipline rather than a mechanism — which is why decisions are recorded as they are
taken, and why a document stating a rule must say how the rule is checked.
- Research must state evidence without identifying the mesh it observed. The shape of a
finding survives anonymisation; the instance does not travel.
- **The objection is answered by indexing, not by location** — the claim being that these
documents remain searchable beside everything else, one source with many surfaces.
**That indexing does not exist.** Checked 2026-08-23, it returns nothing. Until it does, the
objection stands unanswered and this repository is the third knowledge store it was argued
not to be. Recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
## References
- Supersedes the earlier position that documentation lives inside the code repository under a
folder there. That position was never recorded separately and has no record of its own.
- [`README.md`](../README.md) — the public-repository rule in full.
@@ -1,67 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 20. Design is written in two layers: what is, and what is intended
## Context
HQ held only the intended mesh. Every reader had to already know the running system in order
to understand what the decisions were about, and a statement about current behaviour had
nowhere to live except inside a document describing an intention.
The consequence was invisible until looked for: an as-is claim inside a to-be document is
indistinguishable from the intention around it, so the document silently stops being true as
the system moves — and nobody can tell which half went stale.
The mesh also has a large body of shipped behaviour that nobody would choose again. It is not
design in the sense of "what we decided"; it is design in the sense of "what is there", and it
is exactly the part a person changing the system most needs.
## Considered options
1. **One layer, describing the target.** Rejected — the status quo. The running system goes
undocumented and the target document accumulates unmarked claims about it.
2. **One layer, describing what runs, with intentions only in decision records.** Rejected:
a decision record is an argument, not a specification, and a multi-part intention has
nowhere coherent to live.
3. **Two layers, declared per document, never mixed.** Chosen.
## Decision
`03-DESIGN` holds two layers, and every document declares which it is:
| Layer | Describes | Written from |
|---|---|---|
| `00-as-is/` | The mesh that exists | The implementation and the operational record |
| `01-to-be/` | The mesh being built toward | Decision records |
An as-is document **records what is, not what should be** — including behaviour nobody would
choose again. A layer that keeps only the good decisions is a brochure.
When a to-be design ships it **does not move**. Its as-is counterpart is written or updated,
the to-be document's status becomes `implemented`, and both stand: one describing what runs,
the other recording what was intended. Deleting the intention loses the reasoning.
Where implementation and intention disagree, the as-is document records the implementation and
says they disagree.
## Consequences
- A reader can tell, from the folder and from one frontmatter field, whether they are reading
a description or a plan. That distinction was previously unavailable at any price.
- Correcting an as-is document requires evidence from the implementation, not agreement — and
needs no decision record, because nothing was decided.
- Two documents must be kept current per subsystem instead of one. This is the cost, and it is
paid on every ship.
- Something that shipped differently from its design becomes a visible divergence rather than
a silently wrong document, and may deserve an issue.
## References
- [`03-DESIGN/README.md`](../03-DESIGN/README.md) — the layer contract and frontmatter schema.
- [`03-DESIGN/00-as-is/`](../03-DESIGN/00-as-is/) — the first eleven as-is documents, written
2026-08-23 from the monorepo and the operational memory.
@@ -1,11 +1,12 @@
---
topic: how we work
status: accepted
date: 2026-07-10
deciders: jochen
reconstructed: true
---
# 9. The mesh is governed by a constitution, injected where work is decided
# 20. The mesh is governed by a constitution, injected where work is decided
> Reconstructed after the fact from the evidence cited below.
@@ -1,15 +1,16 @@
---
topic: how we work
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 25. HQ is the source of the mesh constitution
# 21. HQ is the source of the mesh constitution
## Context
[ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set,
[ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) established a canonical rule set,
injected into every eligible design session and checked before output is accepted. It lives in
the knowledge base, where the orchestrator reads it.
@@ -39,7 +40,7 @@ directly.
Publishing is a playbook step, not a manual act, and it ends with **reading the page back and
verifying the change is present**. A publish that reported success and did nothing is exactly
the failure class this mesh keeps producing
([ADR 0008](0008-a-failed-step-fails-the-job.md)).
([ADR 0010](0010-delivery.md)).
Section numbering is stable, because the orchestrator and the review fragments cite sections by
number.
@@ -47,7 +48,7 @@ number.
## Consequences
- One source, many surfaces — the same argument HQ's separation already rests on
([ADR 0019](0019-hq-is-its-own-repository.md)), applied to the rules themselves.
([ADR 0019](0019-how-this-repository-works.md)), applied to the rules themselves.
- Each rule keeps the incident that earned it, in a place that is reviewed as a diff.
- An edit to the derived page survives until the next sync and then vanishes. The playbook says
so, and nothing mechanically prevents it.
@@ -63,5 +64,5 @@ number.
- [`00-META/process/05-constitution-sync.md`](../00-META/process/05-constitution-sync.md) —
the sync, including the read-back.
- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it
- [ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) — the governed page and why it
exists.
@@ -1,57 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 21. Every workflow is a playbook, and agents operate through them
## Context
HQ stated a knowledge flow — research becomes design — and nowhere stated how anything moves
along it. What graduation required, who wrote the decision, what closed an effort, what
happened when something shipped: all of it was convention held in one person's head.
A large share of the work here is done by agents. An unwritten convention is not available to
an agent at all, so each one either invents a procedure or asks. Both produce a repository
whose shape depends on who last touched it.
## Considered options
1. **Convention, learned by reading existing documents.** Rejected — the status quo. It
transmits shape but not rules, and it transmits the mistakes along with the pattern.
2. **One long contributing document.** Rejected: it is read once, and the step someone needs
is never the step they are reading.
3. **A playbook per workflow, each with trigger, steps and outputs, wrapped by a thin skill.**
Chosen.
## Decision
Every workflow is a playbook in [`00-META/process/`](../00-META/process/): trigger, who runs
it, steps, outputs. Engineers and agents follow the same playbooks, and **agents must not act
outside them**.
Each playbook is wrapped by a thin skill that defers to it as authoritative and adds only
mechanical scaffolding — next free number, frontmatter block, where the file goes. The
playbook holds the reasoning; the skill holds the steps. When they disagree, the playbook
wins.
## Consequences
- An agent arriving with no context can act correctly, because the procedure is retrievable
rather than remembered.
- The playbooks are themselves reviewable. A bad rule can be found and changed, which is not
true of a convention.
- Duplication between playbook and skill is real, and is managed by making the skill thin and
naming the playbook as authoritative in the skill's first lines. Nothing prevents them
drifting; the constraint is that only one carries reasoning.
- A workflow with no playbook is a workflow agents will get wrong. Adding one is part of
adding the workflow.
## References
- [`00-META/process/00-overview.md`](../00-META/process/00-overview.md) — the five playbooks
and the flow they implement.
- Modelled on the process layer in the sibling HQ repository for the PAPA platform, which
arrived at the same shape and the same thin-skill split.
@@ -1,54 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 22. Status lives in frontmatter; cross-cutting views are generated
## Context
Status was carried in prose — a bold line near the top of a document saying what state it was
in — and indexes were maintained by hand. The decision-record index had already drifted from
the folder it described **after a single addition**, which is about as short a demonstration
as the failure mode offers.
A hand-maintained index is a copy of something the filesystem already knows. It is correct
only for as long as everyone remembers it exists, and its being wrong is silent.
## Considered options
1. **Prose status plus hand-maintained indexes.** Rejected — the status quo, already
demonstrably broken.
2. **A central status file.** Rejected. It centralises the drift rather than removing it: the
file and the documents disagree, and the file is the one people read.
3. **Machine-readable frontmatter per document; every cross-cutting view generated on
demand.** Chosen.
## Decision
Every document carries its state in YAML frontmatter — research overviews, design documents,
decision records, issue reports — with a schema stated in the section README.
**There are no central status files.** Every cross-cutting view — a status matrix, the
decision-record index, the open-issue list — is generated from frontmatter when asked for, and
never written to disk.
Prose does not restate status. One place, and two is one too many.
## Consequences
- A view cannot drift from what it describes, because it does not persist.
- Status becomes queryable. Inconsistencies — a closed effort with nothing in `became:`, an
`implemented` design with no owning repository — are findable mechanically, and the
generator reports them as flags rather than silently rendering around them.
- Frontmatter must be valid and paths in it must resolve, which is now something to check.
- A reader browsing the repository on a forge sees no index. That is the trade: the index is
correct and absent rather than present and wrong.
## References
- [`.claude/skills/hq-status/SKILL.md`](../.claude/skills/hq-status/SKILL.md) — the
generator, including the inconsistencies it flags.
- [`02-DECISIONS/README.md`](README.md) — the hand-written index that drifted, and its removal.
@@ -1,15 +1,16 @@
---
topic: how we work
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
---
# 40. The constitution absorbs what is already enforced
# 22. The constitution absorbs what is already enforced
## Context
[ADR 0025](0025-hq-is-the-source-of-the-constitution.md) makes this repository the source
[ADR 0021](0021-hq-is-the-source-of-the-constitution.md) makes this repository the source
and the knowledge base a derived copy, and playbook
[05](../00-META/process/05-constitution-sync.md) publishes the copy whenever a rule changes.
@@ -97,8 +98,8 @@ trusting the second success message either.
## References
- [ADR 0025](0025-hq-is-the-source-of-the-constitution.md) — source and copy.
- [ADR 0009](0009-the-mesh-is-governed-by-a-constitution.md) — why the copy is injected at all.
- [ADR 0021](0021-hq-is-the-source-of-the-constitution.md) — source and copy.
- [ADR 0020](0020-the-mesh-is-governed-by-a-constitution.md) — why the copy is injected at all.
- [Playbook 05](../00-META/process/05-constitution-sync.md) — the sync this record interrupts.
- [ADR 0034](0034-a-test-defends-a-decision.md), [ADR 0035](0035-a-picture-is-read-from-what-runs.md),
[ADR 0018](0018-the-mesh-creates-no-symlinks.md) — the three rules whose sync surfaced this.
- [ADR 0017](0017-a-test-defends-a-decision.md), [ADR 0018](0018-a-picture-is-read-from-what-runs.md),
[ADR 0012](0012-the-mesh-creates-no-symlinks.md) — the three rules whose sync surfaced this.
@@ -1,11 +1,12 @@
---
topic: how we work
status: accepted
date: 2026-08-26
deciders: jochen
reconstructed: false
---
# 42. The approval is the checkpoint, not the second pair of hands
# 23. The approval is the checkpoint, not the second pair of hands
## Context
@@ -73,5 +74,5 @@ the previous sync reported success and changed nothing.
## References
- [`how-we-build.md`](../00-META/how-we-build.md) §2 — the rule, now carrying this.
- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard a checkpoint is held to: a
- [ADR 0010](0010-delivery.md) — the standard a checkpoint is held to: a
step that reports success without doing anything is the fault, not the shortcut.
@@ -1,60 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 23. Issues have a front door, separate from the operational memory
## Context
Findings that were nobody's task accumulated in a table inside the decision ledger — a package
install reporting success while installing nothing, a manifest key read by no code, an
end-to-end harness dead for months. They were measured, true, and unowned: a table row cannot
be assigned, diagnosed or closed.
The mesh already has an operational memory holding roughly a hundred and thirty entries,
indexed on symptoms. The obvious move — put these there — is wrong, and the reason is the
distinction worth recording.
## Considered options
1. **Leave them in the ledger.** Rejected: a ledger records decisions taken, and these are
the opposite — questions nobody has answered.
2. **Put them in the operational memory.** Rejected. That store answers *how do I fix this
occurrence*; these are *why does the design allow this at all*. Filing them there makes
them findable by symptom and unfindable as open questions, and nothing there has a state
that can be closed.
3. **A numbered issue folder in HQ, deliberately narrow.** Chosen.
## Decision
`04-ISSUES` is the front door for something wrong at the level of **design or governance**:
a rule enforced by nothing, a stated invariant that is false, a failure the design permits to
be silent, or a symptom whose owner cannot be found without the whole mesh in view.
One numbered folder per issue: the report with the symptom as observed and the evidence, and
a diagnosis document carrying the trail, dated, including what was ruled out.
**This is not a second copy of the operational memory.** An issue here is a question HQ must
*answer*; an entry there is an incident someone must *clear*. An issue whose answer is a
general lesson belongs in both — and the operational memory is searched first, because if the
answer is already there this was never an issue.
## Consequences
- A finding gets a number, a state and an owner, and closing it is a visible act.
- The symptom-to-component trail accumulates in a place where the whole mesh is visible, which
is where cross-component diagnosis has to happen.
- The boundary needs judgement on every report, and will sometimes be got wrong. Filing too
narrowly loses a finding; filing too widely rebuilds the operational memory here, which is
the outcome HQ's separation was argued against
([ADR 0019](0019-hq-is-its-own-repository.md)).
- Six issues opened on creation, all previously unowned observations.
## References
- [`04-ISSUES/README.md`](../04-ISSUES/README.md) — the boundary table and the frontmatter
schema.
- [`00-META/process/03-issues.md`](../00-META/process/03-issues.md) — the playbook.
@@ -0,0 +1,95 @@
---
topic: what runs on it
status: accepted
date: 2026-08-30
deciders: jochen
reconstructed: false
---
# 24. Model access is a provision, and a licence is a thing with a name
## Context
Everything in this mesh that thinks needs a model, and there is more than one way to reach one:
| | |
|---|---|
| **hosted services** | several vendors, each with its own account, quota and key |
| **models the mesh runs itself** | open-weight models on a node with the hardware for them |
And the choice is **per consumer, deliberately**: a workstation's own session on one account, a
laptop on the organisation's, two hired workers on the mesh's local model because their work does
not justify paid tokens. Those are three different answers to one requirement, held at once, in
one mesh.
**The existing system has the hard half of this already** — automatic licence refresh and
switching between accounts when one is exhausted — and it works. It is not being replaced because
it was wrong; it is being rebuilt because it lives in a place that cannot express the rest.
## Decision
**Model access is a provision.** A module that needs to think declares `requires: model-access`;
anthropic, openai, grok and a locally-run model are four modules that provide it. That is
[ADR 0009](0009-modules-and-the-graph.md)'s mechanism unchanged, and it buys the things that
mechanism already buys: several implementations of one job, a refusal when more than one could
answer, and choosing by assigning the one you want.
**A locally-run model needs nothing new at all.** It is a mesh-scoped provision on the node with
the hardware — the same shape as a database, including the credential.
### A licence is a named thing, and the name is the operator's
Not an anonymous credential hanging off a provider. *The personal account*, *the organisation's
account* — those are names a person uses, and the mesh has to use them too, because the whole
point is saying **which one** a given consumer uses.
**Many to many.** One provider has several licences; one licence serves several consumers. So it
is **not a claim** — claims are for things only one holder may have, and two machines sharing an
account is the ordinary case rather than a collision.
### Four things this needs that the mesh does not have
Written as gaps rather than as design, because each is a real piece of work and pretending
otherwise is how a plan becomes a surprise.
**1. A provider that is not on a node.** A mesh-scoped provision today is answered by *the machine
running it*, and the reachability rule refuses two ends that share no private network. A hosted
service is on nobody's machine and is reached over the public internet. That is a third scope —
answered by a record rather than by a node — and the reachability rule must not apply to it.
**2. A secret the mesh is given rather than one it mints.** Every credential the mesh handles
today it generated itself, sealed to both ends, and discarded. An API key arrives from a person.
The missing verb is *accept*: take a value, seal it to each holder, and **discard the plaintext**
— because a mesh that keeps operator-supplied keys readably is the arrangement this project
[measured and rejected](0009-modules-and-the-graph.md).
**3. A consumer that is not a machine.** *This worker uses that licence* is a binding to an agent,
not to a node. [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) already says an agent
holds credentials and that delivery follows its node bindings and modality — so what is delivered
still lands on a machine, and what is **chosen** is chosen per agent. The provisions model has no
consumer identity other than a node.
**4. Switching is a reaction, not a declaration.** Everything here is desired state, reconciled by
comparison. A licence that hits its limit and must be swapped is a response to something observed,
and it cannot be expressed as a declaration without the declaration meaning *whatever is working
right now* — which is not a thing anybody declared. **It belongs with observability, changing a
binding**, and the binding is then declared as usual. Saying this plainly is what stops the
declaration language growing a conditional.
## Consequences
- **The refusing rule applies here and will be felt.** A mesh holding three ways to reach a model
refuses every consumer that has not said which — which is correct and is a great deal of
saying-which the first time. The remedy is one assignment per consumer, and the message names
the candidates.
- **A licence outliving its holder is a live credential nobody is watching.** The same rule the
provisioner follows applies: what the mesh granted and no longer grants is withdrawn.
- **Nothing here makes a node authenticate to a model provider.** ADR 0001 holds: an agent does.
What changes is that the mesh can now say *which agent, which licence, which node it lands on*,
which is the fact ADR 0001 records as missing.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — provisions, scope, choosing, and sealed credentials
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — agents hold credentials, not nodes;
`hal/ai` as the context owning provider grants and rotation
@@ -1,76 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 24. The folder numbering is the flow, and decision records run oldest first
## Context
Two orderings were wrong in ways that only show up when someone new reads the repository.
**The folders.** Decisions lived in an unnumbered folder that sorted after the numbered ones,
so the repository's most load-bearing content read as an annex.
The sibling HQ repository for the PAPA platform had already solved this and solved it
crookedly: its design folder existed from its initial commit, and when its decision folder was
finally promoted it took the **next free number** rather than its place in the sequence. That
repository now reads `01 research → 03 decision → 02 design`. A decision precedes the design
it authorises and is numbered after it. By the time this was visible, the design folder was too
settled to renumber.
**The records.** Fourteen decisions had been taken in implementation and never written down —
the broker, the module abstraction, the mesh database, the artifact, the three silos and the
rest. Meanwhile two records existed, holding numbers 0001 and 0002, for decisions taken last.
## Considered options
1. **Match the sibling repository exactly**, inheriting its ordering. Rejected: structural
parity is worth something, but not the cost of copying a scar the other repository would
not choose again.
2. **Leave the folder unnumbered.** Rejected — the annex problem, and it leaves an unexplained
gap for anyone arriving from the sibling repository.
3. **Number by position in the flow, and renumber the records chronologically.** Chosen, on
the grounds that this repository was four commits old and nothing outside it cited a
number. That is the only window in which either renumbering is free.
## Decision
**The numbering is the flow.** Research produces a decision; the decision authorises a design.
So `01-RESEARCH`, `02-DECISIONS`, `03-DESIGN`, `04-ISSUES`. Following the folder numbers walks
the process in the order it happens.
**Decision records are a chronological ledger.** They run oldest first. The fourteen decisions
already taken in implementation were back-filled as records 0001–0014, each dated from the
history, each carrying `reconstructed: true` and saying so in its first lines, and each citing
the commit, pull request or knowledge-base entry it was recovered from. The two existing
records moved to 0015 and 0016.
A reconstructed record is not a transcript. Where the deliberation is not recoverable it states
what the alternatives were and why the chosen one won on the evidence available — not a
discussion that did not happen. Where a date is not establishable it says so.
The foundational folder is `00-META`, matching the sibling repository.
## Consequences
- The repository reads in process order, and the gap at `03` that a reader coming from the
sibling repository would notice is explained by this record.
- The design documents can cite reasoning instead of asserting rules, because the reasoning now
exists.
- Structural divergence from the sibling repository, deliberately, in exactly one place. It is
recorded here so that the difference reads as a choice rather than an accident.
- **Record numbers are now stable and renumbering is over.** This decision spends the one
window that existed; a future record takes the next free number regardless of its date.
- Reconstructed records carry a standing risk: they are the most confident-sounding documents
in the repository and the least directly witnessed. The `reconstructed` flag exists so that
is never invisible.
## References
- The sibling repository's restructure of 2026-07-13 moved its decision folder in a single
commit of twelve renames with no content change, alongside the same status-into-frontmatter
and playbook changes made here.
- [`02-DECISIONS/README.md`](README.md) — the format, and the note on reconstructed records.
@@ -0,0 +1,105 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0019-how-this-repository-works.md
---
# 25. The design record is read where it is written, never copied to be found
## Context
**These documents cannot be found by searching the mesh's memory, and never could.** Checked on
2026-08-23 and again on 2026-08-31, against both the symptom-indexed store and the structured
archive, using a decision record's full title and a distinctive phrase from a design document: no
result, no partial match, no stale copy.
That matters because of what was promised. The objection to giving this material its own
repository was that the mesh already has a knowledge store, and a second one repeats the mistake
that store was created to fix. **The answer offered was indexing rather than location** — that
these documents would be returned beside everything else in a search, so where they were authored
became a separate question. The indexing was never built.
**The claim has since stopped being load-bearing**, which is why this is a decision rather than an
incident. [`README.md`](../README.md) names the gap in the place the claim used to sit, and
[ADR 0019](0019-how-this-repository-works.md)'s reasoning rests on cadence, reviewers and scope —
none of which depend on being searchable from elsewhere. What remained was an unbuilt capability
and an open question, recorded as
[`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md).
**A signpost was added on 2026-08-31 and measured.** One entry in the mesh's memory naming what
lives here and when to come looking. A search for *design records, decisions, repository* returns
it; a search phrased the way somebody actually asks — *why is the mesh built this way* — returns
nothing, because the store matches terms and not meaning. **Reachable is not the same as
surfacing**, and the measurement is what established which one a signpost buys.
## Considered Options
1. **A one-way sync into the mesh's memory.** A job reads this repository on a schedule and writes
the documents into the searchable store. It works with what exists today and needs nothing
built first. **Rejected**, because it creates a second copy of every document, and the failure
mode of a derived copy is the one this repository is least able to tolerate: *the copy that is
searched quietly stops matching the copy that is edited*, and the enforced one wins. A design
record that has silently diverged from the reasoning it claims to carry is worse than one that
cannot be found — the first misleads, the second merely fails.
2. **Leave the signpost and close nothing.** Honest, free, and it keeps the gap visible.
**Rejected as an end state**, though it is what stands until the option below exists. It
answers only for a reader who already suspects these documents exist, which is precisely not
the reader the mesh's memory is designed for.
3. **An agent reads this repository directly, and the search consults it.** Nothing is copied.
**Adopted.**
## Decision
**The design record is read where it is written.** Retrieval is an agent reading this repository,
not a copy living in a second store — and a search of the mesh's memory consults that agent, so
what it knows appears beside ordinary results rather than only when it is asked.
Both halves are the decision. The first alone is merely a reader, and would leave this repository
reachable but not surfacing — the state measured above. **The second half is what discharges the
promise** that these documents are returned beside everything else.
**There is no copy, and that is the point.** No sync, no schedule, no reconciliation, and nothing
that can drift, because there is only ever one of each document. It is also always current,
including for work that is not yet committed.
**The direction of reading is one-way and stays that way.** The agent reads this repository and
answers from it. Nothing flows back: this repository is public, the mesh is not, and a return path
would be how installation-specific detail arrives into documents that must not carry it
([`README.md`](../README.md)).
## Consequences
**This repository stops being a fourth knowledge system, properly.** The original objection was
about adding a knowledge *system*. An agent with read access adds no store at all — which answers
the objection more completely than the indexing that was promised, rather than merely as well.
**ADR 0019's promise is amended, not satisfied.** It said these documents would be *indexed*. They
will not be. They will be *read*, and the search will ask. The commitment that survives is the one
that mattered — that a searcher finds them without already suspecting they exist — and the
mechanism behind it is different from the one named.
**It is gated on an agent that does not exist yet.** Until it does, the signpost is what stands,
and this repository is reachable rather than surfacing. That is a known and stated gap, not a
silent one — and the gap is now a build task with a decided shape rather than an open question.
**The search must degrade honestly.** When the agent cannot be reached, a search has to say that
this material was not consulted. A result set that silently omits it looks identical to one where
nothing matched, and *silence and success must never look alike*
([ADR 0004](0004-a-node-and-how-it-joins.md)) — the rule this repository has now paid for twice.
**A rule states how it is checked, and this one is checkable.** The check is the measurement that
produced this record: search the mesh's memory for a phrase that appears only in a design document
here, and require it back. That check fails today, deliberately, and passing it is what closes
`04-ISSUES/006`.
## References
- [`04-ISSUES/006`](../04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md) —
the gap, the two measurements, and why closing it early was refused
- [ADR 0019](0019-how-this-repository-works.md) — the promise this amends
- [`README.md`](../README.md) — the objection, and the gap named where the claim used to sit
@@ -1,89 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 26. Every decision is a record; there is no ledger
## Context
HQ carried a decision ledger at its root: a chronological table of forty-one numbered
decisions, each with who decided and a pointer to where the reasoning lived. It was created
deliberately, to make decisions findable and to give a home to decisions too small to warrant
a document.
By the time the decision records were back-filled
([ADR 0024](0024-the-numbering-is-the-flow.md)) the ledger had become three things at once,
and only one of them was still needed.
Classified, its forty-one entries were: ten restating a record, eleven restating design
documents, fifteen describing how this repository works — with the reasoning in a README rather
than anywhere citable — three small rules with no home at all, and two superseded stubs.
So the ledger was mostly a copy. Worse, it was a **hand-maintained index**, which
[ADR 0022](0022-status-lives-in-frontmatter.md) had just finished rejecting for the
decision-record index on the grounds that it had drifted after a single addition. Keeping one
copy of that pattern while removing another is not a position.
It had also produced a naming collision that a directory listing makes plain: `DECISIONS.md`
beside `02-DECISIONS/`, holding different things.
## Considered options
1. **Keep the ledger.** Rejected. It duplicates the records, restates status, and is the exact
hand-maintained index this repository decided against elsewhere.
2. **Keep it, renamed, for small decisions only.** Rejected, and this is the option worth
arguing with — it is genuinely useful to record a decision without writing a document. But a
decision small enough to be one table row is almost always a **rule** rather than a
decision, and a rule belongs in [`how-we-build.md`](../00-META/how-we-build.md) where it is
enforced and where its reasoning is kept. That is where the three orphans went.
3. **Every decision is a record; nothing else.** Chosen. This is how the sibling HQ repository
for the PAPA platform works, and it has no ledger of any kind.
## Decision
**If a decision is worth recording, it is worth a record. If it is not worth a record, it is
not recorded.**
`02-DECISIONS` holds every decision. There is no ledger, no index file, and no central status
of any kind. The chronological view — decisions in the order they were taken — is *generated*
from record frontmatter, which is what the ledger was actually for.
Content that was only in the ledger was rehomed rather than dropped:
| Was | Went to |
|---|---|
| Decisions about how this repository works | Records [0019](0019-hq-is-its-own-repository.md)–[0025](0025-hq-is-the-source-of-the-constitution.md) |
| Small rules with no record | [`how-we-build.md`](../00-META/how-we-build.md) — the package rule, and two already there |
| Lab decisions not stated in the design | [`03-DESIGN/01-to-be/01-end-to-end-testing.md`](../03-DESIGN/01-to-be/01-end-to-end-testing.md) |
| "Deliberately not decided" | The research effort and design document each question belongs to |
| Unowned observations | [`04-ISSUES`](../04-ISSUES/) ([ADR 0023](0023-issues-have-a-front-door.md)) |
## Consequences
- One place to look, and nothing to keep in sync. The collision between the ledger and the
record folder is gone.
- Structural parity with the sibling repository on decisions, which
[ADR 0024](0024-the-numbering-is-the-flow.md) deliberately broke on folder numbering. The
divergence is now exactly one thing, and it is the one thing that was argued for.
- **Writing a record is now the only way to record a decision, and a record is more work than
a table row.** The real risk is that a small decision goes unrecorded because nobody wanted
to write a document. The mitigation is that a small decision is usually a rule, and
`how-we-build.md` takes rules cheaply — but this is a cost, not a solved problem, and it is
the thing to watch.
- The chronological view now depends on the generator existing and being run. It did not
before.
- Two superseded ledger stubs had no record of their own. The position that documentation
lives inside the code repository is now recorded only as superseded context in
[ADR 0019](0019-hq-is-its-own-repository.md); the system-container position is explained in
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md). Neither is lost.
## References
- The sibling PAPA HQ repository: root holds only agent instructions and a README; every
decision is a numbered record, and its graduation playbook has no path for an unrecorded
decision.
- [ADR 0022](0022-status-lives-in-frontmatter.md) — the hand-maintained-index argument this
applies consistently.
@@ -0,0 +1,130 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0004-a-node-and-how-it-joins.md
---
# 26. The mesh has a session of its own, and it is the node session's mechanism
## Context
[ADR 0004](0004-a-node-and-how-it-joins.md) gives every node a session: one per node, permanent,
remembering across callers, its system prompt the node's engram, reachable over the broker like
everything else. **Any node can message any node**, and that is called the one part of the system
that is genuinely a mesh — symmetric, with no centre.
**There is no way to address the mesh itself.** A question that spans machines — *what is running
across all of this*, *which nodes are behind*, *why is it built this way* — has to be put to some
node, which then asks the others. That works, and it makes a mesh-wide question **nobody's
question**: every node answers it as a foreigner, from a position where the whole is not in view.
**Three things independently arrived at the same missing piece.**
[ADR 0025](0025-the-design-record-is-read-not-copied.md), taken hours before this one, commits to
an agent that reads the design repository directly and answers into search. That agent has to
exist, run somewhere, and be askable — and nothing in the record says what it is or where it
lives.
[`14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) records, as a gap deliberately
not half-built: *this worker uses that licence is a binding to an agent, not to a node* — and the
provisions model has no consumer identity other than a node. A session that must be assigned a
licence is exactly that consumer, and node sessions are already one.
**And ADR 0004 never said how a session is set up.** It describes behaviour and stops: nothing
states how a session starts, where its context lives, how the engram reaches it, or how a message
off the broker becomes a prompt. There is no design document for it. That gap was invisible until
something had to be built *like* a node session, because describing a second instance of a
mechanism requires the mechanism to have been described once.
## Considered Options
1. **No mesh session; keep relaying through a node.** Costs nothing and works today. **Rejected.**
It leaves mesh-wide questions belonging to nobody, and it does not survive contact with
ADR 0025 — that agent still needs a home, so the thing gets built anyway, unnamed, as an
attachment to whichever node happened to host it.
2. **A new kind of agent, built separately.** Purpose-built for the whole mesh. **Rejected.** It
would hold a session, a memory, a licence and broker plumbing — every one of which the node
session already has. Two implementations of one mechanism drift, and the vocabulary collision
that [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) exists to undo began exactly this
way: two things that were nearly the same, built twice, until neither word meant one thing.
3. **The same mechanism, started in a different context.** **Adopted.**
## Decision
**The mesh has one session, addressed as the mesh, and it is a node session in every respect but
three.**
| | |
|---|---|
| **the context it starts in** | the mesh's, not a machine's — this is the whole of what makes it different |
| **its engram** | the mesh's system prompt, as a node's engram is that node's |
| **its licence binding** | assigned in its own right, not inherited from the machine it runs on |
Everything else is unchanged and deliberately so: it is permanent, it remembers, it is reachable
over the broker, it holds its own tools, and switched off it still answers *I am switched off*
rather than falling silent.
**It runs on the node that holds the control plane** — not for convenience, but because that node
is already the one place excepted from *compromise of a node is compromise of that node*
(ADR 0004). An agent able to reach everything, placed anywhere else, creates a **second** such
place. Putting it where the authority already sits concentrates nothing new.
**It is an addition to per-node messaging and never a replacement.** Every node remains directly
addressable. This is not a preference: ADR 0001 holds that losing the control plane costs *change,
not operation*, and a mesh whose only conversational surface lives on that node would lose the
ability to ask anything while every machine kept running perfectly. **The front door may not be
the single point.**
**It is not an employee** ([ADR 0003](0003-agents-are-persistent-employees.md)). Nobody hires it,
it holds no task queue, it is never drained or reassigned. What it does with work that belongs
somewhere else is **dispatch it** — to node sessions, or to workers — which is what a node session
already does when asked something it does not have.
**It is ADR 0025's reader.** The agent that reads the design repository and answers into search is
this session, not a second one. One agent, one memory, one place to reach; two would both need
that repository and would eventually disagree about what it says.
**"One per node" is about address, not about process count.** ADR 0004's rule — *two and nothing
decides which replies* — forbids ambiguity in who answers when a **node** is addressed. The mesh
session answers when the **mesh** is addressed. The control-plane node therefore hosts two
sessions and no ambiguity, and stating this here is what stops it reading as a contradiction
later.
## Consequences
**The node session's setup must now be designed, and it never was.** This decision is expressed as
*the same as a node session, elsewhere*, which is only meaningful once that mechanism is written
down. The design document covering both is the immediate consequence of this record, not a
follow-up to it.
**A consumer that is not a machine stops being deferrable.** The licence binding above is the gap
`14-model-access.md` names, and it now has two consumers rather than a hypothetical one. Until it
exists, a session's model access can only be expressed as *this module on this machine*, which
cannot say *this node's session uses the personal licence and the mesh's uses the company one* —
the thing the binding is for.
**Symmetry is preserved, and it is worth being precise about why.** ADR 0004's claim is about what
a node can reach, and it is untouched: node-to-node messaging is unchanged, nothing is routed
through the mesh session, and it is a participant rather than a hop. What arrives is a
participant that happens to be the one a person usually addresses.
**Availability degrades to inconvenience rather than to silence** — but only because of the
addition rule above. If that rule is ever relaxed, this consequence inverts, and it inverts
quietly: everything keeps working and nobody can ask about it.
**The surface a person uses is not decided here.** That a board is a good place to talk to it is
likely and is not this record's business; the session is reachable over the broker like everything
else, and what puts a text box in front of it is a separate choice.
## References
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the node session this extends
- [ADR 0025](0025-the-design-record-is-read-not-copied.md) — the reader this session is
- [ADR 0003](0003-agents-are-persistent-employees.md) — the vocabulary this is not
- [`03-DESIGN/01-to-be/14-model-access.md`](../03-DESIGN/01-to-be/14-model-access.md) — *a
consumer that is not a machine*, the gap this makes concrete
@@ -0,0 +1,109 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 27. A provision names what the consumer is coupled to, not the role it plays
## Context
Provisions are named after roles. The catalogue and every test fixture built so far say:
```
provides: database
requires: database
```
**Nothing distinguishes one engine from another.** A module requiring `database` is satisfied by
any module providing `database`, so a module written against PostgreSQL can be matched to a
provider of Microsoft SQL Server, resolve as satisfied, deploy, and fail on its first query.
**The mesh runs several engines** — PostgreSQL, Microsoft SQL Server, MariaDB, and others behind
products that expose their own. This is not a hypothetical collision.
**The failure is in the direction that hides.** Resolution *succeeds*. Nothing is refused, nothing
is logged, and the breakage surfaces later as an error inside an application, on a machine, with
nothing connecting it back to a match made elsewhere by something that thought it had done its
job. **A wrong answer delivered confidently costs more than a refusal**, and the whole point of
refusing on ambiguity ([ADR 0009](0009-modules-and-the-graph.md)) was to not do this.
**How it got in:** every test written for the resolver had exactly one provider of each name, so no
mismatch was expressible and none was caught. The fixtures agreed with the design. That is the same
fault as [`04-ISSUES/005`](../04-ISSUES/005-pipeline-test-harness-unbuildable/00-report.md)'s
imagined output and [`019`](../04-ISSUES/019-a-comment-asserting-a-fact-about-a-machine/00-report.md)'s
unchecked comment, at the level of a name rather than a line.
## Considered Options
1. **Keep role names; let the operator assign correctly.** The mesh would refuse ambiguity when two
providers exist, so a person picks. **Rejected.** It makes correctness depend on somebody
knowing that the module they are assigning speaks a particular dialect — which is exactly the
knowledge the provisioning model exists to remove. And with one provider of each name, nothing
is ambiguous and nothing is asked.
2. **A role name plus a `flavour:` or `engine:` qualifier**, matched as a second field.
**Rejected.** Two fields that must agree is a constraint the resolver has to enforce and a
manifest author has to remember, to express something one field already can. The name is the
contract; splitting it invites a requirement that names a role and forgets the qualifier, which
then matches everything again.
3. **The name says what the consumer is coupled to.** **Adopted.**
## Decision
**A provision is named for the thing a consumer's code is written against.**
```
provides: postgres-database
requires: postgres-database
```
**The test is whether the consumer can tell the difference.** If swapping the provider would break
the consumer, the name must say which provider — because a match that breaks the consumer is not a
match. If the consumer genuinely cannot tell, a role name is correct and better.
| provision | | why |
|---|---|---|
| `postgres-database`, `mssql-database` | **specific** | applications are written against a dialect; a swap breaks them |
| `route` | **role** | the consumer wants its name reachable and does not care what proxies it |
| `resolver` | **role** | the consumer wants names to resolve |
| `artifact-store` | **role** | the consumer fetches by digest over a protocol, and nothing else |
**`database` is not a provision and may not be provided.** There is no context in which an
application talks to a generic database: it talks to PostgreSQL or it talks to SQL Server. A name
that cannot be true of any real consumer should not be expressible.
**This is about coupling, not about products.** Two providers of `postgres-database` — a container
on this node and a managed instance elsewhere — are interchangeable and *should* both match. What
may not be interchangeable is what the consumer's queries are written in.
## Consequences
**Every manifest that names a database changes.** Doing this now costs a rename across a handful of
examples. Doing it after modules are migrated costs it across all of them, plus every deployment
that resolved against the old name.
**Wrong requirements now fail loudly, and at the right moment.** A module requiring
`postgres-database` where only `mssql-database` is provided is unsatisfiable, so it is **refused at
resolution** with both names visible — rather than deployed and broken later. This is the property
that was lost, restored.
**Generic role names are still right, and the rule says when.** This does not push specificity
everywhere; it puts it exactly where a consumer is coupled. Naming `route` after a particular proxy
would be the same error in the other direction, and would prevent a swap that genuinely changes
nothing.
**It is checked, not merely stated** ([`00-META/how-we-build.md`](../00-META/how-we-build.md) §5).
A manifest providing a name known to be engine-generic is refused, naming what to say instead.
Without that, this record is a convention, and a convention is what the previous naming was.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — provisions, and refusing on ambiguity
- [`03-DESIGN/01-to-be/07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md) — *the
provisioning model uses databases, roles and schemas as PostgreSQL means them*, which is this
record's point made about the substrate before it was made about modules
@@ -1,112 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 27. The product is Novox Mesh; Nox is an identity, not a second system
## Context
The name `HAL` was never chosen. This began as a dotfiles repository, the first commits in
February 2026 adopt dotfiles and per-node overrides, and the name arrived with the code — as
recorded in
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md),
most of the current shape is inherited from that origin rather than designed for a mesh. The
name is part of the inheritance.
Three things make it worth changing rather than living with.
**It is borrowed, and borrowed badly.** HAL is the canonical *untrustworthy* machine
intelligence. For infrastructure whose entire proposition is that it manages your machines,
heals itself, and is trusted with credentials, that is an unhelpful flag to fly, and it is not
a name anyone owns.
**There is a name available that is owned.** The company is Novox. A product of Novox should
carry that lineage rather than a film reference.
**A platform and a persona are different things, and one name was doing both.** `HAL` named the
mesh *and*, implicitly, the thing an operator talks to. Those are separate concerns — the
platform is what runs; the persona is who answers.
## Considered options
1. **Keep `HAL`.** Rejected. Every reason to keep it is sunk cost, and the sunk cost is at its
smallest today.
2. **Rename everything to a single new name covering platform and persona.** Rejected: it
repeats the conflation that made `HAL` ambiguous.
3. **Separate the two: a product name and an identity.** Chosen.
## Decision
**The product is `Novox Mesh`**, shortened to `mesh` in internal use — repository names, the
module namespace, environment variables, paths.
**`Nox` is an identity of Novox**, and specifically an **agent identity within the mesh's own
model** — a named participant, exactly as
[ADR 0012](0012-agents-are-persistent-employees.md) defines one. Not a separate product, not a
separate runtime, not a privileged path.
**Nox is the agent of the mesh, not of a node.** This is the part that carries weight:
- **Every node keeps its own identity.** That already exists and stays — a node is a named
participant with its own character, and addressing one directly remains possible and normal.
- **Nox is scoped to the whole mesh.** It is what the mesh is called when the mesh itself
speaks, rather than one machine within it.
- **Nox addresses node identities.** Asking Nox for something that lives on one node is Nox
talking to that node, not a human choosing a machine.
- **A human mostly talks to Nox.** It is the front door.
That last point makes Nox the concrete form of the vision in
[`00-META/mission.md`](../00-META/mission.md): *an agent states an intent and the mesh carries
it out — no console to open, no runbook to follow, no remembering which node holds which
thing.* Nox is who that intent is stated to. The mission described the behaviour; this names
the thing that has it.
Nox holds no private channel. Whatever it can do, it does through the same surfaces every other
agent uses — which is not a naming detail: a persona with its own path would be the one part of
the mesh with no human checkpoint, and the skeleton already rules that out.
`HAL` is retired.
**Timing is the substance of this decision, not an aside.** The skeleton in
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) is not built. Renaming
before it exists costs a search and replace across research documents. Renaming after costs the
same class of migration as everything else this repository is trying to avoid, and would
therefore not happen.
## Consequences
- **The as-is layer keeps `HAL`.** It describes what runs, and what runs is called HAL. The
to-be layer uses `mesh`. The rename is part of the migration, and the two-layer split
([ADR 0020](0020-design-is-written-in-two-layers.md)) is what makes holding both names
coherent rather than confusing.
- **Records 0001–0026 keep `HAL`.** They are immutable and they say what was decided when it
was decided. No record is edited for a name.
- Tier 2 cannot be `mesh-mesh`. The control plane is **`mesh-control`**; `mesh-broker` was
rejected because the substrate already contains a message broker.
- The namespace, environment variable prefix, service names and on-disk paths all change. In
the existing system that is a migration and is not attempted here.
- **`mesh` is a generic word**, and it already means something specific in infrastructure — a
service mesh is a different thing. Recorded as a known trade rather than an oversight: the
full name `Novox Mesh` is distinctive, and the short form is internal.
- The persona has a name before it has behaviour. That is the right order — it is an identity in
a system that already has a model of identities, so it needs no new machinery to exist.
- **Except in one respect, and it is a real gap.**
[ADR 0012](0012-agents-are-persistent-employees.md) binds every agent to a home node, one to
one, with a workspace on that machine. A mesh-scoped agent has no home node by definition, so
the model does not currently have a shape for Nox. Extending it — an agent whose scope is the
mesh rather than a machine — is a decision of its own and is not taken here.
- Two levels of identity now exist where there was one: the node, and the mesh. The distinction
has to stay visible in every surface, or "ask Nox" and "ask a node" collapse into each other
and it stops being clear who is answering.
## References
- [ADR 0012](0012-agents-are-persistent-employees.md) — what an identity is in this system, and
why `Nox` needs no separate mechanism.
- [ADR 0020](0020-design-is-written-in-two-layers.md) — why the as-is and to-be layers can
legitimately use different names for the same system.
- The dotfiles origin, and the naming inheritance it explains:
[`03-DESIGN/00-as-is/10-module-catalogue.md`](../03-DESIGN/00-as-is/10-module-catalogue.md).
-89
View File
@@ -1,89 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 28. HQ is company-scoped; the mesh is its first product
## Context
This repository was `hal-hq` — one product's headquarters, named for the product. Then the
product was renamed ([ADR 0027](0027-the-product-is-novox-mesh.md)), which forced the question
of what the repository is actually the headquarters *of*.
Two facts settled it, and both were checked rather than assumed.
**Novox already delivers other things.** The company's forge organisation holds live projects
beside the mesh, and they are registered as build sources — meaning the mesh already builds and
deploys them. They are not hypothetical future products; they exist and ship today.
**They are tenants, not peers.** They run *on* the mesh. Every one of them is developed,
delivered and hosted by it. So the mesh is not one product among several — it is the ground the
others stand on.
That distinction decides the scope. If the mesh were a product beside others, a per-product HQ
would be right. Because it is the substrate the company operates on, a decision about the mesh
is a decision about how the company works.
## Considered options
1. **`mesh-hq` — one HQ per product.** The safe choice, and the reversible one: a second
product creates its own HQ and shared practice graduates upward later. Rejected, knowingly,
because it models the mesh as a peer of things that are actually its tenants.
2. **A company HQ *and* a product HQ, from the start.** Rejected as ceremony — two repositories
for one operator, and the constitution's own YAGNI rule says not to.
3. **One company-scoped HQ, `novox/hq`, with the mesh as its first product.** Chosen.
## Decision
The repository is **`novox/hq`** — Novox's headquarters, not the mesh's.
It holds the reasoning behind what Novox builds. Today almost all of that is the mesh, because
the mesh is what Novox is building. That is a fact about the present, not a definition of the
repository.
**The scope of each document is fixed now, so the eventual split is mechanical rather than
archaeological:**
| Scope | Documents | Moves if products separate? |
|---|---|---|
| **Company** | [`how-we-build.md`](../00-META/how-we-build.md), [`process/`](../00-META/process/), [`repos.md`](../00-META/repos.md), this record and [0019](0019-hq-is-its-own-repository.md)–[0027](0027-the-product-is-novox-mesh.md) | No — they stay at the top |
| **Product (mesh)** | [`mission.md`](../00-META/mission.md), [`context.md`](../00-META/context.md), [`effect.md`](../00-META/effect.md), `01-RESEARCH`, `03-DESIGN`, `04-ISSUES`, records 0001–0018 | Yes — into a product section |
The folders are **not** restructured now. One product's content under a company name is
correct while there is one product's worth of it, and nesting before there is anything to nest
is the ceremony option 2 was rejected for.
## Consequences
- Engineering practice has a home that does not belong to the mesh. `how-we-build.md` — never
write to production directly, migrations for schema changes, runtime evidence for behavioural
criteria — is true of any Novox project, and its being in a mesh repository was always a
slight mislabelling.
- The constitution derived from it ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md))
can legitimately govern work outside the mesh. Under a product HQ it could not have, without
either duplicating or reaching across repositories.
- **The bet is not entirely forward-looking, and that is worth being honest about.** Novox
already has work that is *not* a mesh tenant — client engagements and at least one product
that is developed outside it. So the company genuinely has a scope wider than the mesh
**today**, which strengthens the case for a company HQ and simultaneously means the split in
the table above is closer than "some day". The table is not a precaution; it is a plan whose
trigger already half-exists.
- What has *not* happened yet is any of that work needing the constitution. That is the actual
trigger ([ADR 0025](0025-hq-is-the-source-of-the-constitution.md)): the moment something
outside the mesh must be governed by the same rules, product-level content moves down a level
and this repository becomes what its name already claims.
- A new repository was created rather than the old one transferred, because the forge predates
the transfer API. The original was verified to contain nothing the new one lacks — every ref
an ancestor, no tags, issues, pull requests, releases or wiki content — and then removed.
- The mesh's own documents now live one conceptual level below the repository they are in. A
reader arriving at `01-RESEARCH` should understand it as the mesh's research, not Novox's.
Nothing in the folder names says so, and that is the cost of not restructuring.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name that forced the question.
- [ADR 0019](0019-hq-is-its-own-repository.md) — why HQ is a repository at all. Unchanged; only
its scope moves.
@@ -0,0 +1,118 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 28. The substrate supplies the control plane and nothing else
*Corrects one row of [ADR 0006](0006-the-substrate-and-the-control-plane.md) and makes explicit
something it left unsaid. The rest of that record stands.*
## Context
ADR 0006 defines the substrate by a circularity: **what the control plane needs in order to run,
and cannot ask itself for, because it is not running yet.** Two questions, and both must be
answered *yes* for something to be substrate.
Its membership table admits the object store on this line:
| role | product | |
|---|---|---|
| object store | **MinIO** | it cannot grant itself a bucket |
**That answers the second question and assumes the first.** It is true that a control plane cannot
grant itself a bucket. Nothing establishes that it needs one.
**It does not.** Verified 2026-08-31 against `mesh-control`: no S3 client, no bucket, no object
storage of any kind outside comments. Artifacts reach nodes as content-addressed blobs in the OCI
registry, and the code records the decision and its reasoning:
> One store, and it is the registry the bootstrap already pulls from. An OCI registry is a
> content-addressed blob store that happens to also understand images… The alternative considered
> was a second store beside it — S3-shaped, buckets, signed URLs. It is the right answer for
> objects that are *mutable*, or need per-reader access, or are not build output. None of that
> describes a digest-pinned archive, and standing up a second service to hold one kind of
> immutable blob means two things to run, two things to back up and two ways for an artifact to be
> missing.
**The row is inherited from the system being replaced**, where an object store distributed module
tarballs. Here nothing does, and the row was never re-tested against the definition it sits under.
**A second thing ADR 0006 never says:** whether a substrate service and a module of the same
product are the same instance. It says the substrate is *not the control plane* and *not a place
for logic*, and stops. The question is not idle — an application wanting a database, on a mesh
whose substrate is already running PostgreSQL, has an obvious wrong answer available.
## Considered Options
1. **Leave the object store as substrate, unused.** Harmless-looking. **Rejected.** A membership
list that includes something nothing needs is a list that has stopped being derived from its
test, and the next member is admitted by precedent instead of argument. It also mandates that
every mesh run a service no mesh uses.
2. **Applications share the substrate's instances.** One PostgreSQL, one of everything.
**Rejected**, below.
3. **The substrate is exactly what the control plane consumes; everything else is a module.**
**Adopted.**
## Decision
**The object store is not substrate.** It fails the first half of the test: the control plane does
not need one. An object store is an ordinary module, required through the module graph like
anything else, and a module wanting one depends on a module providing one.
**The substrate has four members, not five**: a relational store, a message bus, an image registry,
and conditionally an identity provider. The registry stays — the control plane genuinely cannot
deliver an artifact without somewhere to put it.
**A substrate service and a module of the same product are different instances, and are not
shared.** The mesh's own PostgreSQL and a PostgreSQL a workload was given are two servers, two
containers, two lifecycles.
Three reasons, and the first is the one that matters:
**The substrate is not in the module graph.** It is raised from the pinned bundle the host carries,
before any mesh exists to declare it. A workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move — which is every property the
provisioning model exists to provide.
**It would put workload data in the control plane's own store.** The mesh's contexts own their
stores exclusively ([ADR 0008](0008-a-context-owns-its-store.md)). An application sharing that
server can exhaust it, lock it, or fill its disk, and the failure is the control plane going down
— which is the one failure that makes every other one harder to fix.
**They are bounded differently.** The substrate is sized, backed up and upgraded as part of
bootstrapping a mesh. A workload's database follows the workload — moved with it, destroyed with
it, restored with it.
## Consequences
**Migrating an object store is ordinary module work**, not substrate work. It was previously going
to be done as part of completing the substrate, which would have been the wrong shape and would
have coupled every mesh to a service the mesh does not use.
**A mesh with no workload needing one runs no object store at all.** That is the correct outcome
and was not previously available.
**Two PostgreSQL containers on a node that hosts both is expected**, not duplication to be
optimised away. Anyone tidying them together should find this record first.
**"Substrate by role and ordinary by delivery" loses one of its two members.** ADR 0006 uses that
phrase of the object store and the registry — things that are substrate but provisioned once a
control plane exists. It now describes the registry alone.
**The definition is applied, not just stated.** Both halves of the circularity test are asked of
each member, and *cannot grant itself one* is not sufficient on its own — it is true of almost any
service, which is what made it possible to admit a member on that half alone.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects one row of
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store exclusively
- `mesh-control internal/builder/registry.go` — where artifacts go, and why not S3
@@ -0,0 +1,112 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 29. A network is a shape, because an action cannot be undone
## Context
**A module of several containers has no way to let them reach each other by name.** A container
declaration carries a `network` field, and it only ever *joins* one that already exists — it was
added so the control plane could reach the store and the broker on the machine it was raised on.
Nothing in the vocabulary **creates** a network.
Without one, containers on a machine share the runtime's default bridge, which gives addresses and
no name resolution between them. So a module that is several containers can only be written by
publishing ports onto the machine and pointing its own parts at the host — which puts a module's
private wiring on the machine's own address space, where anything else on the machine can reach it
and any other module can collide with it.
**This is the gap a mail system meets and nothing else so far does**
([`00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) 3.3). It is being taken now
rather than then, because 3.3 is the task most likely to send work back into the declaration
language and the least useful place to discover it.
**Adding a shape is not a small change, and the host says so** — the vocabulary is asserted
against a stated number, with the reason written into the failure: *every addition widens what a
compromised control plane can express, so a change here is a decision.* The host applies what it
is told; the only bound on a hostile control plane is what the language can say
([ADR 0004](0004-a-node-and-how-it-joins.md)).
## Considered Options
1. **An `action` that creates the network.** The vocabulary already has one, the bundle already
uses seven of them, and `docker network create` with a `verify` is exactly the shape an action
takes. Nothing would need adding. **Rejected**, on removal:
> An action has no footprint the host can undo — it ran, and whatever it did belongs to
> whatever it acted on.
A network made this way **leaks when the module is unassigned**, and the mesh cannot tell: the
record says an action ran, and there is nothing to reverse. Unassigning a module would leave a
network behind on every machine it was ever on, and the only way to find them would be to go
and look. *A resource the mesh can create and never clean up is one it should not create.*
There is a second reason, and it is the one that generalises: an action is opaque. **The mesh
cannot tell what an action did**, so a network created by one is not a thing the mesh knows
about — it cannot be reported, counted, or reasoned about, and a module could not require one.
2. **Publish ports on the machine instead.** No new shape, and it works today. **Rejected.** It
makes a module's internal wiring part of the machine's address space: two modules that each
want a database on a fixed port collide, and anything else on the machine can reach what was
meant to be private. It also makes the module's manifest depend on what else is installed,
which is the thing provisioning exists to remove.
3. **`network` as a ninth shape.** **Adopted.**
## Decision
**`network` joins the vocabulary, and the vocabulary is nine shapes.**
```
{"id": "internal", "type": "network", "name": "mail"}
```
**A name and nothing else.** Not a driver, a subnet, an address range or a gateway: every one of
those is a thing a module would have to know about the machine it lands on, and a module that
names a subnet is a module that collides with whatever else chose the same one. The runtime picks;
the mesh names.
**It is created if absent and removed when no longer declared** — an ordinary shape, with the same
lifecycle as a directory. That is the whole reason it is a shape.
**Declared before the containers that join it.** Resources are applied in the order the module
wrote them, and orphans are removed in **reverse** — so a network written first is created first
and removed last, after the containers attached to it are gone. This is not a new rule; it is the
existing one, and it happens to be exactly right here. A network written *after* its containers
would fail to remove while they still hold it, and that failure is reported rather than silent.
**What it does not do:** it does not reach across machines. A network is one machine's, like
everything else the host applies. Modules on different machines reach each other over the private
network the mesh already provides ([ADR 0007](0007-connectivity.md)), and a shape that tried to
span machines would be a second overlay with a worse contract.
## Consequences
**The vocabulary is nine, and the count moves with a record.** The test that asserts it names this
one, so the next person to change it finds the argument rather than a number to edit.
**A compromised control plane can now create and destroy networks on a machine.** Stated plainly
because that is the cost, and the bound is the point: it can create a named network and remove
one, and it can do neither to anything it did not declare. It cannot inspect, attach to, or
reroute what is already there — those would be different shapes, and are not being added.
**A multi-container module becomes expressible**, which unblocks 3.3 and, less obviously, makes
several smaller modules simpler: anything that is a service plus a sidecar currently has to
publish a port to talk to itself.
**Nothing is required to use it.** A module of one container declares no network and joins none,
exactly as now. The substrate keeps using `host`, which is a runtime-provided network and not one
the mesh creates.
## References
- [ADR 0005](0005-the-node-host.md) — the host's vocabulary, and why each shape is a decision
- [ADR 0004](0004-a-node-and-how-it-joins.md) — what may be pushed is bounded by form, not by trust
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — 1.3,
and the mail system at 3.3 that this is for
@@ -1,101 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 29. The lab's first scenario has no pipeline, and the lab comes first
## Context
[The lab design](../03-DESIGN/01-to-be/01-end-to-end-testing.md) opens with *"what is under
test is a module; the mesh is the harness"*, and everything follows from that: a scenario has
its own forge, its own coordinator, and its own delivery cascade ending in verify. The verdict
*is* a pipeline result.
That is the right design for testing a module against the mesh that exists. It is unusable for
the thing now being built.
**The new mesh has no coordinator.** Tier 0 is a host binary and tier 1 is a pinned bundle
([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md)). A scenario that
requires a forge, a coordinator, a cascade and a meshware daemon cannot exercise them, because
all four are tier 2 and do not exist yet.
**And the sequencing was written backwards.** [Research 009](../01-RESEARCH/009-migration/00-overview.md)
placed the lab at phase B, as verification of tiers already built. But tier 0 is the component
that takes over a machine's packages, services and network — it cannot be developed against a
machine anyone needs. It needs somewhere disposable to exist **before** it is written, not
after.
## Considered options
1. **Develop tiers 0 and 1 against a real machine; add the lab afterwards.** Rejected twice
over. Developing something that reformats a machine, against a machine that is in use, is
how a machine is lost. And it would leave the bootstrap path exercised only when performed
for real — which is precisely the property that makes the current first-node script the
least-tested code in the system.
2. **Build the full lab first.** Impossible, not merely unwise: the full scenario needs a
coordinator, a forge and a delivery cascade, all of which are tier 2. It cannot precede the
tiers it is meant to test.
3. **Two scenario classes, the smaller one first, the larger a superset.** Chosen.
## Decision
The lab has **two scenario classes**, and the first has no pipeline in it at all.
| | **Bootstrap scenario** | **Full scenario** |
|---|---|---|
| Contains | one or more virtual machines, the host binary, a pinned substrate bundle | a complete mesh: forge, coordinator, delivery, modules |
| Verdict from | what the host reports about the state it reconciled | a pipeline result ending in verify |
| Exercises | tiers 0 and 1 | tiers 2 and above, and modules |
| Exists to | develop the mesh | test what runs on it |
The bootstrap scenario is a **strict subset** of the full one — the same virtualisation, the
same networking, the same scenario lifecycle, simply stopping before a control plane exists.
Nothing forks, which is the same rule the existing design already holds itself to.
**The lab is built first**, ahead of tier 0, and [research 009](../01-RESEARCH/009-migration/00-overview.md)
is resequenced accordingly. It is the environment everything else is developed inside.
Of the runner's two candidate jobs, this settles their order: **scenario lifecycle is needed
immediately** — something must materialise, snapshot and destroy a mesh before anything else
can be written. **Assertion execution comes later**, with the full scenario, because a
bootstrap scenario's assertions are about the state a single host reconciled and are small
enough to state directly.
## Consequences
- **The hardest path to test becomes the one exercised most.** Raising a node from nothing is
currently a script that runs when a node is created and is otherwise never touched. Under
this decision it is the inner development loop for every change to tiers 0 and 1.
- The first thing built is small: virtualisation, a network, a way to place a binary, and a way
to snapshot and reset. No forge, no coordinator, no pipeline, no modules.
- The full scenario becomes reachable by *addition* rather than by rework, because it differs
only in what is placed inside the machines.
- The lab acquires a second audience. It was designed for a module author and now also serves
whoever is building the mesh itself — which is the same "one runner, two callers" argument
the design already makes, extended one step.
- **A stale claim in the design is corrected.** It argues that scenarios are *"affordable with
system containers and would not be with virtual machines — the unit choice is what makes the
gate possible at all."* [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) superseded that:
a lab node is a virtual machine, and the scale argument for system containers was found to
have been invented rather than required. The design text did not follow the decision. It does
now.
- The lab's home is `novox/mesh-lab`, recorded in
[ADR 0030](0030-the-repository-structure.md) — written after this record, because this one
needed a repository that no decision had yet named.
- The bootstrap scenario's fidelity is its whole value, and also its risk: if it diverges from
how a real node is raised, it certifies something that does not happen. That is the same
hazard the existing design names for the full scenario, and the same answer applies —
nothing new drives it, and what runs is the real thing.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — a lab node is a virtual machine.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
observation this rests on: a scenario needing only tiers 0 and 1 is one machine and a pinned
bundle, which is also exactly the bootstrap path.
- [Research 009](../01-RESEARCH/009-migration/00-overview.md) — the migration sequence this
reorders.
@@ -0,0 +1,96 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 30. Data outlives the mesh that declared it
## Context
**The conversion runs on live services holding real data**, and starts on the node that holds all
of it. Identity, mail, everything. The requirement stated plainly: a data directory may be
*moved*, and may never be *lost*.
**The host deleted them.** A directory that stopped being declared was an orphan, and an orphan
directory was removed with `os.RemoveAll` — everything under it — while the report said
`removed`. A module unassigned took its database's files with it, and nothing anywhere said what
had been in there.
Reproduced before it was fixed: assign a module, let a service write into its directory, unassign
the module, and the file is gone.
**A directory stops being declared for ordinary reasons**, which is what makes this sharp rather
than theoretical. A module unassigned from a node. A manifest edited to move a data folder — the
exact operation the conversion needs. A resource renamed. A typo. **Every one of those is a normal
day's work, and every one of them was destructive.**
**The removal order was already right, and that is what makes a fix possible.** Everything the
mesh puts inside a directory is itself a declared resource, and orphans are removed in reverse
declaration order — so by the time a directory is reached, what the mesh wrote there is already
gone. Anything still present was put there by something else.
## Considered Options
1. **A `keep` flag on the directory.** A module declares which of its directories hold data, and
the host leaves those. **Rejected.** It is safe only when somebody remembered, and the failure
of forgetting is total and silent. A rule that protects data only when it was asked to is not
a rule about data, it is a rule about attentiveness — and this is the one place in the system
where being wrong does not recover.
2. **Never remove a directory.** Simple and unarguably safe. **Rejected**, narrowly: every module
ever assigned would leave its directories behind for ever, and a machine that accumulates
things nobody can account for is one where nobody can tell what is still in use. The clean-up
that is genuinely the mesh's is worth keeping.
3. **Remove a directory only when it is empty.** **Adopted.**
## Decision
**A directory that still holds anything is kept, and the mesh says so.** An empty one is removed.
**This is the host's existing line applied to the one shape where getting it wrong is
unrecoverable** — *it removes what it made and leaves what it merely configured*
([ADR 0005](0005-the-node-host.md)). An empty directory is what the host made. A full one is not.
**No flag, no declaration, nothing to remember.** Emptiness is the test, and it is derived from
the removal order rather than asserted: the mesh's own contents are gone by then, so what remains
is by definition something nobody declared.
**It is reported, not silent.** The outcome is `kept`, naming how many items are inside and saying
they are for a person to deal with. A directory quietly left behind is how a machine accumulates
things nobody can account for — which is the objection to option 2, and it is answered by saying
so rather than by deleting.
**Files are unchanged.** A declared file is the mesh's own — it wrote it, it owns it, and losing a
configuration file is not the failure this is about. The distinction is deliberate: **directories
hold what other things produced; files are what the mesh itself put there.**
## Consequences
**Moving a data directory is now safe by default.** The manifest changes, the old path stops being
declared, and the data stays where it is until somebody has looked at it. That was the operation
most likely to destroy something during the conversion, and it is now the operation that does the
least.
**Unassigning a module leaves its data.** Correct, and it means unassignment is no longer a way to
clean up — removing data is a person's act, done knowingly. Given what unassignment did before,
that is the trade being made and it is the right way round.
**A machine can accumulate directories nobody removed.** Accepted, and mitigated by saying so
every time rather than by a periodic sweep. A sweep would be the deletion this record exists to
prevent, on a timer, with nobody watching.
**It is not a backup, and must not be mistaken for one.** This stops the mesh destroying data. It
does nothing about a disk, a mistaken `rm`, or a service corrupting its own store. The conversion
still needs backups taken and **restored** before anything is moved — a backup nobody has restored
is a belief, not a copy.
## References
- [ADR 0005](0005-the-node-host.md) — the host removes what it made
- [`03-DESIGN/01-to-be/00-work-breakdown.md`](../03-DESIGN/01-to-be/00-work-breakdown.md) — the
conversion this was found by planning
@@ -1,99 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 30. The repository structure, and the rule that names them
## Context
The tiers are settled ([research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md))
and the product is named ([ADR 0027](0027-the-product-is-novox-mesh.md)), but the repositories
themselves were only ever sketched in research. Two consequences had already appeared.
[ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) makes the lab phase 0 of the entire
migration and **could not say where it lives**, because no record named a repository.
And the research contradicted an accepted record: it listed `mesh-hq` for this repository, while
[ADR 0028](0028-hq-is-company-scoped.md) had decided `novox/hq` and explicitly rejected that
name. A design resting on research is a design resting on something that can change without a
decision.
There is also an implied naming rule that has never been written down. ADR 0027 says
repository names take `mesh`; ADR 0028 gives this repository no prefix at all. Both are right,
for a reason neither states.
## Considered options
Only the naming rule had genuine alternatives; the tier repositories follow from the tiers.
1. **No prefix — `novox/host`, `novox/control`.** The organisation already says Novox, so the
prefix reads as stutter. Rejected once it was established that Novox delivers more than the
mesh: with several products the prefix is not stutter, it is the product namespace doing
real work, and the forge has no nested groups to do it instead.
2. **An organisation per product — `novox-mesh/host`.** Puts the product boundary where the
forge's only real grouping primitive lives, so permissions and teams attach to it. Rejected
for now as premature: no per-product access boundary exists yet, and it costs `novox-`
repeated across every organisation.
3. **Product-prefixed repositories in the company organisation.** Chosen.
## Decision
**The naming rule:** a repository that belongs to a product carries that product's prefix. A
repository that is company-scoped does not.
That is why this one is `hq` and the mesh's are `mesh-*`. Both records were already correct;
the rule connecting them is stated here.
**The repositories:**
| Repository | Tier | Holds |
|---|---|---|
| `novox/mesh-host` | 0 | the node host — the one binary installed by hand |
| `novox/mesh-substrate` | 1 | the four pinned services, as declarations |
| `novox/mesh-control` | 2 | the control plane and its contexts |
| `novox/mesh-surfaces` | 3 | tools, web, cli — thin, no logic |
| `novox/mesh-sdk` | — | contracts shared across tiers: types, not behaviour |
| `novox/mesh-lab` | — | the lab: scenario lifecycle, networking, placement |
| `novox/hq` | — | this repository. Company-scoped ([ADR 0028](0028-hq-is-company-scoped.md)) |
**The lab is its own repository.** Its lifecycle differs from everything else in the list: it
is never shipped to a node, it outlives any single tier, and it drives virtualisation on a
workstation — which nothing else in the mesh does. Putting it inside the host would couple
development tooling to a shipped component; putting it inside the control plane would make the
bootstrap scenario depend on a tier that does not exist when it is needed.
**Tier 4 is deliberately not decided here.** Whether the catalogue is one repository, one per
domain, or one per application remains open from
[ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) and is blocked on
[research 005](../01-RESEARCH/005-domain-grouping/00-overview.md): how many repositories hold
domains cannot be answered before what the domains are. Recording the gap is the point —
`mesh-catalog` appears in the research sketch and is **not** decided by this record.
## Consequences
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) can name its target. Phase 0 has
a home, which was the immediate blocker.
- The research sketch stops being load-bearing. It remains what it is — a sketch — and the
design layer can now cite a record instead.
- **Seven repositories where there is currently one**, for a mesh that today lives in a single
monorepo. That is the cost, and it is not small: seven release cadences, seven sets of
dependencies, and cross-repository changes that were previously one commit. The offsetting
argument is the tier rule — a boundary that only points downward is enforceable across
repositories and merely conventional inside one.
- The prefix will read as redundant for as long as the mesh is the only product with
repositories. That is accepted deliberately: the alternative is renaming everything at the
moment a second product appears, which is the class of migration this project is trying to
stop performing.
- Nothing is created yet. This records what the repositories *are*; creating them is part of
phase 0 and after.
## References
- [ADR 0027](0027-the-product-is-novox-mesh.md) — the product name the prefix comes from.
- [ADR 0028](0028-hq-is-company-scoped.md) — why this repository has no prefix.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lab, and why it is first.
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md) — the tiers, and the
sketch this supersedes as a source.
@@ -0,0 +1,72 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0006-the-substrate-and-the-control-plane.md
---
# 31. The control plane authenticates nobody, so identity is a module
## Context
[ADR 0006](0006-the-substrate-and-the-control-plane.md) left one member of the substrate
conditional, and said exactly why:
| role | product | |
|---|---|---|
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
[`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md) carried it as an open question —
*whether identity is the fifth* — noting it followed from a decision nobody had taken.
**The decision is taken: the control plane does not delegate authentication.** There is no mesh
identity provider.
**Nothing in the mesh's own machinery ever needed one.** A node proves itself with a keypair it
generated, over a broker account issued at enrolment
([ADR 0004](0004-a-node-and-how-it-joins.md)). Declarations are verified by signature. None of
that touches an identity provider, and the conditional was never about machines — it was only ever
about whether a *person* signing in to a mesh surface would be authenticated by something else.
## Decision
**Identity is a module**, like the mail system and the forge. It runs *on* the mesh, not *of* it
([ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)) — a provider other modules require,
which is the ordinary shape and needs nothing new to express.
**So the substrate is three, and no longer conditional**: a relational store, a message bus, and
an image registry. Together with
[ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md), which removed the
object store, the list is settled and every member is there for the same reason — the control
plane needs it and cannot ask itself for it.
**A mesh that wants no identity provider runs none.** That is now expressible, and was not while
it sat in the substrate as a maybe.
## Consequences
**The last open question about substrate membership is closed.** Both halves of ADR 0006's test
now have an answer for every candidate, and the answer for identity is *the control plane does not
need it*.
**It does not settle how a person signs in to a mesh surface**, and that is deliberately left
open. What is settled is that whatever answers it is not part of what must exist before the mesh
does — so it can be decided late, changed, or replaced, which is precisely what being substrate
would have prevented.
**It becomes a real test of the module graph.** An identity provider is a module that *other
modules require* — the object store already consumes it — so it exercises the provider chain more
seriously than anything ported so far, where the provider was written alongside its consumer.
**Ordering follows from it rather than from preference.** Anything requiring identity has to move
after it, which is a dependency the graph can state rather than something a person has to
remember.
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the conditional this closes
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the other member
removed, and the test applied properly
- [ADR 0004](0004-a-node-and-how-it-joins.md) — how a node proves itself, which needs none of this
@@ -1,87 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 31. The lab provides the underlay; the mesh builds the overlay
## Context
[Research 004](../01-RESEARCH/004-lab-network/analysis.md) worked out the topology a lab has to
reproduce: a routable segment using documentation addresses, a household segment behind NAT, a
router that forwards exactly one port so a *published-but-NATed* node is real, and a machine
that can attach to either segment or detach entirely.
It also records what makes that topology **mean** something, and this is where a boundary has
to be drawn. Hub election is by convention rather than by flag — the hub is the node whose
profile is server and whose overlay address begins `10.10.0.1`. Direct peering depends on two
nodes sharing a site. Names resolve from mesh configuration on each node.
Those are all facts the *mesh* establishes. The question is whether a scenario declares them.
It is tempting to say yes, because a scenario that hands you a working overlay is a scenario
you can start testing against immediately.
## Considered options
1. **The lab configures the overlay too** — assign the overlay addresses, elect the hub, write
the peer configuration, seed the names. Rejected, and the reason is the whole point of the
lab: **a lab that builds the overlay certifies its own work.** If the mesh's peering logic
is broken, a scenario that pre-built the peering still comes up green. The most valuable
thing the lab can test is precisely the part this would replace.
2. **The lab provides nothing but bare machines** — no addressing, no segments, no NAT. Also
rejected. Then the scenario cannot reproduce *published but behind NAT*, which research 004
identifies as the case that only exists in production today, and the lab loses its reason to
use virtual machines at all.
3. **The lab provides the underlay; the mesh builds the overlay.** Chosen.
## Decision
**A scenario declares the underlay** — the facts a machine would have before any of our
software touched it:
- which segments exist, and their address ranges
- which machine sits on which segment, at which address
- what NAT sits between them, and which ports are forwarded through it
- which machines are detached, and can be attached or detached during a run
**A scenario declares nothing about the overlay** — no overlay addresses, no hub, no peering,
no names, no certificates. Those are the mesh's job, and a scenario that supplied them would be
testing itself.
The rule stated in one line: **a scenario provides what a hosting provider and a home router
would provide, and nothing our software is responsible for.**
## Consequences
- **The overlay becomes a thing under test rather than a fixture.** Whether peers form,
whether the hub is elected, whether a NATed node's endpoint is learned — all of it is
observed rather than arranged. That is the class of fault research 004 says is discoverable
only in production today.
- The lab stays small, and stays honest. It needs to know about virtualisation, bridges,
addresses and NAT. It never needs to know what a mesh node is.
- A scenario cannot assert "the overlay came up" as a precondition, because it is an outcome.
A bootstrap scenario that wants a working overlay has to wait for one and check, which is
the correct shape.
- **The address ranges are load-bearing, not cosmetic.** The routable segment uses RFC 5737
documentation space specifically because the mesh's own code decides *public versus private*
by matching the address — a private range there makes the hub test as unreachable, and the
mesh silently never forms. Research 004 calls this the single most important fact in the
document, and the declaration format has to make getting it wrong hard.
- The router is a machine the lab materialises without being asked, because NAT requires
somewhere to run. That is an implicit machine in an otherwise explicit declaration, and it is
worth knowing about rather than discovering.
- **Host capability profiles are detected, not declared** — a consequence of
[research 006](../01-RESEARCH/006-mesh-from-scratch/code-skeleton.md), and consistent here: a
scenario does not say what a machine is allowed to do, it provides a machine. Which leaves an
open question: a lab machine is always privileged, so the `user` and `edge` profiles have no
scenario that exercises them yet.
## References
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology, the documentation
ranges, and the hub-election and peering conventions this deliberately does not touch.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the two scenario classes this
declaration has to serve without forking.
@@ -1,71 +0,0 @@
---
status: accepted
date: 2026-08-23
deciders: jochen
reconstructed: false
---
# 32. A scenario is an isolated address space, and the lab never reaches into it over IP
## Context
A scenario declares literal addresses —
[the declaration](../03-DESIGN/01-to-be/02-scenario-declaration.md) is full of them, and it has
to be, because reproducing *published but behind NAT* means saying which address the world sees.
That raises a question the declaration left open: **two scenarios at once.** Several agents
working means several scenarios, and the lab design already calls that a requirement. But two
scenarios built from the same declaration want the same addresses, and there are only three
documentation ranges in existence.
## Considered options
1. **Allocate addresses from a pool at raise time**, rewriting the declaration's literals.
Rejected. It makes the addresses in a declaration a fiction, so a scenario reproducing a
specific topology no longer reproduces it; it breaks the RFC-range validation, since
allocated addresses would have to come from somewhere real; and the numbers a person reads
in the file stop being the numbers they will see in a capture.
2. **One scenario at a time.** Rejected — it is the requirement, not an inconvenience. A gate
an agent has to queue for is a gate that gets bypassed.
3. **Give each scenario its own network stack, so the addresses do not collide.** Chosen.
## Decision
**A scenario is a closed address space.** Every segment materialises as its own isolated link,
belonging to one scenario instance. Two scenarios raised from the same declaration hold the same
addresses and never meet, because nothing joins their links.
The declaration therefore keeps its literal addresses, and they mean exactly what they say.
**The consequence that constrains everything else: the lab never reaches into a scenario over
IP.** It talks to a machine through the virtualisation layer's own channel — the same way one
executes a command in a container without the container being routable.
That is not a preference. If the lab reached machines by address, the workstation running it
would need a route into each scenario, and two scenarios carrying the same prefix would give it
two routes to the same destination. Concurrency would be impossible, and it would fail in the
worst available way: not with an error, but by one scenario's traffic arriving in another.
## Consequences
- Scenarios are concurrent by construction, with no allocation, no bookkeeping and no limit
beyond the machine's capacity.
- The three documentation ranges stop being a scarce resource. Every scenario may use all of
them, because no two scenarios share a link.
- **The lab cannot use IP to check anything**, which is more of a constraint than it first
appears: *"can this machine reach that one"* has to be asked **from inside the scenario**, by
executing on a machine, rather than probed from outside. That is the honest way to ask it
anyway — reachability from the workstation is not the question.
- A scenario is a unit that can be paused, snapshotted and destroyed whole, because nothing
outside holds a reference into it.
- The lab needs a scenario **instance** identity distinct from the scenario name in the
declaration: the declaration is a kind, and several instances of one kind may exist.
- **The workstation is not on the scenario's network, so it is not a node in it.** Anything a
developer wants to reach — a web interface, a database — needs an explicit, deliberate
forward out of the scenario, which is a feature rather than a gap: nothing leaks by default.
## References
- [ADR 0031](0031-the-lab-provides-the-underlay.md) — the declaration whose literal addresses
this preserves.
- [ADR 0029](0029-the-labs-first-scenario-has-no-pipeline.md) — the lifecycle jobs this shapes.
@@ -0,0 +1,79 @@
---
topic: how we work
status: superseded
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0031-the-control-plane-authenticates-nobody.md
superseded-by: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 32. The local account owns the mesh; a surface delegates to a module
## Context
[ADR 0031](0031-the-control-plane-authenticates-nobody.md) settled that the control plane
authenticates nobody, and deliberately left one thing open: **how a person signing in to a mesh
surface is authenticated.** This answers it, and answers a question 0031 did not ask — *who owns
the mesh at all.*
**There was no answer, and the absence was invisible** because every operation so far has been run
by the person sitting at the machine. Nothing had to say whether that was the design or the
circumstance.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login,
and there is nothing else to hold.
**No mesh user model.** No accounts, no roles, no grants, nothing to administer. A person with a
shell on a node can do anything the mesh can do there, because that is already true and pretending
otherwise would be a boundary that does not exist.
**This follows from what was already decided rather than adding to it.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is therefore the security boundary*. A user model inside that boundary would guard nothing: anyone
who could be stopped by it could equally read the node's key off the disk.
**The board is different, and the difference is the network.** A surface reachable by a browser
has to know who is asking, because the people reaching it are not, by construction, people with a
shell on the machine. **So the board delegates to an OAuth provider** — which is a module.
## What this does not change
**The identity provider is still not substrate** (ADR 0031). A *surface* delegating
authentication is not *the control plane* delegating it. The control plane runs, applies
declarations and reaches nodes with no identity provider in existence; only the board needs one,
and only to decide whose browser it is talking to.
The test is unchanged and still answers no: *does the control plane need it in order to run?*
## Consequences
**The board depends on a module, and says so.** An ordinary edge in the graph, which means the
board cannot come up before the provider it authenticates against — stated as a dependency rather
than discovered as an outage.
**Moving the identity provider takes the board with it.** During that module's own conversion the
board is unavailable, and that is acceptable: it is a surface, nothing depends on it, and a brief
interruption is the trade already accepted everywhere else. Nothing that keeps a service serving
goes through it.
**Anyone with a shell on a node has full authority there.** Written down rather than left implied,
because it is the sentence that decides who gets an account on a machine. The protection is the
machine's own login, and the overlay that keeps the machine unreachable from outside
([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** Deliberate, and the cost of
having no user model: there is no way to give a person authority over one node without giving them
a shell there. If that is ever wanted, it is a new decision and not a gap in this one.
## References
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody; this answers what it left open
- [ADR 0004](0004-a-node-and-how-it-joins.md) — no authorisation between nodes, and why the mesh
boundary is the security boundary
- [`03-DESIGN/01-to-be/11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — the surface this is
about
@@ -1,81 +0,0 @@
---
status: accepted
date: 2026-08-24
deciders: jochen
reconstructed: false
extends: 0016-a-lab-node-is-a-virtual-machine.md
---
# 33. A router is scenery, not a node — so it is a container
## Context
[ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) settles that **a lab node is a virtual
machine**, and its reasoning is fidelity: a node boots a stock image and runs the real install,
so it has to be a real machine or the thing under test is not the thing that ships.
A scenario also needs routers. NAT, port forwarding, policy between segments and mapping
expiry are all things a router does, and until one is materialised a multi-segment scenario
raises isolated islands
([03-DESIGN/01-to-be/02-scenario-declaration.md](../03-DESIGN/01-to-be/02-scenario-declaration.md)).
The declaration already implies them: a gateway is *the one implicit machine in an otherwise
explicit declaration*.
The question is whether ADR 0016 binds those too.
## Considered options
1. **A router is a node, so it is a virtual machine.** Consistent, and pays for a consistency
nobody needs. A router boots in roughly ten seconds against a container's one; a
six-segment scenario wanting three routers spends thirty seconds per raise on scenery.
2. **The hypervisor provides NAT** — bridges with translation switched on, and its own
forwarding primitives. Rejected on a stronger ground than speed: it makes the *lab* provide
what the declaration is supposed to own, and it cannot express a mapping that expires, a
gateway that refuses to forward, or policy between siblings. The model would shrink to fit
the tool.
3. **A router is scenery, and scenery is a container.** Chosen.
## Decision
**ADR 0016 binds nodes. A router is not a node.**
Nothing under test runs on a router. It is not a participant, it holds no identity, the mesh
never installs anything on it, and no assertion is ever made about its internals. It exists so
that packets between machines behave the way they behave in the world — which is the definition
of scenery.
So a router is a **system container**, and the fidelity argument does not reach it: what a
router must reproduce is kernel behaviour — translation, connection tracking, filtering,
forwarding — and a container has the same kernel.
**Verified before deciding, not assumed.** In a plain unprivileged container:
| Needed for | Works |
|---|---|
| routing at all | `net.ipv4.ip_forward`, `net.ipv6.conf.all.forwarding` |
| `nat:` | nftables masquerade, rules accepted and listed back |
| `mapping_ttl:` | `nf_conntrack_udp_timeout`, `nf_conntrack_tcp_timeout_established` |
No privileged mode, no nesting, no capability grants.
## Consequences
- A raise stops paying a boot per router. Scenery costs about a second where a node costs ten,
and a scenario's cost tracks the machines actually under test.
- **The distinction is now load-bearing and has to stay legible.** *Node* means something under
test; *scenery* means something that makes the test real. If anything is ever installed on a
router by the mesh, it has become a node and this decision no longer covers it.
- Routers and nodes are different kinds of thing in the lab's own model, which is a small extra
concept — justified by it being true, rather than by the saving.
- A container shares the host kernel, so a scenario cannot reproduce a router running a
*different* kernel from the workstation. Nothing currently wants that; if something does, that
router becomes a virtual machine and this record needs revisiting rather than bending.
- The gateway stays implicit in the declaration. A scenario declares `gateway:` on a segment and
never names the machine that serves it — which is right, because it is not a machine the
scenario has anything to say about.
## References
- [ADR 0016](0016-a-lab-node-is-a-virtual-machine.md) — what a lab *node* is, unchanged.
- [Research 004](../01-RESEARCH/004-lab-network/analysis.md) — the topology needing a router,
and why *published but behind NAT* only exists in production today.
@@ -0,0 +1,91 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0028-the-substrate-supplies-the-control-plane-and-nothing-else.md
---
# 33. The substrate is a store and a broker
## Context
Third correction to one table in one day, all found the same way: by asking whether **both** halves
of the substrate test were actually answered for a given member, or only the second.
The test ([ADR 0006](0006-the-substrate-and-the-control-plane.md)) is *what the control plane needs
in order to run, and cannot ask itself for, because it is not running yet.* ADR 0006 admits the
image registry on this line:
| role | product | |
|---|---|---|
| image registry | **an OCI registry** | it cannot grant itself a repository |
**That is the second half again.** It is true that a control plane cannot grant itself a
repository. Nothing establishes that it needs one *in order to run*.
**Counted rather than argued.** `substrate-first-node.lock` — the only bundle there is, and what a
first node actually becomes — raises twelve resources, and no registry is among them:
```
container runtime · the store · one database per context · the schemas
· the broker's certificate · the broker · the control plane
```
The registry arrives afterwards, as an ordinary module the mesh assigns. That is what the lab
asserts, in those words: *the mesh runs its own artifact store.*
**ADR 0006 half-said this already**, calling the registry *substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.* A member that is provisioned by the
thing it supposedly precedes is not a member; the phrase was carrying a contradiction rather than
resolving one.
**The registry is a closer call than the object store, and the difference is worth keeping.** The
control plane never touches an object store at all — no client, no bucket, ever
([ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md)). It genuinely
*uses* the registry: the builder pushes to it, hosts pull from it, and nothing reaches a machine
without it. **So the registry is a real dependency of the mesh operating, and not of the control
plane starting** — and it is the second that the word substrate means.
## Decision
**The substrate is two things: a relational store and a message bus.** Both are in the bundle,
both must exist before the control plane's first instruction, and neither can be asked for.
**The registry is an ordinary module.** The mesh cannot deliver anything without one, and it
installs one the way it installs everything else. The first node's chicken-and-egg is already
solved and needs nothing from this list: it fetches upstream images directly, then runs a registry
of the mesh's own.
**The test is applied to both columns, every time.** *Cannot grant itself one* is true of almost
any service and settles nothing on its own. It is what admitted the object store, and then the
registry, and both were removed by asking the other question.
## Consequences
**The substrate is now exactly what the bundle raises**, which is the strongest form this list can
take: it can be checked by counting rather than by reading an argument. A member that is not in
the bundle is not substrate, and the two statements cannot drift apart.
**A mesh that builds nothing still needs a registry** — to receive anything at all — but it needs
it as a module, on its own schedule, replaceable. That was already true and was obscured by the
list.
**The word may now be doing too little work.** "Substrate" for *a database and a broker* is a term
of art for two things everybody can name. Renaming is not taken here and is worth considering
separately; what this record fixes is the membership, not the vocabulary.
**Three removals from one table in one day is itself the finding.** Each member was admitted on the
half of the test that is easy to answer, and the design read plausibly throughout. The rule that
comes out of it is not about substrates: **a test with two conditions is a test only when both are
asked.**
## References
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — the definition, and the table this
corrects a second row of
- [ADR 0028](0028-the-substrate-supplies-the-control-plane-and-nothing-else.md) — the object
store, removed for the same reason
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — identity, which was conditional and
is now a module
@@ -0,0 +1,81 @@
---
topic: how we work
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0032-the-local-account-owns-the-mesh.md
---
# 34. The local account owns the mesh, and a web application's login is not that
*Supersedes [ADR 0032](0032-the-local-account-owns-the-mesh.md), which decided the right thing and
described it wrongly. The decision below is unchanged; what it said about the board was an
invention.*
## Context
ADR 0032 answered *who owns the mesh* — the account that installed the host — and then framed the
board as **a surface that delegates authentication**, a category it made up for the occasion. It
does not need one.
**The board is a web application.** It has a login, provided by the identity module, in the way
every web application has a login. That is a fact about an application, not a property of the
mesh, and giving it a name in the mesh's vocabulary implied a relationship that is not there.
The cost of the invented category was not cosmetic. It made the identity module look like part of
the mesh's own authority — something the mesh *depends on* to know who anybody is — when the truth
is that the mesh knows nothing about people at all, and one of the applications running on it has
a login.
## Decision
**The account that installed the host owns the mesh on that node.** Authority is a local login.
There is nothing else to hold, no user model, no roles, and nothing to administer.
**This follows from what was already decided.**
[ADR 0004](0004-a-node-and-how-it-joins.md) says there is no authorisation between nodes — every
node is the operator's own, so a message from one is a message from them, and *the mesh boundary
is the security boundary.* A user model inside that boundary would guard nothing: anyone it could
stop could read the node's key off the disk.
**A web application's login is its own business.** The board authenticates its users through the
identity module. So might anything else the mesh runs. **None of that is mesh authority**, and the
mesh does not learn who anybody is from it.
## The line this draws, which is the reason to write it down
**Signing in to an application must not, on its own, become authority over the mesh.**
Today it cannot: the board reads and does not act
([`11-a-board.md`](../03-DESIGN/01-to-be/11-a-board.md) — *not the way to change things*). Looking
at a page tells you what is true and changes nothing.
**The moment the board can assign a module, whoever it lets in has mesh authority** — and it would
arrive as a feature rather than as a decision. That is the failure this record exists to make
visible, because it is the kind that is only obvious afterwards.
So: **a surface that can change the mesh is a change to who owns the mesh**, and is taken as one.
Not forbidden — wanting to manage nodes from a browser is reasonable — but not something that
turns up in a pull request titled *add assign button*.
## Consequences
**The identity module is not special.** Not substrate ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)),
not part of the mesh's authority, and nothing about the mesh stops working when it is down. Some
applications cannot be logged into, which is what it means for an application's login provider to
be unavailable.
**Anyone with a shell on a node has full authority there.** Unchanged from ADR 0032, and still the
sentence that decides who gets an account on a machine. The protection is the machine's own login
and the overlay that keeps it unreachable from outside ([ADR 0007](0007-connectivity.md)).
**A node cannot be operated by somebody without a login on it.** The cost of having no user model.
If that is ever wanted, the paragraph above says what it costs.
## References
- [ADR 0032](0032-the-local-account-owns-the-mesh.md) — superseded; same decision, invented category
- [ADR 0004](0004-a-node-and-how-it-joins.md) — the mesh boundary is the security boundary
- [ADR 0031](0031-the-control-plane-authenticates-nobody.md) — the control plane authenticates
nobody
@@ -0,0 +1,115 @@
---
topic: what runs on it
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0034-the-local-account-owns-the-mesh.md
---
# 35. One implementation, several surfaces, and what that costs
## Context
The mesh is operated from a command line today. It needs to be operable from a browser and from a
model's tools as well, and the three must not be three different systems.
**The pattern is already in the code and unnamed.** `board` serves HTTP by calling the same
functions the CLI calls; it holds nothing and decides nothing. What follows makes that the rule
rather than a property of one command.
**The board is a presentation layer over the control plane.** Not an application beside it holding
a database credential — the thing that shows what the control plane knows, and asks it to do what
a person asked for.
## Decision
**The logic lives once, in the context that owns it. A surface is an adapter with no decisions in
it.**
| surface | for |
|---|---|
| **command line** | a person on a machine, and the recovery path below |
| **HTTP** | the board, and anything else that speaks to the mesh over a network |
| **model tools** | an agent asking the mesh to do something |
**Every surface refuses identically, because the refusal is not in the surface.** An assignment
that cannot be satisfied is refused by the same resolution whichever way it arrived. The moment a
surface can accept something another would reject, the mesh has two answers to one question and
people learn which to trust.
**Reading and doing are both exposed.** The HTTP surface is not read-only: managing the mesh from
a browser is the point. This takes the decision
[ADR 0034](0034-the-local-account-owns-the-mesh.md) said had to be taken deliberately —
**a browser login now carries authority over the mesh** — and takes it knowingly rather than
letting it arrive with a feature.
**The networked surfaces authenticate through an OAuth2 identity provider.** Named by protocol
rather than by product, like every other dependency the mesh takes — AMQP for the bus, S3 for an
object store, OCI for the registry
([ADR 0006](0006-the-substrate-and-the-control-plane.md)). What fills the role today is a module
running Keycloak; what the control plane knows is that it validates a token against a provider
speaking OAuth2, and replacing that provider is a migration rather than a redesign.
**The command line does not authenticate at all**: it is already behind the machine's own login,
which is what owns the mesh (ADR 0034).
## What this is not: a kernel every module imports
**The shared library is the failure this project was started over**, and the difference has to be
stated or it will be rebuilt. The old one is 155 files and 34,636 lines *containing code from
every context* — work-domain logic sitting in the kernel every module imports, each piece landing
there to avoid a cycle between two modules that both needed it.
**Shared surfaces are not a shared library.** What is shared here is that three adapters call the
same functions. Those functions stay in the context that owns them — provisioning's logic in
provisioning, identity's in identity — and no module imports another's. A surface may call many
contexts; a context still may not reach into another's store
([ADR 0008](0008-a-context-owns-its-store.md)).
The test, when something is about to be put "somewhere shared": *does this belong to a context, or
does it only belong to the surface?* If it belongs to a context it goes there, even if two
surfaces want it.
## The loop this creates, and the way out
**The control plane's networked surfaces will depend on a module the control plane assigns.**
An identity provider is an ordinary module ([ADR 0031](0031-the-control-plane-authenticates-nobody.md)). When
it is down, or being migrated, or misconfigured, the HTTP and tool surfaces cannot authenticate
anybody — including the person trying to fix it.
**The command line is the way out, and it is why local ownership matters more rather than less.**
It authenticates through nothing, needs no network, and is available on the machine to the account
that owns the mesh. **A mesh must always be operable by somebody standing at it.**
So the rule: **no capability exists only behind an authenticated surface.** Anything the board can
do, the command line can do. That is not a courtesy to CLI users; it is the recovery path, and a
capability that exists only over HTTP is one that disappears exactly when identity does.
## Consequences
**Identity is still not substrate**, and the test still answers no: the control plane runs, applies
declarations and reaches nodes with no identity provider in existence
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). What is unavailable without it is two
surfaces, not the mesh.
**Whoever the identity provider admits has authority over the mesh.** That is now a real perimeter
with real consequences, where before it guarded a page that only read. Who may log in, and to
which realm, becomes a decision about the mesh rather than about an application.
**A surface must not grow an opinion.** The likely erosion is a validation added to the board
because it was quicker there — and then the CLI accepts something the board rejects, or worse the
reverse. Adapters hold no decisions.
**Three surfaces over one implementation is a cost paid three times if it is not one
implementation.** The reason to write this down now is that the second surface is the cheapest
moment to get it right, and the third is where the drift usually starts.
## References
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh, and the
line this record deliberately crosses
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the shared library this must not
become
- [ADR 0008](0008-a-context-owns-its-store.md) — a context owns its store, which a surface does
not change
@@ -1,75 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
---
# 36. A node is a managed machine, and disconnection is a situation
## Context
[Research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) left open: *"does an
unprivileged node earn a place in the inventory, or only a presence? Decides whether 'node'
means one thing or two."*
The question came from requirement 6 — *Arch Linux only for now; ideally any device, including
phones, on lighter terms* — and from the observation that some machines cannot be fully
managed. A phone will not run the host. A laptop is absent for days.
The question assumed the answer was a **class**: full nodes and lesser ones, with the
inventory recording the first and merely acknowledging the second.
## Considered options
1. **Two classes — nodes and presences.** An unprivileged device gets a lighter record and a
reduced contract. Rejected: it makes "node" mean two things, so every context that reasons
about nodes acquires a branch, and the branch is invisible in the type. The mesh already has
one instance of this shape and it is the one this repository keeps writing issues about —
a declared thing that is only sometimes honoured.
2. **One class, membership by capability.** Everything is a node; what it can do is a property.
Chosen.
## Decision
**A node is a managed machine inside the mesh.** Not a device that is merely known about, not
an unprivileged something. If the mesh does not manage it, it is not a node — it is a client, a
peer, or a thing on the network, and those want their own names rather than a weakened version
of this one.
**A disconnected node is still a node, in a different situation.** Reachability is state, not
class. A node that is switched off, roaming, or behind a connection that has dropped has not
become a lesser kind of thing; it has a last-known state and a pending set of declarations.
The distinction the original question reached for is real, but it is **capability**, not kind —
what this machine can be asked to do — and that belongs in the host's profile, not in the
definition of a node.
## Consequences
- **The inventory has one shape.** No branch, no second record type, no context that must ask
which kind it is holding.
- **Local state is structural, not a convenience.** If disconnection is an ordinary situation
rather than an exception, the host's store is authoritative while disconnected by design —
it is what makes the situation ordinary. This promotes `store/` from a component to a
requirement.
- **Absence is not failure.** A node that has not been seen is in a state, and the mesh must be
able to say which. Anything that treats unreachable as broken will be wrong most of the time
about a laptop.
- **Devices that cannot be managed do not become nodes by being lenient about the word.** A
phone that cannot run the host is not a node under this record. Whether the mesh should reach
such devices at all, and as what, is not decided here and needs its own record if it is
wanted.
- **The reduced-contract idea is not lost, it is relocated.** What a given node can be asked to
do is its profile — the host's capability detection — and varies per machine without varying
what a node is.
## References
- [Research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) — the open question, and
requirement 6 that raised it.
- [ADR 0015](0015-mesh-brokers-nodes-host-agents-think.md) — *nodes host*; this says what a
node is.
- [Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) —
capability as something detected rather than assumed, which is where the reduced contract
now lives.
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-08-31
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0035-one-implementation-several-surfaces.md
---
# 36. Bootstrap ends at a usable mesh, and the first credential comes from a person
## Context
Bootstrap currently ends when the control plane starts
([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). That is a mesh that runs and
cannot yet be used by anybody who is not standing at the machine: the networked surfaces need an
OAuth2 identity provider ([ADR 0035](0035-one-implementation-several-surfaces.md)), the provider is
a module, and no module has been assigned.
**So bootstrap should go further** — through the identity provider and the first login — and stop
at a mesh somebody can actually use.
**One thing in the way, and it is not incidental.** The mesh has never held a readable secret. The
sealing code says what it does and why:
> Make generates a secret and seals it to both ends, **keeping no readable copy.**
An initial administrator's credential is the first value a **person must read**. Everything else
the mesh generates is something no human ever sees, and everything a human provides is something
the mesh immediately stops being able to read.
## Considered Options
1. **The mesh generates it and prints it once**, to the terminal of whoever ran the bootstrap.
Convenient, and needs no prompt. **Rejected.** It would give the control plane a plaintext
secret for the first time — briefly, and only to one terminal, but the capability would then
exist. *An exception made for one case does not stay one*: the next credential that is awkward
to supply gets printed too, and the property that a copy of the mesh's database is a copy of
nothing stops being checkable by reading the code.
2. **No password: a one-time link that lets the operator set their own.** The nicest to use.
**Rejected for now** — it needs a mechanism that does not exist, and the thing it improves is
one prompt, once, on a new mesh.
3. **The operator supplies it.** **Adopted.**
## Decision
**Bootstrap runs to a usable mesh**: the substrate, the control plane, the identity provider as an
ordinary module, its realm and client provisioned, an administrator able to log in, and the
networked surfaces available.
**The administrator's credential is supplied by the person doing the bootstrap**, on standard
input and not echoed — the path that already exists for a model-access key. The mesh seals it and
cannot read it afterwards.
**What is created is an account in the identity provider, not a user of the mesh.** The mesh still
has no user model and gains none here ([ADR 0034](0034-the-local-account-owns-the-mesh.md)). What
this produces is the first login for the applications that have one.
**The provisioning is ordinary.** A realm, a client and a first account are what an identity
module's provisioner makes from what the mesh granted it — the same shape as a database and a
bucket, which are built and proven.
**The surfaces arrive when their dependency does.** The command API is not started with the
control plane and then broken until identity exists; it becomes available once it can authenticate,
the way anything else waits for a provider.
## Consequences
**The mesh still never holds a readable secret**, and that sentence needs no exception clause.
That is the whole reason for the prompt.
**An unattended bootstrap is still possible, and the value still comes from outside.** Automation
supplying the credential is the operator supplying it. What is refused is the *mesh inventing*
one — so an unattended bootstrap with no credential provided produces a mesh with no
administrator, which is correct rather than broken.
**Bootstrap gains an interactive step**, and it is the only one. Worth stating because a bootstrap
that cannot run without a person is a real constraint on how a node is stood up, and this is
deliberate rather than an oversight.
**The identity provider is still not substrate.** It is assigned by the control plane, so it comes
after it, and a thing that comes after cannot be a thing that must exist before
([ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)). Bootstrap running through it does not
move it: bootstrap is a sequence, the substrate is a dependency.
**And the recovery path is unchanged.** When the identity provider is broken later — which is the
failure that matters, not the one at first start — the command line still works, because it
authenticates through nothing (ADR 0035).
## References
- [ADR 0035](0035-one-implementation-several-surfaces.md) — the surfaces, and why the command line
must keep working
- [ADR 0034](0034-the-local-account-owns-the-mesh.md) — the local account owns the mesh; this adds
no user model
- [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) — what must exist before the control
plane, which this does not change
@@ -1,97 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
---
# 37. The host applies; it does not decide
## Context
The skeleton absorbs overlay membership, packet filtering, package management, service
supervision, the container runtime and filesystem management into tier 0, and
[research 006](../01-RESEARCH/006-mesh-from-scratch/00-overview.md) called this *"the
skeleton's biggest unproven claim. A binary whose whole argument is that it has no dependencies
now carries six concerns."*
That claim has now been measured against the monorepo's `main`:
[`host-size.md`](../01-RESEARCH/006-mesh-from-scratch/host-size.md).
The measurement says the question asked about the wrong axis.
## Considered options
1. **Absorb the six concerns as they are.** What the skeleton literally proposes. Rejected on
evidence: two of the ten modules implementing them open a direct connection to the control
plane's database and compute their own configuration. Absorbing those unchanged puts a
Postgres client and knowledge of the mesh schema inside tier 0 — an upward dependency, and
the tier rule is the whole of the bootstrap argument.
2. **Leave them as modules.** Keeps the tier rule trivially, and keeps the fault that prompted
the skeleton: four modules constituting *how a node is reachable* with no relationship the
mesh can see, so one intent is expressed four times
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md) finding 4 measures this and
finds it is the only place in the catalogue where the shape genuinely occurs).
3. **Split each concern: decide centrally, apply locally.** Chosen.
## Decision
**The host has one concern: apply declared state on this machine.** The six are not six
concerns it carries; they are instances of the one.
Each divides:
- **Deciding** — what this node's overlay, names, exposure, filtering, packages and services
*should be*. This needs every other node, and belongs to the control plane.
- **Applying** — putting that on the machine. This needs root and locality, and belongs to the
host.
**The host never queries the mesh database.** A host that reads the control plane's schema is
tier 0 depending on tier 2, and the tiers stop being a bootstrap answer the moment that is
permitted once.
## Why the evidence supports it
**Size was the wrong worry.** The ten modules total 2 755 lines. The machinery that already
applies state on a node — `meshware`, `env-sync`, `config-sync` — is 3 059. Everything being
absorbed is smaller than what already exists to apply it. The host is not a new large thing; it
already exists, spread across three core modules and unnamed.
**Eight of the ten are already pure appliers.** They receive derived state and put it on the
machine. Absorbing them moves code that has no dependency to move.
**The split has already been happening, unnamed.** `dnsmasq-app` needs the same mesh-wide data
as `wireguard` and does not query for it. Its own comments record why: the values were
*"duplicated by hand on all four nodes"* until someone derived them centrally, after a rename
meant editing four override rows nobody knew about. That is this decision, reached once by
fixing a bug.
## Consequences
- **Two modules must be split before they can be absorbed**, and they are the two hardest.
`wireguard` needs every node's key, address, site and endpoint reachability; `traefik` needs
certificates and every node's exposed names. The measurement says the design is right; it
does not say the migration is cheap, and this record does not claim it is.
- **The overlay and firewall modules stop existing** as the skeleton says — but the reason is
now sharper than "they are host concerns". The host holds membership and applies filtering;
the control plane decides policy; swappable backends stay modules.
- **Six vocabularies remain.** Zero dependencies, but the host must still know what a WireGuard
peer, an nftables rule, a package, a unit, a container and a dataset *are*. That surface is
the residue of the original worry and is not measured by anything here.
- **A dependency-direction lint is now load-bearing**, not a nicety. This record is a rule
about direction, and per this repository's own standard a rule states how it is checked: an
upward import fails the build. A tier rule enforced by intention is the same as no tier rule.
- **What the host carries versus what it finds is still open.**
[Issue 007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md) — the
host manages `wg`, `nft`, `pacman`, `docker`; it does not contain them, and *installed* is
not the same as *usable*.
## References
- [`host-size.md`](../01-RESEARCH/006-mesh-from-scratch/host-size.md) — the measurement.
- [Research 005](../01-RESEARCH/005-domain-grouping/analysis.md) — reachability as the only
measured co-change cluster in the catalogue.
- [ADR 0003](0003-the-mesh-database-is-the-source-of-truth.md) — what the control plane decides
from.
- [ADR 0030](0030-the-repository-structure.md) — `mesh-host` as tier 0.
- [ADR 0008](0008-a-failed-step-fails-the-job.md) — the standard the direction lint is held to.
+103
View File
@@ -0,0 +1,103 @@
---
topic: building it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 37. Where a module lives
## The question
The mesh's own module descriptions currently sit in `examples/` inside the control plane, beside
the small programs that hand out logins. That was fine while there were three of them. It is
wrong now, and the name is doing active harm: everything in `examples/` reads as a sketch, and one
of them shipped naming a container image that nothing in the repository builds. A directory called
*the catalogue* would have made *does this actually work* the obvious question to ask of it.
So: **one repository holding the modules we ship?** And if so, where does everything that is not
ours go?
## What a module actually is, counted
The system being replaced has **126 modules** on its main branch. The shape of them is the whole
argument, so it is measured rather than assumed:
| | count | what it is |
|---|---|---|
| **the module is software** | 47 | its own source tree lives inside the module — a daemon, a service, a library |
| **helper scripts only** | 44 | no application of its own; scripts it runs at install time or offers to an agent |
| **a description and nothing else** | 35 | a package to install and some files to write |
**Two thirds of modules contain code.** The largest is a shared library of 182 source files. A
speech-capture module carries a complete daemon — audio capture, mixing, transcription, a model
runner. Treating a module as *a description of something else* is true of barely a quarter of them.
That kills the simplest answer. A catalogue cannot be "a folder of manifests" when most modules
are programs.
## The four kinds, which want different homes
**1. What the mesh is made of.** The control plane, the host, the shared library, the board.
These are not modules that happen to be ours; they are the mesh, expressed as modules so it can
install itself. They belong in the repositories that build them, which already exist.
**2. Something the world made, that we describe.** A forge, a mail system, an identity provider,
a media server. Nobody upstream ships a description; somebody has to write one, and it is the same
description for everybody who runs it. **This is what a catalogue is for.** It is also where the
small programs that create accounts belong, because such a program is part of describing that
service, not part of the mesh.
**3. Something we wrote, that runs somewhere.** An application, a site, a side project. The
description belongs **with the code, at the root of its own repository**, because the two change in
the same commit. A repository that gains an environment variable and a description that gains it
elsewhere will drift, and there is no mechanism that could stop it. This is already how it works
and it should stay that way.
**4. A package and some files.** A tool, a font, a shell. Thirty-five of these, and each is a few
lines. The catalogue.
## The proposal
**A `mesh-catalog` repository** holding kinds 2 and 4: descriptions of software we did not write,
and the programs that provision it. Not kind 1, which is the mesh itself. Not kind 3, which lives
with its own code.
**The mesh's list of modules is not this repository.** It is a table in the control plane, filled
by adding a description to a running mesh. The catalogue is a *source* to add from — one of
several, and the mesh already records which: every module carries where it came from, the branch
followed there, and the commit its description was read at. **Nothing needs inventing to support
modules from anywhere**; a repository of our own is simply the source we curate.
**A description is checked by the tool, not by a test that imports the tool.** Today a test in the
control plane parses the example manifests by reaching into the control plane's internals, and
another reads the control plane's own build file to check every image a module names can be built.
Two jobs tangled. A `module check` command on the control plane's binary would let the catalogue
hold data validated from outside, and would give the same check to somebody describing their own
application in their own repository — which is the case that matters most and currently has no
check at all.
## What this costs, and the argument against
**It is early.** Ten modules exist, four of them ours. Moving ten files is a morning; moving a
hundred is a week — but the hundred is not here yet, and splitting now adds a second repository to
release across before there is anything to release.
The counter is that the tangle is already producing faults rather than merely threatening to. A
manifest naming an unbuildable image, and a test reading a build file two directories up, are both
symptoms of one repository doing two jobs. And the moment the first module is adopted on a real
machine, the descriptions stop being examples and become the thing deployments come from. **That
is the moment this becomes urgent, and it is close.**
## What it does not settle
**Where a provisioning program's image is published**, and how a description pins it. A description
names an image by digest; the image is built from the catalogue; the catalogue must therefore both
produce an image and refer to it, which is the same knot the bootstrap has and solves by writing
the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
@@ -1,106 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
extends: 0037-the-host-applies-it-does-not-decide.md
---
# 38. A node joins by linking first, and the mesh finishes the job
## Context
[ADR 0037](0037-the-host-applies-it-does-not-decide.md) settles that the host applies and the
control plane decides. That leaves the case where there is no control plane to decide: the
first node, which must raise a mesh from nothing, and the second, which must join one.
Raised by the operator: *"shouldn't the host have two modes — one for the initial node, setting
up the mesh, so we know the full state; then when adopting a second node, we enter the mesh
early and let our first node take over the mesh-related work? The host should only set up the
bare minimum for the other nodes in the mesh to complete adoption."*
The instinct is right and it is the resolution of the gap ADR 0037 leaves open. The framing
needs one correction, and the correction comes from what the mesh already does.
**Today there are three hand-run shell paths**: `install.d/adopt.sh` (183 lines),
a separate first-node bootstrap, and `install.d/rescue.sh`. The skeleton names the cause —
*"the first node is raised by a special script that exists only because of the circularity"* —
and Move 1 exists to remove it. Three paths that do nearly the same thing, maintained
separately, run by hand, outside anything that checks them.
**That is the two-mode problem, already at its worst.** A decision that gives the host two
modes risks rebuilding `adopt.sh` and `bootstrap.sh` inside the binary, where they will drift
in exactly the same way and be harder to see.
## Considered options
1. **Two modes — genesis and join.** What was proposed. Rejected as a *structure* while adopted
as an *intent*: two modes is two code paths, the first is exercised once per mesh and the
second constantly, so the rarely-run one rots. The current three scripts are the evidence.
2. **One path, and the first node is special-cased inside it.** The conditional moves rather
than disappearing, and now it is scattered instead of named.
3. **One behaviour, two sources of declaration.** Chosen.
## Decision
**The host has one behaviour: apply the declaration it has.** What differs between the first
node and the fiftieth is not what the host *does* but **where the declaration comes from** —
and, exactly as in [ADR 0036](0036-a-node-is-a-managed-machine.md), that is a situation rather
than a class.
| Situation | Declaration comes from |
|---|---|
| no mesh reachable | the pinned bundle the host carries (`substrate.lock`) |
| mesh reachable | the control plane, over the link |
**The first node is not a different kind of node.** It is a node whose mesh is not up *yet*. It
applies the bundle it carries, the control plane comes up on top of it, and from that moment it
takes declarations like everything else. Its specialness is temporary and self-erasing, which
is the property `adopt.sh` and the bootstrap script do not have.
**A joining node does the minimum to be reachable, and nothing else.** It establishes identity
and a route to the control plane — the `link` — and then stops deciding. Everything after that
arrives as declarations.
**The minimum is deliberately small:** an identity, an address, and one peer to reach. A
joining node does **not** compute the overlay. It needs a single peer to reach the mesh; the
full peer set is derived centrally and pushed down, like everything else.
## Why this resolves what 0037 left open
ADR 0037 records that `wireguard` and `traefik` are the two modules that must be split before
they can be absorbed, and that they are the hardest because they need mesh-wide state.
**A joining node never needs that state.** The hard part of the overlay — every node's key,
address, site and endpoint reachability — is only needed to compute the *whole* mesh, which is
the control plane's job. The node needs one peer. The rest arrives.
So the migration ADR 0037 calls expensive is smaller than it looked, and this record is what
makes it smaller.
## Consequences
- **Adoption stops being a script.** The three hand-run paths collapse into the host: joining
is establishing a link, and rescue is a node whose local state is discarded so the mesh can
re-derive it. Whether rescue is fully covered by this is not decided here.
- **The bundle is a fallback, not a mode.** It is what a host applies when nothing better is
available, which also covers a node that has been disconnected for a long time — ADR 0036's
ordinary situation.
- **The rarely-run path is now the common one.** The first node exercises the same code every
other node exercises constantly. That is the whole reason for choosing this over two modes.
- **The link becomes the security boundary.** Everything a node applies arrives through it, so
what may be pushed, and how a joining node proves it is entitled to join, is its own
question — taken up by [ADR 0039](0039-the-link-is-the-security-boundary.md).
- **The bundle must be able to raise the substrate alone.** Whether one host can bring up the
four pinned services with no mesh present is Move 1 of the skeleton and remains unproven.
This record depends on it and does not establish it.
## References
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — the split this completes.
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — situation rather than class, applied here
to the first node.
- [Research 006, Move 1](../01-RESEARCH/006-mesh-from-scratch/skeleton.md) — the pinned bundle,
and the special script it exists to remove.
- [`00-as-is/05-runtime-and-installation.md`](../03-DESIGN/00-as-is/05-runtime-and-installation.md)
— how a node comes into being today.
@@ -0,0 +1,82 @@
---
topic: what runs on it
status: proposed
date: 2026-09-01
deciders: jochen
reconstructed: false
rests-on: 02-DECISIONS/0009-modules-and-the-graph.md
---
# 38. The mesh assigns the port, and a module does not care
## The problem, as met
A database module cannot start on a machine that runs the control plane. The mesh keeps its own
store there and holds 5432; the module publishes 5432. Nothing notices until a container runtime
three layers down says `port is already allocated`
([`028`](../04-ISSUES/028-two-things-want-one-port-and-nothing-says-so/00-report.md)).
A module cannot fix this by choosing better, because **a module cannot know what else is on the
machine.** It is written once and assigned anywhere. Any number it picks is a guess about a
machine it has never seen, and two modules guessing the same number is not a mistake either of
them made.
## The number is written three times, and nothing makes them agree
Every module says its port in three places:
| where | for | example |
|---|---|---|
| `listens` | the rule set that lets traffic in | `{port: 5432, from: mesh}` |
| `serves` | what a consumer must know to connect | `{port: 5432}` |
| a container's `ports` | what the runtime publishes | `"5432:5432"` |
They agree today because one person wrote all three. Nothing checks it. A module whose `serves`
said 5432 and whose container published 5433 would resolve, compose, apply, and hand every
consumer a port that answers nothing.
## The decision
**The mesh assigns the machine-side port, and the module says only what it needs.** A module
declares that a container port must be reachable and what it is for. Which number the machine uses
is the mesh's to choose, because the mesh is the only thing that knows what else is there.
**One source, and the other two are derived.** `serves` carries the assigned port so a consumer is
told where to connect without the module having written it down; the rule set is computed from the
same assignment. Three copies become one fact.
**An assignment is made once and kept**, exactly as a credential is. A port that moved on every
push would restart both ends each time and would hand consumers a number that was true when it was
read.
## Some ports cannot move, and that is a claim
Mail is 25, submission is 587, IMAP over TLS is 993. A mail system on a strange port is not a mail
system. So a module may say a port is **fixed by the protocol** rather than assigned.
**A fixed port is exactly a claim** — the thing the mesh already has for what is singular on a
machine: one seat, one display server, one artifact store. Two modules wanting 25 on one machine is
the same shape as two wanting the seat, and gets the same answer: the second is refused, by name,
when it is assigned rather than when it is applied.
That is why this does not need a new mechanism so much as it needs the existing one pointed at
ports.
## What follows
- **A module becomes portable in a way it was not.** Two databases on one machine stop being a
collision and become two assignments.
- **The substrate has to be visible.** The mesh cannot assign around its own store while it has
never heard of it. What the bundle holds must be written down somewhere the assignment can read
— which the bundle does not say today.
- **A refusal can be useful.** *25 is held by the mail system on this machine* is a sentence a
person can act on. `port is already allocated` is not.
- **`serves` stops being written by hand**, which is a small vocabulary change with a large
consequence: what a consumer is told is now derived from what actually happened.
## What this does not settle
**Whether a module should publish to the machine at all.** Assignment makes publishing safe; it
does not make it necessary. Consumers could instead reach a provider on the module's own network by
name, with nothing published — which would make the question moot for anything inside the mesh, and
would still leave it for anything reached from outside.
@@ -1,190 +0,0 @@
---
status: accepted
date: 2026-08-25
deciders: jochen
reconstructed: false
extends: 0038-a-node-joins-by-linking-first.md
---
# 39. The link is the security boundary
## Context
[ADR 0038](0038-a-node-joins-by-linking-first.md) makes the link the one channel a node takes
declarations from, and names the gap it leaves: *"everything a node applies arrives through it,
so what may be pushed, and how a joining node proves it is entitled to join, is now a question
worth its own record."*
This is that record. It is a design decision about a boundary that does not exist yet — but
what it replaces is measured, and that is the argument.
Settled as: **a node owns no password. It owns an identity, and that identity is what it
presents to the broker.**
### What adoption does today
`install.d/adopt.sh` asks the operator to paste credentials in by hand:
```
The meshware module needs registry database and minio credentials.
REGISTRY_DB_PASSWORD=<postgres password from novox>
REGISTRY_MINIO_PASSWORD=<minio password from novox>
```
plus an `NPM_TOKEN` for the private registry. These are not adoption-time credentials that are
then discarded: `wireguard` and `traefik` open a `pg` connection on every reconcile
([ADR 0037](0037-the-host-applies-it-does-not-decide.md)).
**So every node permanently holds a credential to the control plane's database, and to the
object store.** They are the same credentials on every node. There is no rotation —
[`00-as-is/06`](../03-DESIGN/00-as-is/06-configuration-and-secrets.md) records that *"there is
no mechanism that rotates one and informs everything holding it. Where a rotation has been
done, it has been done by hand, and doing it wrong has taken services down."*
Compromise of any node is therefore compromise of the mesh's database, and there is no
mechanism to recover from it.
### The link is not new
Written first as though the link were a thing to build. It is not.
[ADR 0001](0001-nodes-communicate-over-a-broker.md) already has it: *every node connects
outbound to a single broker; nothing ever connects to a node*, each node declaring an exchange
named for itself and consuming from its own queue
([`00-as-is/01`](../03-DESIGN/00-as-is/01-mesh-and-transport.md)).
That is already outbound-only, already per-node addressed, and already the one channel
everything arrives through. **This record is not proposing a channel. It is proposing that the
channel carry per-node identity instead of one shared credential.**
The same as-is records the fault, for the broker rather than the database: *"the broker is a
single point of failure and a single point of trust. Its credential is mesh-wide, so rotating
it is a mesh-wide operation, and doing it wrong has taken the broker down."*
## Considered options
1. **Keep shared credentials, scope them per node.** Least change: give each node its own
database role. Rejected — it makes the blast radius smaller without changing its shape, and
it keeps tier 0 speaking the control plane's schema, which ADR 0037 forbids for reasons that
are not about security at all.
2. **Accept the exposure as the cost of simplicity.** A shared credential is one thing to
understand and nothing to build, and the objection to replacing it is real: mutual
authentication fails opaquely, and a node that cannot link is harder to debug than a node
with a wrong password. Rejected on the ground that the simplicity is what makes it
unrotatable — the credential cannot be changed *because* everything holds the same one, so
the arrangement's convenience and its unfixability are the same property.
3. **Mutual authority on a node-initiated link, with the node holding nothing but its own
identity.** Chosen.
## Decision
**The link is the only way anything reaches a node**, and four properties make it a boundary
rather than a pipe.
### It is outbound and node-initiated
The node dials the control plane. Nothing dials a node. This is not only defensive — it is what
the topology already requires: most nodes sit behind a household connection with no forwarded
port ([research 004](../01-RESEARCH/004-lab-network/00-overview.md)), so an inbound control
channel would work for the hosted node and not for the rest, and the difference would be
invisible until it mattered.
A node therefore has **no listening control surface at all**.
### A node holds its own identity and nothing else
No shared secret, no credential to anything it does not own. A node's identity authenticates it
to the control plane and grants access to nothing else.
**Compromise of a node is compromise of that node.** That is the property today's arrangement
does not have, and it is the main reason for this record.
### Authority is mutual
The node proves it may join, and **the control plane proves it is the mesh**. One-way is not
enough here: the host applies whatever the link delivers, so a node that cannot tell the mesh
from something impersonating it will apply that something's declarations. Given ADR 0038, an
attacker who can answer a joining node's first call owns the machine.
### What may be pushed is bounded by form, not by trust
The control plane may push **declarations of known shape** and nothing else. It may not push a
command to run. The host's vocabulary is finite, versioned and auditable, and anything outside
it is refused rather than best-effort interpreted.
**Stated honestly: this bounds form, not impact.** A compromised control plane can declare
harmful state — a malicious package, an open firewall — and the host will apply it faithfully,
because that is what it is for. What the property buys is that the blast radius is describable:
it is exactly what the declaration language can express, which can be reviewed. An arbitrary
command channel has no such bound. This is a real limit and not a defence-in-depth story.
### Joining is a deliberate, bounded act
A joining node presents a **one-time, short-lived enrolment token** issued by the mesh for that
purpose, and exchanges it for its own durable identity. The token grants exactly one thing:
the right to become a node. It is not a credential to any service, it does not persist after
exchange, and it expires whether used or not.
This replaces hand-carried shared secrets with a thing that is useless once used and useless
after a while.
## What this actually costs
The objection to weigh is overhead, and it is smaller than it looks because most of it is
already running.
| Property | Where it comes from |
|---|---|
| outbound, node-initiated | already true — ADR 0001 |
| per-node addressing | already true — per-node exchange and queue |
| per-node credential | a broker user per node; the broker already has users, virtual hosts and per-queue permissions |
| mutual authority | transport-level certificates on a connection that already exists |
| bounded by form | already true — three message shapes and only three |
| **enrolment** | **the one genuinely new mechanism** |
And ADR 0037 subtracts rather than adds: under it a node holds **no** database credential at
all, so this record replaces three hand-carried shared secrets with one per-node identity that
grants only identity.
**It must fail legibly.** A boundary that refuses a node without saying why is worse than the
credential it replaced, because a wrong password at least announces itself. A node that cannot
link must report which side rejected it and on what grounds, in terms someone can act on. This
is `how-we-build` §5 applied to a security mechanism: a refusal that proves only that something
went wrong is transport reported as effect.
## Consequences
- **ADR 0037 removes a standing exposure as a side effect.** Its rule — the host never queries
the mesh database — was chosen for tier discipline. It also removes the reason every node
holds the database password. Worth recording because the two arguments are independent and
both hold.
- **Rotation becomes possible and is still not designed.** Per-node identities can be revoked
individually, which is what makes rotation tractable at all. The mechanism —
what rotates, on what trigger, and how holders learn — is **not decided here** and remains
the open weakness `00-as-is/06` records.
- **The enrolment token has to come from somewhere.** Issuing it is a control-plane operation
and the first node has no control plane, so the first node's identity is self-issued and
becomes the root of trust when the mesh comes up. **That is a real asymmetry** — the one
place ADR 0038's "no special first node" does not fully hold — and it is named here rather
than hidden.
- **A declaration vocabulary is now a security artefact, not only a design one.** Every
addition widens what a compromised control plane can express. That is a reason to keep it
small and a reason for additions to be reviewed as such.
- **Offline nodes need identities that survive disconnection.** Per
[ADR 0036](0036-a-node-is-a-managed-machine.md) disconnection is ordinary, so an identity
that must be refreshed to remain valid would make a laptop fail for being a laptop. What
expires and what does not is **not decided here**.
- **This is a boundary that does not exist yet.** Nothing in the current mesh implements any of
it, and the migration from shared credentials to per-node identity touches every node and the
substrate. No estimate is offered.
## References
- [ADR 0038](0038-a-node-joins-by-linking-first.md) — the link, and the gap this fills.
- [ADR 0037](0037-the-host-applies-it-does-not-decide.md) — why the host stops holding database
credentials at all.
- [ADR 0036](0036-a-node-is-a-managed-machine.md) — disconnection as ordinary, which constrains
what may expire.
- [`00-as-is/06-configuration-and-secrets.md`](../03-DESIGN/00-as-is/06-configuration-and-secrets.md)
— secrets today, and the absence of rotation.
- [Research 004](../01-RESEARCH/004-lab-network/00-overview.md) — why most nodes cannot accept
an inbound connection.
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
---
# 39. What the SDK holds, and what it refuses
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
The earlier "repository structure" decision (folded in consolidation; see the reconciliation note
above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor)
named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- The earlier "repository structure" decision — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record (folded in consolidation;
nearest survivor [ADR 0015](0015-applications-live-in-their-own-repository.md)).
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — the decomposition this serves: code
belongs to the boundary that owns it.
- [ADR 0005](0005-the-node-host.md) — why the host mirrors the contracts instead of
importing the SDK.
+101
View File
@@ -0,0 +1,101 @@
---
topic: what runs on it
status: accepted
date: 2026-09-03
deciders: jochen
reconstructed: false
extends: 0009-modules-and-the-graph.md
---
# 40. What a module is
_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._
## Context
[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note
above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1. **Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2. **Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3. **Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
([research 005](../01-RESEARCH/005-domain-grouping/analysis.md)).
This is [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md)'s rule made general — "names the protocol, not
the product; a database names the engine because the app targets it" — with the reason stated: the
contract sits where the coupling is.
### What is not a module
- A **library** (built against, never deployed — [ADR 0039](0039-what-the-sdk-holds-and-refuses.md)).
- A **control-plane context** (the mesh itself — [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md)).
A **swappable machine mechanism** (a firewall — ufw, nftables) *is* a module implementing a
capability. The host hardcodes no firewall, supervisor, package manager or runtime; it owns only the
generic apply primitives and platform detection, so it runs where none of those exist — an Android
phone has no ufw, systemd, pacman or Docker.
## Consequences
- **Supersedes the earlier "grouped by domain" decision** (folded in consolidation; see the
reconciliation note above). Modules are
organised by their relationships (seats, provisions), not grouped into domain folders.
- **Refines [ADR 0009](0009-modules-and-the-graph.md).** Everything the mesh runs and integrates is
a module — but a module is defined by the *software it delivers*, not by being a bucket of features.
- The target is a **self-fulfilling mesh**: declared wants bound to swappable modules, provisioners
wiring credentials, nothing hardcoded. The control plane's whole job is the binding.
- Converting a module from the old system includes pulling its per-module code out of the shared SDK
([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)) and shipping its provisioner as an adapter to a
mesh interface — a de-coupling, not just a move.
## References
- [ADR 0009](0009-modules-and-the-graph.md) — everything is a module; this says what one is.
- [ADR 0001](0001-mesh-brokers-nodes-host-agents-think.md) — contexts are the mesh, not modules.
- The earlier "grouped by domain" decision — superseded (folded in consolidation; see the note above).
- [ADR 0027](0027-a-provision-names-what-the-consumer-is-coupled-to.md) — protocol-not-product, generalised here.
- [ADR 0039](0039-what-the-sdk-holds-and-refuses.md) — per-module code lives in the module.
- [research 011](../01-RESEARCH/011-the-module-graph/00-overview.md) — the graph of these relationships.

Some files were not shown because too many files have changed in this diff Show More